Drawing method and related equipment
By obtaining element description information in images and using artificial intelligence technology to generate maps, the problem of computer resource waste is solved and more efficient and accurate mapping is achieved.
Patent Information
- Application Number
- CN202410387220.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-30
- Publication Date
- 2025-10-10
AI Technical Summary
During the map construction process, the computer resources required for image stitching are too much, resulting in serious waste of computer resources.
By obtaining element description information in the image and using artificial intelligence technology to generate maps, computer resource consumption is reduced.
It reduces the consumption of computer resources in the mapping process and improves the timeliness and accuracy of mapping.
Smart Images

Figure CN120765797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer technology, and more particularly to a mapping method and related equipment. Background Art
[0002] Currently, when building a map, multiple images of an environment are collected and stitched together to create a basemap. Based on this basemap, a map is then generated that reflects the elements within the environment. However, because images contain rich information and therefore require large amounts of data, stitching together multiple images to create a basemap consumes significant computing power. Consequently, building a map using this approach requires significant computer resources. Summary of the Invention
[0003] The present application provides a mapping method and related equipment, which can greatly reduce the computer resources consumed in the mapping process.
[0004] This application provides the following technical solutions:
[0005] In a first aspect, the present application provides a mapping method that can use artificial intelligence technology in application scenarios that require mapping. In this method, an execution device obtains at least one set of images of a first environment, and each of the at least one set of images includes at least one image of the first environment; the execution device can obtain first description information of each element in the first environment based on the at least one image of the first environment included in each set of images, and then determine a map corresponding to all elements in the first environment based on the first description information of each element in the first environment. Exemplarily, the execution device can be an image acquisition device located on the edge side, such as an image acquisition device on the side of a vehicle, drone, robot, handheld device, or other type of terminal, etc., or the execution device can also be a server that is communicatively connected to the image acquisition device.
[0006] Exemplarily, each first description information can indicate the category and position of an element in the first environment. The elements in the first environment can be understood as independent objects in the first environment. When a certain first description information is the first description information of a line element in the first environment, the first description information of the line element is also used to indicate the shape of the line element. For example, the shape of the line element may include the direction, inclination or other shape information of the line element.
[0007] In this implementation, after obtaining at least one group of images of the first environment, first description information of the elements in the first environment is first obtained based on each group of images of the first environment. The first description information indicates the category and position of the elements in the first environment. For line elements existing in the first environment, the first description information also indicates the shape of the line elements, that is, the first description information includes key information related to the composition obtained from the image, and then a map corresponding to the elements in the first environment is generated based on the first description information obtained from each group of images of the first environment. Since the amount of data of the first description information obtained from at least one group of images of the first environment will be much smaller than the amount of data of the aforementioned at least one group of images, the method of generating a map based on the obtained first description information can greatly reduce the computer resources consumed by the mapping process; in addition, after the computer resources consumed by the mapping process are reduced, it is conducive to deploying the mapping method provided by the present application on the terminal side, so that the terminal side can realize the mapping of elements in the environment more timely, which is conducive to obtaining a more accurate map in a timely manner.
[0008] In one possible implementation, if the first environment is a road environment, the elements in the first environment include road elements. For example, the road elements may include at least one of the following: lane lines, road lines, road surface markings, traffic lights, poles, road signs or other road elements, etc. The map corresponding to the elements in the first environment is a map containing road elements.
[0009] If the first environment is a shopping mall (that is, an example of an indoor environment), the elements in the first environment may include stores in the shopping mall and roads in the shopping mall, and the “map corresponding to the elements in the first environment” may be expressed as a navigation map of the shopping mall. If the first environment is an environment where power lines are located, the elements in the first environment may include power lines and telephone poles, and the “map corresponding to the elements in the first environment” may be expressed as a map containing the power lines. If the first environment is a construction site, the elements in the first environment may include buildings, and the “map corresponding to the elements in the first environment” may be expressed as a map corresponding to the buildings in the construction site. If the first environment is an environment where a forest is located, the elements in the first environment may include trees, and the “map corresponding to the elements in the first environment” may be expressed as a map containing trees in the forest. In this implementation method, a specific application scenario of the present solution is provided, which improves the degree of integration between the present solution and the specific application scenario.
[0010] In one possible implementation, if the first environment is a road environment, the road elements in the road environment may include a first road element and a second road element. The first road element refers to a line-shaped road element in the road environment, and the second road element is a road element different from the first road element in the road environment. If each piece of first description information is expressed in a vector format to indicate the position (optionally, the shape) of an element in the first environment, the first description information of the first road element may include a function expression indicating the position and shape of the first road element and the category of the first road element. The first description information of the second road element may include the coordinates of a point indicating the position of the second road element and the category of the second road element.
[0011] For example, the aforementioned function expression may be a polynomial. For example, if a first road element is a straight line, a linear polynomial may be used to indicate the position and shape of the first road element. For another example, if a first road element is an arc, a quadratic polynomial, a cubic polynomial, a quartic polynomial, or other types of polynomials may be used to indicate the position and shape of the first road element. The specific polynomial to be used may be determined based on the specific application scenario.
[0012] Optionally, the first description information may also include other attribute information of the road element. If a road element is a lane line, the first description information of the lane line may further indicate at least one of the following information: whether the lane line is solid or dashed, whether the lane line is single-line or double-line, the color of the lane line, or other attribute information. If a road element is a road sign, the first description information of the road sign may further indicate what information the road sign indicates. For example, if the road sign is a left-pointing arrow, the first description information of the road sign may also indicate that a left turn is permitted. For another example, if the road sign is 30, the first description information of the road sign may also indicate a speed limit of 30 km / h, etc. If a road element is a road sign, the first description information of the road sign may also indicate what information the road sign indicates, etc.
[0013] In this implementation, a function expression is used to indicate the position and shape of line-shaped road elements in the road environment, and the coordinates of a point are used to indicate the position of other road elements other than line shapes. Since the function expression can not only accurately reflect the position and shape of the road elements, but also consumes very few computer resources, the coordinates of the point can correspondingly accurately reflect the position of the road elements, and the coordinates of the point consume very few computer resources. That is, it can not only accurately describe the road elements in the road environment, which is conducive to generating accurate maps in subsequent steps, but also greatly reduces the consumption of computer resources.
[0014] In one possible implementation, the execution device obtains first description information of elements in the first environment based on at least one image of the first environment included in each group of images, which may include: the execution device obtains M subsets based on at least one image of the first environment included in each group of images, and then obtains M groups of description information based on the M subsets.
[0015] Each of the M subsets includes at least one image of the first environment, and M is an integer greater than or equal to 2; there may be intersections between different subsets in the M subsets, or there may be no intersection at all between different subsets in the M subsets.
[0016] Each group of description information in the M groups of description information includes a first description information of at least one element in the first environment, that is, each group of description information in the M groups of description information includes at least one first description information. Since the description information of different groups in the M groups of description information may include the first description information of the same element (hereinafter referred to as the "target element" for the convenience of description), that is, the M groups of description information can include at least two first description information of the same target element. The at least two first description information of the target element come from different groups in the M groups of description information, and at least one of the aforementioned target elements can exist in the first environment.
[0017] The execution device determines the map corresponding to the element in the first environment based on the first description information, which may include: the execution device generates second description information for each target element based on at least two first description information of each target element, and then determines the map corresponding to the element in the first environment based on the second description information of each target element; wherein the second description information indicates the category and position of the target element, and when the target element is a line element, the second description information also indicates the shape of the target element.
[0018] Exemplarily, if the target element is a line element in the first environment, after the execution device obtains at least two first description information corresponding to the same line element, that is, obtains at least two lines based on the aforementioned at least two first description information, it can sample each of the at least two lines to obtain at least two discrete point sets corresponding one-to-one to the at least two lines, and perform spatial clustering on the at least two discrete point sets. Since errors may occur when the execution device assigns identification information to the elements in the first environment, after performing spatial clustering on the at least two discrete point sets, the outlier discrete point sets can be eliminated, and then new lines can be fitted based on the remaining discrete points. The second description information can be determined based on the new lines. The second description information includes the position and shape of the new line, that is, the position and form of the target element. In addition, the execution device can also retain the category of the target element carried by the first description information of the target element (optionally, other attribute information) in the second description information of the target element, so that the second description information of the target element also indicates the category of the target element (optionally, other attribute information).
[0019] If the target element is an element in the first environment that uses the coordinates of a point to indicate its position, then after obtaining at least two first description information of the same target element, the execution device determines at least two groups of points corresponding to the same target element based on the aforementioned at least two first description information, each group of points including at least one point. The execution device can directly eliminate outlier points from the at least two groups of points, and then determine the final coordinates of the point corresponding to the target element based on the remaining points (that is, the coordinates of the point included in the second description information of the target element); for example, a weighted average can be performed based on the coordinates of the remaining points to obtain the final coordinates of the point, or the aforementioned "weighted average" can also be replaced by finding the median or other calculation methods.
[0020] Since each image of the first environment obtained is easily affected by the acquisition position and acquisition angle, there is inconsistency between frames, that is, different first description information of the same element in the first environment may be obtained based on different images. In this implementation method, each group of images of the first environment is divided into M subsets, and a group of description information is obtained using each subset of the M subsets. Then, at least two first description information of the same target element in the first environment can be obtained from the M groups of description information, and the second description information of the target element is generated based on the at least two first description information of the target element. This is conducive to eliminating the impact of inconsistency between frames, and is conducive to making the position information carried by the second description information of the element in the first environment more accurate, and thus is conducive to obtaining a more accurate map.
[0021] In a possible implementation, the number of the at least one group is N groups, N is an integer greater than or equal to 2, and the N groups of images of the first environment are obtained by performing image acquisition on the first environment at least twice, and there are different trajectories in at least two image acquisition trajectories corresponding to the at least twice image acquisition. For example, if the N groups of images are collected by a vehicle, the "image acquisition trajectory" can be understood as a driving trajectory of the vehicle; if the N groups of images are collected by a robot, the "image acquisition trajectory" can be understood as a moving trajectory of the robot; if the N groups of images are collected by a drone, the "image acquisition trajectory" can be understood as a flight trajectory of the drone; if the N groups of images are collected by a handheld device, the "image acquisition trajectory" can be understood as a moving trajectory of the handheld device, and the like.
[0022] There is an image acquisition trajectory when each group of images of the road environment is collected, and the map obtained based on the images collected once can still be affected by the collection angle. In the implementation, the number of the at least one group of images is N, and N is an integer greater than or equal to 2, that is, the first environment is collected by using different image acquisition trajectories, and then the final map is obtained according to the at least two groups of images, which is beneficial to eliminate errors in the single observation process, and is beneficial to make the final obtained map more accurately reflect the real environment, that is, to obtain a more accurate map.
[0023] In a possible implementation, the first environment is a road environment, and the elements in the first environment include road elements. The execution device determines the map corresponding to the elements in the first environment according to the first description information, which can include: the execution device determines one map information corresponding to the road elements according to the first description information corresponding to each group of images in the N groups of images, that is, N map information can be obtained based on the N groups of images, and the N groups of images correspond to the N map information one by one. For any one of the N map information (for convenience of description, hereinafter referred to as "first map information"), the execution device can divide the map space indicated by the first map information into a plurality of spatial regions based on the first image acquisition trajectory, and obtain third description information of at least one instance in each spatial region of the plurality of spatial regions according to the first map information, that is, third description information of a plurality of instances in the plurality of spatial regions can be obtained.
[0024] The first map information corresponds to the first image acquisition trajectory in the at least two image acquisition trajectories, that is, the first image acquisition trajectory is the image acquisition trajectory of the group of images used to obtain the first map information. Optionally, there can be an overlapping region between different spatial regions in the plurality of spatial regions. Alternatively, there can be no overlapping region between different spatial regions in the plurality of spatial regions.
[0025] For example, the instances in this application can also be understood as "rigid body instances", that is, when fusing N map information, each rigid body instance is regarded as a minimum adjustment unit, that is, in this implementation, the fusion is performed at the granularity of "instance". Particularly, since after the first map information is divided into multiple spatial regions, the multiple road elements in the first map information are also divided into respective spatial regions, an instance can represent a road element in one of the multiple spatial regions (for example, representing a second road element); or, since a linear road element may run through multiple spatial regions, an instance can also represent a partial line segment of a linear road element (that is, the first road element) in a spatial region.
[0026] The third description information indicates the category and location of the instance. When the instance represents a partial line segment of a line-shaped road element within a spatial area, the third description information also indicates the shape of the instance. Optionally, the third description information also indicates other attribute information of the instance.
[0027] After obtaining the third description information of multiple instances in multiple spatial areas, the execution device can determine the final position information corresponding to the first instance based on at least two third description information corresponding to the first instance, and the final position information corresponding to the multiple instances in multiple spatial areas is used to obtain a map corresponding to the elements in the first environment.
[0028] In which, the multiple instances in the multiple spatial regions include a first instance, and each of the at least two third description information corresponding to the first instance is derived from one of the at least two map information. Exemplarily, since the N map information is obtained based on N sets of images of the same road environment, that is, the N map information reflects the same road environment, the N map information can be understood as N observation results of the same road environment. Therefore, among the multiple instances (also referred to as rigid body instances) obtained based on the N map information, there are N instances corresponding to the same second road element (or the same line segment of the same first road element) in the road environment, that is, the same second road element (or the same line segment of the same first road element) in the road environment has N third description information. "At least two third description information corresponding to the first instance" can be understood as at least two third description information corresponding to the same second road element (or the same line segment of the first road element) in the road environment. In detail, "fusing the at least two third description information corresponding to each first instance" can be understood as fusing the multiple third description information corresponding to each second road element in the road environment, and fusing the multiple third description information corresponding to each line segment of each first road element.
[0029] Exemplarily, after obtaining the final positions of the plurality of line segments, the execution device can connect the plurality of discrete line segments into a continuous line; optionally, the execution device can also perform smoothing processing on the continuous line; and then the execution device can determine the position and shape of the second road element in the road environment, and the second map corresponding to the road element in the road environment can include the position and shape of the second road element in the road environment, and can also include the category of the second road element. The second map corresponding to the road element in the road environment can also include the final position information and the category of the first road element.
[0030] In the present implementation, for a line-shaped road element (such as a lane line, a road line, etc.), since the line-shaped road element is generally long, it is difficult to directly adjust the line-shaped road element at the granularity of the entire line to achieve perfect fitting between the final position information of the line-shaped road element and the actual position information thereof. In the present application, the entire line is divided into a plurality of line segments, and then fusion is performed at the granularity of the line segment, which is beneficial to make the fitting more accurate, thereby being beneficial to obtain more accurate positions, i.e., to obtain a more accurate map.
[0031] In a second aspect, the present application provides a mapping device, which can use artificial intelligence technology in an application scenario requiring mapping. The mapping device comprises: an acquisition module configured to acquire at least one set of images of a first environment, each set of images comprising at least one image of the first environment; and the acquisition module is further configured to acquire first description information of an element in the first environment according to the at least one image of the first environment included in each set of images, wherein the first description information indicates the category and position of the element in the first environment, and in the case where there is a line element in the first environment, the first description information of the line element further indicates the shape of the line element; and a determination module configured to determine a map corresponding to the element in the first environment according to the first description information.
[0032] In the second aspect of the present application, the mapping device can also perform the steps performed by the execution device in the various possible implementation manners of the first aspect. For the meanings of the terms in the second aspect of the present application and the various possible implementation manners of the second aspect, and the beneficial effects brought by each possible implementation manner, reference can be made to the descriptions in the various possible implementation manners of the first aspect, which will not be repeated here.
[0033] In a third aspect, the present application provides a device, which comprises a processor and a memory, the processor is coupled to the memory, the memory is configured to store a program, and the processor is configured to execute the program in the memory, so that the device performs the method of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a vehicle comprising a processor and a memory, wherein the processor is coupled to the memory, the memory being used to store programs; and the processor being used to execute the programs in the memory, so that the vehicle executes the method described in the first aspect above.
[0035] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first aspect above.
[0036] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes a program. When the program runs on a computer, it enables the computer to execute the method described in the first aspect above.
[0037] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the terminal device or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A schematic diagram of the structure of the artificial intelligence main framework provided in the embodiment of the present application;
[0039] Figure 2 An architectural diagram of the mapping system provided for this application;
[0040] Figure 3 A schematic diagram of a flow chart of a mapping method provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram provided in an embodiment of the present application that uses vector expression to indicate elements in a first environment;
[0042] Figure 5 A schematic diagram of inter-frame inconsistency provided in an embodiment of the present application;
[0043] Figure 6 A schematic diagram of multiple spatial regions provided in an embodiment of the present application;
[0044] Figure 7 A schematic diagram of the center points of multiple examples provided in the embodiments of the present application;
[0045] Figure 8 A schematic diagram of a graph created when optimizing using a graph optimization algorithm provided in an embodiment of the present application;
[0046] Figure 9 A comparative diagram before and after deduplication provided in an embodiment of the present application;
[0047] Figure 10 A schematic diagram of two stacked map information provided in an embodiment of the present application;
[0048] Figure 11 A schematic diagram of the structure of a mapping device provided in an embodiment of the present application;
[0049] Figure 12 A schematic diagram of the structure of the device provided in the embodiment of the present application;
[0050] Figure 13 A schematic structural diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all of the embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0052] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances. This is merely a way of distinguishing when describing objects with the same properties in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, so that a process, method, system, product or apparatus that includes a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to these processes, methods, products or apparatuses.
[0053] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information (such as the indication information described below) is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated; it is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance, for example, the indication of specific information can be achieved with the help of the arrangement order of each information agreed in advance (such as predefined by the protocol), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific method of indication. It is understandable that, for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.
[0054] First, the overall workflow of the artificial intelligence system is described. Figure 1 , Figure 1 A structural diagram of the artificial intelligence main framework provided for the embodiment of the present application is provided below. The above artificial intelligence theme framework is elaborated from the two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, the data undergoes a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.
[0055] (1) Infrastructure
[0056] The infrastructure provides computing power support for artificial intelligence systems, enabling communication with the outside world and providing support through the basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically adopt hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs); the basic platform includes related platform guarantees and support such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to obtain data, and this data is provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0057] (2) Data
[0058] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0059] (3) Data processing
[0060] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0061] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0062] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0063] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0064] (4) General ability
[0065] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0066] (5) Smart products and industry applications
[0067] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart manufacturing, smart transportation, smart homes, smart medical care, smart security, smart driving, smart cities, etc.
[0068] The method provided in this application can be applied to various application scenarios that require mapping. For example, in intelligent driving, a map corresponding to the road environment needs to be deployed in the vehicle, and then the map can be used to provide navigation functions, or the map can also be used to provide assisted driving functions, etc., and it is necessary to generate a map that can reflect the road elements in the road environment.
[0069] For example, in the field of intelligent manufacturing, robots can be used to survey construction sites to generate maps corresponding to the buildings on the construction site. This map can be used to reflect information such as the height and shape of the buildings under construction, thereby facilitating the monitoring of the construction progress of the construction site.
[0070] For example, in the smart city sector, vehicles can be used to generate maps of outdoor road environments, while robots and / or handheld devices can be used to perform indoor mapping, generating maps corresponding to the indoor environment. This allows for seamless indoor and outdoor mapping. Alternatively, mapping can be performed inside shopping malls to facilitate navigation within the mall.
[0071] For example, in the field of smart terminals, drones can be used to map power lines so that they can be repaired based on the maps corresponding to the power lines; or, drones + mechanical vehicles can be used to map forests to obtain maps corresponding to the forests. This map is used to reflect information such as the number, density, and height of trees in the forest to facilitate the management of trees in the forest; or, drones can be used to map mines to obtain maps corresponding to the mines, and so on.
[0072] In all the above application scenarios, it is necessary to map the elements in the environment to obtain a map corresponding to the elements in the environment. The current mapping method is often to collect multiple images of the environment, stitch the multiple images together to obtain a base map, and then generate a map corresponding to the elements in the environment based on the base map. However, since the process of "stitching the multiple images of the environment to obtain the base map" consumes a lot of computing power, the above method of mapping requires a lot of computer resources. To solve the above problem, the present application discloses that after acquiring at least one set of images of a first environment, the execution device can first obtain first description information of the elements in the first environment based on at least one image of the first environment included in each set of the at least one set of images. The first description information is used to indicate the category and position of each element in the first environment. If there are line elements (i.e., line-shaped elements) in the first environment, the first description information of the line elements can also indicate the shape of the line elements (e.g., the inclination of the line). That is, the first description information includes key information related to the composition obtained from the image; and then the execution device can generate a map corresponding to the elements in the first environment based on all the obtained first description information. Since the data volume of the first description information obtained from at least one group of images of the first environment will be much smaller than the data volume of the aforementioned at least one group of images, the method of generating a map based on the obtained first description information can greatly reduce the computer resources consumed by the mapping process.
[0073] Optionally, since a machine learning model may be used in the specific implementation process of the mapping method provided in this application, before describing the specific implementation process of the method provided in this application in detail, please refer to Figure 2 , Figure 2 An architectural diagram of the mapping system provided in this application, such as Figure 2 As shown, the mapping system 200 includes a training device 210 , a database 220 , an execution device 230 and a data storage system 240 , and the execution device 230 includes a calculation module 231 .
[0074] The database 220 stores a training data set. During the training phase of the machine learning model 201, the training device 210 generates the machine learning model 201 and iteratively trains the machine learning model 201 using the training data set, thereby obtaining a trained machine learning model 201. The "trained machine learning model 201" may also be referred to as the "trained machine learning model 201." The machine learning model 201 may be specifically represented by a neural network or a non-neural network model.
[0075] The machine learning model 201 obtained by the training device 210 after the training operation can be deployed to the computing module 231 of the execution device 230. The execution device 230 can call data, code, etc. in the data storage system 240, or store data, instructions, etc. in the data storage system 240. The data storage system 240 can be located in the execution device 230, or the data storage system 240 can be an external memory relative to the execution device 230.
[0076] During the application stage of the machine learning model 201, the execution device 230 can determine whether the vehicle deviates when passing a fork in the road based on the image or point cloud data of the vehicle's surrounding environment through the trained machine learning model 201.
[0077] In some embodiments of this application, please refer to Figure 2 , the execution device 230 and the client device can be integrated into the same device, so that the user can directly interact with the execution device 230. For example, when the client device is a vehicle, the execution device 230 can be a module in the vehicle's host processor (Host CPU) that uses a machine learning model to perform data processing. The execution device 230 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the vehicle. The GPU or NPU is mounted on the host processor as a coprocessor, and the host processor assigns tasks.
[0078] It is worth noting that Figure 2 This is only a schematic diagram of the architecture of the mapping system provided by an embodiment of the present invention, and the positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in other embodiments of the present application, the execution device 230 and the client device can be independent devices. The execution device 230 is configured with an input / output (I / O) interface to exchange data with the client device. The client device sends at least one set of images of the first environment to the execution device 230 via the I / O interface. After generating a map corresponding to the elements in the first environment, the execution device 230 can return the results of the yaw detection to the client device via the I / O interface. The machine learning model 201 in the computing module 231 is used in the process of generating the aforementioned map.
[0079] For details, please refer to 3. Figure 3 A schematic diagram of a flowchart of a mapping method provided in an embodiment of the present application. The mapping method provided in an embodiment of the present application may include:
[0080] 301. Acquire at least one set of images of a first environment, where each set of images in the at least one set of images includes at least one image of the first environment.
[0081] In an embodiment of the present application, the execution device can obtain at least one set of images of the first environment; illustratively, in one case, the execution device can be a terminal-side device, for example, the execution device and the image acquisition device can be the same device, then the aforementioned at least one set of images can be acquired by the execution device, for example, the execution device (that is, the image acquisition device) can be a vehicle, drone, robot, handheld device or other type of terminal-side image acquisition device, etc.; alternatively, the execution device can also be a server, then the execution device can receive the aforementioned at least one set of images sent by the image acquisition device.
[0082] For example, the first environment may be a road environment, an indoor environment, a construction site, an environment where power lines are located, an environment where forests are located, an environment where mines are located, or other types of environments, etc. The specific environment of the "first environment" may be determined based on the actual application scenario.
[0083] Optionally, at least one image included in each of the above-mentioned groups of images may include multiple images of the first environment captured sequentially; for example, if the first environment is a road environment, multiple images of the road environment may be captured while the vehicle is driving in the road environment; for another example, if the first environment is an indoor environment, multiple images of the indoor environment may be captured while the mobile robot is moving in the indoor environment; for another example, if the first environment is an environment where power lines are located, a drone may be used to fly in an environment where power lines are deployed, and multiple images of the environment where the power lines are located may be captured during the flight, and so on. The specific details may be determined in combination with the actual application scenario, and examples will not be given one by one here.
[0084] Optionally, the number of the at least one group is N groups, where N is an integer greater than or equal to 2, and the N groups of images of the first environment are obtained by performing at least two image capture operations on the first environment, and there are different trajectories in at least two image capture trajectories corresponding to the at least two image capture operations.
[0085] For example, if N groups of images are collected by a vehicle, the "image collection trajectory" can be understood as the vehicle's driving trajectory; if N groups of images are collected by a robot, the "image collection trajectory" can be understood as the robot's moving trajectory; if N groups of images are collected by a drone, the "image collection trajectory" can be understood as the drone's flight trajectory; if N groups of images are collected using a handheld device, the "image collection trajectory" can be understood as the moving trajectory of the handheld device, and so on. It should be noted that the specific meaning of the "image collection trajectory" can be determined in combination with the actual application scenario. The examples here are only for the convenience of understanding this solution and are not used to limit this solution.
[0086] Among them, the image acquisition device (that is, the above-mentioned vehicle, drone, handheld device) can be deployed with a positioning system, and the positioning system can be used to determine the image acquisition trajectory. For example, the positioning system can be a global positioning system (GPS), or for example, the positioning system can be a Beidou positioning system, or other types of positioning systems, etc.
[0087] Furthermore, the image acquisition trajectory may include the posture information of the image acquisition device at the image acquisition point corresponding to each image. For example, the posture information may include the displacement of the image acquisition device on the x-axis, the displacement of the y-axis, and the displacement of the z-axis; optionally, the posture information may also include the rotation angle of the image acquisition device around the x-axis, the rotation angle around the y-axis, and the rotation angle around the z-axis.
[0088] 302. Based on at least one image of the first environment included in each group of images, obtain first description information of elements in the first environment, wherein the first description information indicates the category and position of the elements in the first environment, and the first environment includes line elements, and the first description information of the line elements further indicates the shape of the line elements.
[0089] In an embodiment of the present application, the elements in the first environment can be understood as independent objects in the first environment; for example, if the first environment is a road environment, the elements in the first environment may include road elements, for example, road elements may include at least one of the following: lane lines, road lines, road surface markings, traffic lights, poles, road signs or other road elements, etc., wherein the road surface markings in this application may be lane lines and ground markings other than road lines, for example, an arrow on the road surface for indicating a left turn, a sign on the road surface for indicating a right turn, a number on the road surface for indicating a speed limit, an arrow-shaped diversion sign or other types of road surface markings, etc., and the examples are not exhaustive here.
[0090] If the first environment is a shopping mall (i.e., an example of an indoor environment), the elements in the first environment may include shops in the shopping mall and roads in the shopping mall. If the first environment is an environment where power lines are located, the elements in the first environment may include power lines and utility poles. If the first environment is a construction site, the elements in the first environment may include buildings. If the first environment is an environment where a forest is located, the elements in the first environment may include trees, etc. The specific expression of "elements in the first environment" can be determined in combination with actual application scenarios, and will not be repeated in detail in the embodiments of this application.
[0091] Optionally, each first description information may be in the form of a vector expression to indicate the position (optionally, also including the shape) of an element in the first environment. For example, when the first environment is a road environment, the road elements in the road environment may include at least one first road element and at least one second road element; wherein the first road element represents a line-shaped road element, such as a road line, a lane line, or other line-shaped road element; and the second road element is an element in the road environment that is different from the first road element, such as a road marking, a traffic light, a pole, a road sign, or other road element.
[0092] For any first road element (a line-shaped road element) in the first environment, a function expression can be used to indicate the position and shape of the first road element. That is, the first description information of each first road element can include the function expression and the category of the first road element. For example, a function expression can be used to indicate the position and shape of a lane line, where the shape of the lane line may include its direction, inclination, or other shape information. Exemplarily, the function expression can be a polynomial. For example, if a first road element is a straight line, a linear polynomial can be used to indicate the position and shape of the first road element. For another example, if a first road element is an arc, a quadratic polynomial, a cubic polynomial, a quartic polynomial, or other type of polynomial can be used to represent the position and shape of the first road element. The specific type of polynomial to be used can be determined based on the specific application scenario.
[0093] For any second road element in the first environment, the coordinates of a point can be used to indicate the location of the second road element. That is, the first description information of each second road element includes the coordinates of the point and the category of the second road element. The aforementioned coordinates can be three-dimensional or two-dimensional, and the specific coordinates can be determined based on actual circumstances. For example, the coordinates of the center point of a road sign can be used to indicate the location of the road sign, or points on the outline of the road sign can be used to indicate the location of the road sign; the coordinates of the center point of a traffic light can be used to indicate the location of the traffic light; the coordinates of the center point of a road sign can be used to indicate the location of the road sign, or multiple points within the road sign can be used to indicate the location of the road sign; the coordinates of the highest point of a pole can be used to indicate the location of the pole, and so on. The examples provided here are provided for the purpose of facilitating understanding of this solution and are not intended to limit this solution.
[0094] For example, the category of a road element may be a lane boundary line, a lane center line, a road line, a road marking, a traffic light, a pole, a road sign, or other categories, etc., which may be determined based on actual conditions.
[0095] For a more intuitive understanding of this solution, please refer to Figure 4 , Figure 4A schematic diagram of an embodiment of the present application using vector expression to indicate elements in a first environment is provided. Figure 4 After visualizing each first description information, Figure 4 Taking the first environment as a road environment as an example, the first description information may indicate three-dimensional spatial position information, wherein: Figure 4 The dot matrix in the figure represents the traffic lights and road signs in the road environment. Figure 4 The isolated points in the figure represent the poles in the road environment. In order to distinguish the isolated points from the dot matrix, Figure 4 The isolated points are circled with white circles, and Figure 4 The points not surrounded by white circles are all dot matrix. Figure 4 The vertical lines in the middle represent lane lines and road lines in the road environment. Figure 4 The horizontal lines in the middle represent road markings. It should be understood that Figure 4 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0096] In the embodiment of the present application, a function expression is used to indicate the position and shape of line-shaped road elements in a road environment, and the coordinates of a point are used to indicate the position of other road elements other than line shapes. Since the function expression can not only accurately reflect the position and shape of the road elements, but also consumes very few computer resources, the coordinates of the point can correspondingly accurately reflect the position of the road elements, and the coordinates of the point consume very few computer resources. That is, it can not only accurately describe the road elements in the road environment, which is conducive to generating accurate maps in subsequent steps, but also greatly reduces the consumption of computer resources.
[0097] For another example, if the first environment is a shopping mall, the coordinates of a point can be used to indicate the location of a store in the mall, the roads in the mall can be regarded as line elements, and function expressions can be used to indicate the location and shape of the roads in the mall; the categories of elements in the first environment may include stores and roads. For another example, if the first environment is an environment where power lines are located, the power lines can be regarded as line elements, and function expressions can be used to indicate the location and shape of the power lines, and the coordinates of a point can be used to indicate the location of the utility poles; the categories of elements in the first environment may include power lines and utility poles. For another example, if the first environment is a construction site, function expressions can be used to indicate the shape and height of a building, for example, the building can be regarded as a polygon, and the shape and height of the building can be indicated by a set of polynomials; the categories of elements in the first environment may include buildings. For another example, if the first environment is an environment where a forest is located, the height of the tree can be regarded as a line element, and the shape of the tree (including the height and the growth direction of the tree) can be indicated by a function expression, and the coordinates of a point can be used to indicate the location of the tree; the categories of elements in the first environment may include trees, etc. It should be noted that the examples here are only used to prove the feasibility of this solution in various application scenarios.
[0098] Optionally, the first description information may also include other attribute information of elements in the first environment. For example, if the first environment is a road environment, when a certain road element is a lane line, the first description information of the lane line may also indicate at least one of the following information: whether the lane line is a solid line or a dashed line, whether the lane line is a single line or a double line, the color of the lane line, or other attribute information. When a certain road element is a road sign, the first description information of a road sign may also indicate what information the road sign indicates. For example, if the road sign is a left arrow, the first description information of the road sign may also indicate that a left turn is allowed. For another example, if the road sign is 30, the first description information of the road sign may also indicate a speed limit of 30KM / h, etc. When a certain road element is a road sign, the first description information of the road sign may also indicate what information the road sign indicates, etc. The examples given here are only for the convenience of understanding this solution. The specific attribute information included may be determined in combination with the actual needs in the actual application scenario.
[0099] Exemplarily, in one implementation, step 302 may include: the execution device obtains M subsets based on at least one image of the first environment included in each group of images, each subset of the M subsets includes at least one image of the first environment, M is an integer greater than or equal to 2, for example, the value of M can be 2, 3, 4, 5 or other values, etc., which are not exhaustive here.
[0100] Optionally, different subsets within the M subsets may intersect. For example, a set of images includes image 1, image 2, image 3, image 4, image 5, image 6, image 7, image 8, and image 9, all captured sequentially from a first environment. Three subsets are obtained based on the aforementioned nine images: subset 1 includes images 1 through 5, subset 2 includes images 3 through 7, and subset 3 includes images 5 through 9. It should be understood that the aforementioned examples are merely for facilitating understanding of this solution and are not intended to limit this solution. Alternatively, different subsets within the M subsets may not intersect at all, a specific determination of which may be made based on actual circumstances.
[0101] Then, the execution device can obtain a set of descriptive information based on each of the M subsets, that is, it can obtain M sets of descriptive information based on the M subsets. Exemplarily, the execution device can input each subset into a machine learning model to obtain a set of descriptive information generated by the machine learning model. The machine learning model in this application can be a convolutional neural network, a fully connected neural network, a neural network based on an attention mechanism, a residual neural network, a support vector machine, or other types of machine learning models, etc.
[0102] Optionally, before inputting each of the M subsets included in each group of images into the machine learning model, the execution device may first align at least one image included in each of the M subsets; illustratively, since each group of images may have a corresponding image acquisition trajectory, the execution device may align the images in each of the M subsets included in each group of images according to the image acquisition trajectory corresponding to each group of images, to obtain M aligned subsets, and then input each of the M aligned subsets into the machine learning model. Alternatively, the execution device may not align each of the aforementioned subsets, but instead input each subset and the image acquisition trajectory corresponding to the subset into the machine learning model.
[0103] Among them, each group of description information in the M groups of description information includes a first description information of at least one element in the first environment, that is, each group of description information in the M groups of description information includes at least one first description information. Since the description information of different groups in the M groups of description information may include the first description information of the same element (hereinafter referred to as the "target element" for the convenience of description), that is, the M groups of description information can include at least two first description information of the same target element. The at least two first description information of the target element come from different groups in the M groups of description information, and at least one of the aforementioned target elements can exist in the first environment.
[0104] Exemplarily, M is 3, and the three groups of description information include the first group of description information, the second group of description information, and the third group of description information. The first group of description information includes the first description information 1 of element 1, the first description information 1 of element 2, and the first description information 1 of element 3. The second group of description information includes the first description information 2 of element 3, the first description information 2 of element 4, and the first description information 2 of element 5. The third group of description information includes the first description information 3 of element 3, the first description information 3 of element 4, the first description information 3 of element 5, and the first description information 3 of element 6. Then, element 3, element 4, and element 5 are all target elements. The three groups of description information include 3 first description information of element 3, 2 first description information of element 4, and 2 first description information of element 5. It should be understood that the example here is only for the convenience of understanding the meaning of "M groups of description information" and is not used to limit this solution.
[0105] In another implementation, step 302 may include: the execution device inputs each group of images as a whole into the machine learning model to obtain a set of description information generated by the machine learning model, and the set of description information includes a first description information for each element in the first environment.
[0106] 303. Determine a map corresponding to the element in the first environment according to the first description information.
[0107] In an embodiment of the present application, the execution device can generate a map corresponding to the elements in the first environment based on the first description information obtained in step 302. For example, if the first environment is a road environment, the "map corresponding to the elements in the first environment" can be expressed as a map containing road elements. If the first environment is a shopping mall, the "map corresponding to the elements in the first environment" can be expressed as a navigation map of the shopping mall. If the first environment is an environment where power lines are located, the "map corresponding to the elements in the first environment" can be expressed as a map containing power lines. If the first environment is a construction site, the "map corresponding to the elements in the first environment" can be expressed as a map containing buildings in the construction site. If the first environment is an environment where a forest is located, the "map corresponding to the elements in the first environment" can be expressed as a map containing trees in the forest, etc. The specific types of maps generated can be determined in combination with the actual application scenario. In an embodiment of the present application, a specific application scenario of the present solution is provided, which improves the degree of integration between the present solution and the specific application scenario.
[0108] Regarding the specific implementation of step 303, exemplarily, in one case, the at least one group of images of the first environment obtained in step 301 includes only one group of images. Exemplarily, in one implementation, if the execution device obtains M groups of description information based on each group of images in step 302, and the M groups of description information include at least two first description information for each target element in at least one target element, then the execution device may process each target element in a manner that: the execution device generates second description information of the target element based on the at least two first description information of the target element; wherein the second description information of the target element indicates the category and position of the target element, and in the case where the target element is a line element, the second description information of the target element also indicates the shape of the target element. Optionally, the second description information of the target element also indicates other attribute information of the target element. The meaning and specific expression of the "second description information" can refer to the above description of the "first description information", the difference being that "elements in the first environment" in the above description are replaced with "target elements".
[0109] Regarding the specific implementation process of "the execution device obtains at least two first description information of the same target element from M groups of description information", exemplarily, the execution device can assign an identification information to each element in the first environment based on the image acquisition trajectory corresponding to each group of images, that is, in each group of images, the same element in the first environment will be assigned the same identification information, and different elements will be assigned different identification information. The execution device can then determine the identification information corresponding to each first description information in the M groups of description information, thereby being able to group and cluster according to the identification information corresponding to each first description information, that is, clustering at least two first description information corresponding to the same identification information, thereby obtaining at least two first description information corresponding to the same element.
[0110] Exemplarily, based on a set of images, it is determined that a first environment includes elements 1, 2, 3, 4, and 5. Element 1 is assigned identification information 1, element 2 is assigned identification information 2, element 3 is assigned identification information 3, element 4 is assigned identification information 4, and element 5 is assigned identification information 5. Based on the aforementioned set of images, a total of three groups of description information are obtained, including the first group of description information, the second group of description information, and the third group of description information. The first group of description information includes the first description information 1 of element 1, the first description information 1 of element 2, and the first description information 1 of element 3. The first description information 1 of element 1 corresponds to identification information 1, the first description information 1 of element 2 corresponds to identification information 2, and the first description information 1 of element 3 corresponds to identification information 3. The second group of description information includes the first description information 2 of element 3, the first description information 2 of element 4, and the first description information 2 of element 5. The first description information 2 of element 3 corresponds to identification information 3, the first description information 2 of element 4 corresponds to identification information 4, and the first description information 2 of element 5 corresponds to identification information 5. The third group of description information includes the first description information 3 of element 3, the first description information 3 of element 4, the first description information 3 of element 5, and the first description information 3 of element 6. The first description information 3 of element 3 corresponds to identification information 3, the first description information 3 of element 4 corresponds to identification information 4, the first description information 3 of element 5 corresponds to identification information 5, and the first description information 3 of element 6 corresponds to identification information 6. The execution device can then obtain the three first description information of element 3 from the three groups of description information based on identification information 3, obtain the two first description information of element 4 from the three groups of description information based on identification information 4, and obtain the two first description information of element 5 from the three groups of description information based on identification information 5. It should be understood that the examples here are only used to facilitate understanding of this solution and are not intended to limit this solution.
[0111] Regarding the specific implementation method of "the execution device generates second description information of the target element based on the at least two first description information of the target element", exemplarily, if the target element is a line element in the first environment, then after obtaining at least two first description information corresponding to the same line element, that is, obtaining at least two lines based on the aforementioned at least two first description information, the execution device can sample each of the at least two lines to obtain at least two discrete point sets corresponding one-to-one to the at least two lines, and perform spatial clustering on the at least two discrete point sets. Since errors may occur when the execution device assigns identification information to elements in the first environment, after performing spatial clustering on the at least two discrete point sets, outlier discrete point sets can be eliminated, and then a new line can be fitted based on the remaining discrete points. The second description information can be determined based on the new line, and the second description information includes the position and shape of the new line, that is, the position and shape of the target element. In addition, the execution device can also retain the category of the target element (optionally, other attribute information) carried by the first description information of the target element in the second description information of the target element, so that the second description information of the target element also indicates the category of the target element (optionally, other attribute information).
[0112] If the target element is an element that uses the coordinates of a point in the first environment to indicate the position, such as a road sign, a road sign, a traffic light or a pole in a road environment, or a store in a shopping mall, or a tree in a forest, etc., no exhaustive list is given here. After the execution device obtains at least two first description information of the same target element, that is, based on the aforementioned at least two first description information, it determines at least two groups of points corresponding to the same target element, and each group of points includes at least one point. The execution device can directly eliminate outlier points from the at least two groups of points, and then determine the final coordinates of the point corresponding to the target element based on the remaining points (that is, the coordinates of the point included in the second description information of the target element); for example, a weighted average can be performed based on the coordinates of the remaining points to obtain the final coordinates of the point, or the aforementioned "weighted average" can also be replaced by finding the median or other calculation methods, etc., which are not exhaustive in the embodiments of the present application.
[0113] It should be noted that "the execution device generates the second description information of the target element according to the at least two first description information of the target element" is conducive to eliminating the inconsistency between frames. For a more intuitive understanding of this solution, please refer to Figure 5 , Figure 5 A schematic diagram of inter-frame inconsistency provided in an embodiment of the present application is provided. Figure 5 Including two sub-diagrams on the left and right, Figure 5 The left and right sub-diagrams of the diagram represent images collected at different locations on the same road. Figure 5The additional lines in the left and right sub-diagrams of the diagram represent lane lines recognized based on the image of the first environment. Figure 5 From the left and right schematic diagrams, we can see that for the same two lane lines in the same road environment, the squares of the lane lines recognized based on the two different images are different, and based on Figure 5 The distance between the two lane lines identified in the left sub-diagram is 4.0 meters (m). Figure 5 The distance between the two lane lines identified in the right sub-schematic diagram is 3.7 meters (m). Therefore, for the same road element in the same road environment, the first description information obtained based on different images may be different. This phenomenon is referred to as inter-frame inconsistency in this application. This inter-frame inconsistency is caused by factors such as shooting angle and shooting position. It should be understood that Figure 5 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0114] Since at least one target element can be determined based on M groups of identification information, the execution device can repeat the above steps at least once to generate second description information for each target element in at least one target element based on at least two first description information corresponding to each target element in the first environment. The execution device then generates a first map corresponding to the elements in the first environment based on the second description information of the target elements, which may include: the execution device determines the first map corresponding to the elements in the first environment based on the second description information of each target element in the at least one target element; illustratively, the "first map corresponding to the elements in the first environment" may include the second description information of all target elements in the at least one target element.
[0115] Optionally, the execution device can also determine at least one single element in the first environment based on the M group identification information, wherein the M group identification information includes a first description information for each single element, that is, the single element and the target element are different elements in the first environment, and the difference between "single element" and "target element" is that the M group identification information includes at least two first description information of the target element, while the M group identification information only includes one first description information of the single element, then the "first map corresponding to the elements in the first environment" can also include the first description information of each single element in at least one single element.
[0116] Optionally, if each piece of first description information is expressed in a vector format to indicate an element in the first environment, then the “first map corresponding to the element in the first environment” may be expressed in a vector map format.
[0117] Since each image of the first environment obtained is easily affected by the acquisition position and acquisition angle, there is inconsistency between frames, that is, different first description information of the same element in the first environment may be obtained based on different images. In the embodiment of the present application, each group of images of the first environment is divided into M subsets, and a group of description information is obtained using each subset of the M subsets. Then, at least two first description information of the same element (that is, the target element) in the first environment can be obtained from the M groups of description information, and the second description information of the target element is generated based on the at least two first description information of the target element. This is conducive to eliminating the impact of inconsistency between frames, and is conducive to making the position information carried by the second description information of the element in the first environment more accurate, and thus is conducive to obtaining a more accurate map.
[0118] In another implementation, if the execution device obtains a set of description information based on each group of images in step 302, and the set of description information includes a first description information for each element in the first environment, the execution device can determine the first map corresponding to the elements in the first environment based on the set of description information, and the "first map corresponding to the elements in the first environment" may include the set of description information.
[0119] In another case, the number of at least one group of the at least one group of images obtained in step 301 is N, where N is an integer greater than or equal to 2. Taking the first environment as a road environment as an example for illustration, step 303 may include: executing the device to generate a piece of map information corresponding to a road element in the road environment based on the first description information corresponding to each of the N groups of images, that is, N pieces of map information corresponding one-to-one to the N groups of images can be obtained. It should be noted that the specific implementation of "executing the device to generate a piece of map information corresponding to a road element in the road environment based on the first description information corresponding to each group of images" can be found in the above description of how to generate a first map corresponding to an element in the first environment. The difference is that "the first map corresponding to the elements in the first environment" in the above description is replaced with "map information corresponding to the road elements", and the specific implementation is not repeated here. In addition, the specific expressions of "map information corresponding to the elements in the first environment" and "the first map corresponding to the elements in the first environment" are similar and are not repeated here.
[0120] After obtaining N map information corresponding to the N groups of images, the execution device needs to fuse the N map information to obtain a second map corresponding to the road elements in the road environment, that is, to obtain a final map corresponding to the road elements.
[0121] Exemplarily, in one implementation, for any one piece of map information among N pieces of map information (hereinafter referred to as "first map information" for the convenience of description), the execution device can divide the map space indicated by the first map information into multiple spatial areas based on the first image acquisition trajectory, and obtain third description information of at least one instance in each of the aforementioned multiple spatial areas based on the first map information, that is, it can obtain third description information of multiple instances in multiple spatial areas.
[0122] The first map information corresponds to the first image acquisition track of the at least two image acquisition tracks, that is, the first image acquisition track is the image acquisition track of a group of images used to obtain the first map information. Optionally, there may be overlapping areas between different spatial areas in the multiple spatial areas. Alternatively, there may be no overlapping areas between different spatial areas in the multiple spatial areas. For a more intuitive understanding of this solution, please refer to Figure 6 , Figure 6 A schematic diagram of multiple spatial regions provided in an embodiment of the present application, Figure 6 The lines in represent lane lines (ie, an example of a road element) included in the first map information. Figure 6 Each rectangular box in represents a spatial region. Figure 6 The lines in represent lane lines and road lines, such as Figure 6 As shown, there are overlapping areas between different rectangular frames, that is, there may be overlapping areas between different spatial regions in multiple spatial regions. It should be noted that Figure 6 The spatial region is shown in a top view, but the actual spatial region may be a three-dimensional spatial region, and Figure 6 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0123] For example, the instances in this application can also be understood as "rigid body instances", that is, when fusing N map information, each rigid body instance is regarded as a minimum adjustment unit, that is, in this implementation, the fusion is performed at the granularity of "instance". Particularly, since after the first map information is divided into multiple spatial regions, the multiple road elements in the first map information are also divided into respective spatial regions, an instance can represent a road element in one of the multiple spatial regions (for example, representing a second road element); or, since a linear road element may run through multiple spatial regions, an instance can also represent a partial line segment of a linear road element (that is, the first road element) in a spatial region.
[0124] The third description information indicates the category and location of the instance. When the instance represents a partial line segment of a line-shaped road element within a spatial area, the third description information also indicates the shape of the instance. Optionally, the third description information also indicates other attribute information of the instance.
[0125] Exemplarily, if an instance represents a second road element, the third description information of the instance may include the second description information of the second road element, that is, the position and category of the second road element (optionally, other attribute information). For example, the position of the instance can be expressed using posture information, and the initial posture information of the instance can be expressed as (x, y, z, 0, 0, 0), where the aforementioned x, y, and z represent the coordinates of the instance in the x-axis direction, y-axis direction, and z-axis direction, respectively, and the aforementioned 0, 0, 0 represents that the initial rotation angles of the instance around the x-axis direction, y-axis direction, and z-axis direction are all 0 degrees.
[0126] If an instance represents a partial line segment of a first road element within a spatial region, the third description information of the instance may include the initial pose information of the center point of the aforementioned line segment, a function expression of the first road element to which the aforementioned line segment belongs, and the category of the first road element (optionally, also including other attribute information). For example, the initial pose information of the center point of the line segment can be expressed as (x, y, z, 0, 0, 0), where the aforementioned x, y, and z represent the coordinates of the instance in the x-axis direction, y-axis direction, and z-axis direction, respectively, and the aforementioned 0, 0, 0 represents that the initial rotation angles of the instance around the x-axis direction, y-axis direction, and z-axis direction are all 0 degrees.
[0127] For a more intuitive understanding of this solution, please refer to Figure 7 , Figure 7 A schematic diagram of the center points of multiple instances provided in the embodiments of this application, Figure 7 The four lines in represent the four lane lines (ie, an example of a road element) included in the first map information. Figure 7 Each point in represents the center point of a line segment in a lane line. It should be understood that Figure 7 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0128] Optionally, after obtaining the third description information of at least one instance in each of the plurality of spatial regions based on the first map information, the execution device may further obtain first relative pose information between adjacent instances in the plurality of instances obtained based on the same first map information. For example, the first relative pose information may be represented as (x, y, z, 0, 0, 0).
[0129] The execution device can perform the above operation on each of the N map information, thereby dividing each of the N map information into multiple spatial areas and obtaining third description information of at least one instance in each spatial area; optionally, first relative posture information between adjacent instances in the multiple instances obtained based on each map information is also obtained.
[0130] Since the N map information are obtained based on N sets of images of the same road environment, that is, the N map information reflect the same road environment, the N map information can be understood as N observation results of the same road environment. Therefore, among the multiple instances (also called rigid body instances) obtained based on the N map information, there are N instances corresponding to the same second road element (or the same line segment of the same first road element) in the road environment, that is, the same second road element (or the same line segment of the same first road element) in the road environment corresponds to N third description information, where N is an integer greater than or equal to 2.
[0131] In order to obtain at least two third description information corresponding to the same second road element (or the same line segment of the same first road element) in the road environment, since at least two instances corresponding to the same second road element (or the same line segment of the same first road element) in the road environment both represent the second road element (or the same line segment of the first road element), for the convenience of description, "obtaining at least two third description information corresponding to the same second road element (or the same line segment of the same first road element) in the road environment" can also be referred to as "obtaining at least two third description information corresponding to the first instance", and one first instance corresponds to one second road element (or a line segment in a first road element).
[0132] Exemplarily, the execution device may perform a pairing operation using a preset algorithm based on the third description information of each instance obtained from the N map information, so that at least two instances representing the same second road element (or the same line segment of the same first road element) can be assigned to the same group, thereby determining an instance combination corresponding to each second road element in the road environment (i.e., an example of the first environment), and determining an instance combination corresponding to each line segment of each first road element in the road environment. The instance combination in this application includes at least two instances. Since each instance has corresponding third description information, at least two third description information corresponding to each second road element in the road environment and at least two third description information corresponding to each line segment of each first road element in the road environment are determined. Exemplarily, each of the at least two third description information is derived from different map information in the N map information. The preset algorithm may be a spatial search algorithm or other type of algorithm, etc., which are not exhaustively listed in the embodiments of this application.
[0133] After obtaining at least two pieces of third description information corresponding to each first instance, the execution device may fuse the at least two pieces of third description information corresponding to each first instance to determine final position information corresponding to each first instance. The final position information corresponding to multiple instances in multiple spatial regions is used to obtain a second map corresponding to the elements in the first environment. More specifically, "fusing at least two pieces of third description information corresponding to each first instance" can be understood as fusing multiple pieces of third description information corresponding to each second road element in the road environment, as well as fusing multiple pieces of third description information corresponding to each line segment of each first road element.
[0134] Regarding the specific implementation process of "fusing at least two third description information corresponding to any first instance", exemplarily, after the execution device obtains at least two third description information corresponding to the first instance, since one third description information is used to indicate the position of a second road element, or to indicate the position and shape of a line segment of a first road element, it can determine the second relative posture information between the positions indicated by different third description information in the aforementioned at least two third description information based on the at least two third description information corresponding to the first instance. Exemplarily, the second relative posture information can specifically include: relative displacement on the x-axis, y-axis and z-axis, and relative rotation angles on the x-axis, y-axis and z-axis.
[0135] It should be noted that if the first instance represents a second road element, since the coordinates of a point are used to indicate the position of the second road element, and there are no relative rotation angles on the x-axis, y-axis, and z-axis between different points, when the first instance represents the second road element, the relative rotation angles on the x-axis, y-axis, and z-axis included in the second relative pose information can all be 0. For example, if the first instance represents a second road element, that is, each piece of third description information corresponding to the first instance uses the coordinates of a point to indicate the position of the second road element, then the point-to-point iterative closest point (ICP) method can be used to determine the second relative pose information between the positions indicated by different third description information corresponding to the first instance.
[0136] If the first instance represents a line segment of the first road element, the position of the line segment is indicated by the coordinates of the line segment's center point and a function expression of the first road element to which the line segment belongs. Therefore, different line segments may have relative rotation angles along the x-axis, y-axis, and z-axis. For example, if the first instance represents a line segment of the first road element, a point-to-line ICP method can be used to determine the second relative pose information between the positions indicated by different third description information corresponding to the first instance.
[0137] Alternatively, the above-mentioned "point-to-point ICP" and "point-to-line ICP" can also be replaced by traditional geometric ICP, etc. The specific method used in the process of "determining the second relative posture information between the positions indicated by different third description information in at least two third description information corresponding to the same second road element (or the same line segment in the same first road element)" can be determined in combination with the actual application scenario.
[0138] After obtaining the second relative posture information between the positions indicated by different third description information in at least two third description information corresponding to each first instance, and the first relative posture information between adjacent instances in multiple instances obtained based on each map information in N map information, the execution device can use an optimization algorithm to optimize the information used to indicate the position in each third description information in at least two third description information corresponding to each first instance, and can obtain at least two optimized third description information corresponding to each first instance. Exemplarily, the optimization algorithm can be a graph optimization algorithm, a filtering algorithm, or other types of optimization algorithms. Further, the graph optimization algorithm can specifically include a factor graph optimization algorithm, and the filtering algorithm can include a Kalman filtering algorithm, etc. The specific algorithm to be used can be determined in combination with the actual application scenario.
[0139] For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 A schematic diagram of a graph created when optimizing using a graph optimization algorithm provided in an embodiment of the present application, Figure 8 In this paper, we take the factor graph optimization algorithm as an example, that is, Figure 8 The factor graph established in is Figure 8 As shown, the factor graph is composed of two rows of nodes and edges. Each node represents an instance (also called a rigid body instance). The instances represented by the nodes in the first row are obtained based on map information 1, and the instances represented by the nodes in the second row are obtained based on map information 2. The aforementioned nodes can also be called vertices.
[0140] Among them, the edge connected to only one rigid body instance is called a priori edge, and the prior edge is used to provide the initial position of the node; to expand on this, when a certain instance represents a second road element, the prior edge is used to provide the coordinates of the point in the third description information of the instance, or, when a certain instance represents a line segment in the first road element, the prior edge is used to provide the coordinates of the point (that is, the center point of the line segment) in the third description information of the instance and the function expression (that is, the function expression of the line segment). The edges between different nodes in the same row are called odometry edges, and the odometry edges are used to provide the first relative pose information. The edges between nodes in different rows are called registration edges, and the registration edges are used to provide the second relative pose information. The meanings of "first relative pose information" and "second relative pose information" can be found in the above description and will not be repeated here. After establishing Figure 8 After obtaining the factor graph shown in , the factor graph can be used to optimize the initial position provided by each prior edge to obtain the optimized position corresponding to each instance. This also optimizes each piece of third description information. The optimized third description information includes the aforementioned optimized position. For example, a least squares problem can be established based on the factor graph, and the process of "solving the optimized position" can be understood as the process of solving the least squares problem.
[0141] It should be noted that Figure 8 The example shown in the figure is based on two map information. This is only for the convenience of understanding this solution. In actual application environments, a factor graph can also contain nodes derived from more map information and edges between each node. The specific details can be determined based on the actual application scenario.
[0142] Since each of the at least two optimized third description information mentioned above indicates a location of the first instance, the execution device can also perform a deduplication operation based on the at least two optimized third description information corresponding to the first instance, thereby obtaining the final location information corresponding to the first instance; the optimized third description information also includes the category of the road element, and optionally, the optimized third description information also includes other attribute information of the road element.
[0143] Optionally, different road elements can use different deduplication methods. For example, if the first instance represents a line segment in the first road element (such as a lane boundary line, a lane centerline, and a road line) in the road element, the execution device can determine at least two line segments stacked together based on the at least two optimized third description information corresponding to the first instance. It can arbitrarily select a line segment from the at least two line segments and directly delete the remaining line segments, retaining only the selected line segment, thereby obtaining the final position information of the line segment; or, it can arbitrarily select a line segment from the at least two line segments, perform expansion processing on the selected line segment, remove other line segments in the expansion area, and then retain only the selected line segment, thereby obtaining the final position information of the line segment. After the execution device performs the above operations on multiple line segments in the road environment (that is, an example of the first environment), it can obtain multiple discrete line segments, that is, obtain the final position information of multiple line segments.
[0144] For another example, if the first instance represents a ground sign in a road element, since the coordinates of a point are used to express the position of the ground sign, and the distribution characteristics of the ground sign are that it is naturally divided by lane lines in the horizontal direction, the position indicated by the optimized third description information of the ground sign will have a relatively small error in the horizontal direction; if the first instance represents a ground sign in a road element, the execution device can determine the coordinates of at least two points based on at least two optimized third description information corresponding to the first instance, and the coordinates of each of the at least two points are used to indicate the position of the center point of the ground sign; the execution device can use a neighborhood search algorithm, a non-maximum suppression algorithm or other algorithms to perform a deduplication operation on the coordinates of the at least two points, thereby obtaining the coordinates of a point corresponding to the ground sign, that is, obtaining the final position information of the ground sign. Optionally, when the neighborhood search algorithm is used, the constraints of the line-shaped road elements on the ground need to be met, that is, when the line-shaped road elements on the ground are found, the horizontal search will be stopped.
[0145] For example, if the first instances represent first road elements other than ground markings in the road elements, such as signal lights, poles, road signs, or other types of first road elements (hereinafter referred to as "ground road elements" for convenience of description), and the positions of the aforementioned ground road elements are expressed by the coordinates of the points, the coordinates of the points corresponding to each first instance can be weighted and averaged to obtain the final position information corresponding to the first instance, for example. The execution device repeats the aforementioned operation at least once to obtain the final position information of each ground road element. Alternatively, after the execution device obtains at least two third description information of each ground road element in all ground road elements, the execution device can determine the positions of all points representing the ground road elements. The execution device can use a clustering algorithm to re-cluster all points, and weighted average the positions of the points in each group obtained after clustering to obtain the final position information of a ground road element. The execution device performs the weighted average operation on each group after clustering to obtain the final position information of each ground road element. Optionally, in view of the characteristics that the ground road elements are more densely distributed at intersections and more sparsely distributed at non-intersections, a smaller clustering search radius can be used at intersections and a larger clustering search radius can be used at non-intersections.
[0146] For a more intuitive understanding of the present scheme, please refer to Figure 9 , Figure 9 A comparison diagram before and after deduplication provided by the embodiments of the present application, Figure 9 including left and right two sub-diagrams, Figure 9 The left sub-diagram of is a diagram before deduplication, Figure 9 The right sub-diagram of represents a diagram after deduplication, Figure 9 The right sub-diagram of is a specific road environment as a background, wherein, Figure 9 There are multiple groups of points stacked together and lines stacked together in the left sub-diagram of, and refer to Figure 9 The right sub-diagram of, after the deduplication operation is performed, the redundant points and lines are deleted. It should be understood that, Figure 9 The instances in are only for convenient understanding of the present scheme, and are not used to limit the present scheme.
[0147] Regarding the process of determining the second map based on the optimized third descriptive information, illustratively, after obtaining the final positions of multiple line segments, the execution device may connect the multiple discrete line segments into a continuous line; optionally, the execution device may also smooth the aforementioned continuous line; thereby, the execution device can determine the position and shape of a second road element in the road environment. The second map corresponding to the road element in the road environment may include the position and shape of the second road element in the road environment, the category of the second road element, and optionally, other attribute information of the second road element. Optionally, the execution device may also assign directional information to the continuous line; and the second map corresponding to the road element in the road environment may also include the directional information of the second road element.
[0148] The second map corresponding to the road elements in the road environment may also include the final position information and category of the first road elements, and optionally, other attribute information of the first road elements. For example, the second map corresponding to the road elements in the road environment may also include the final position information and category of the ground signs, and optionally, other attribute information of the ground signs. For another example, the second map corresponding to the road elements in the road environment may also include the final position information and category of the ground road elements, and optionally, other attribute information of the ground road elements. It should be noted that the specific information included in the "other attribute information" has been explained in the above description and will not be repeated here.
[0149] When capturing each set of images of the road environment, there will be an image capture track. However, the map obtained based on the images captured in a single time may still be affected by the capture perspective. In the embodiment of the present application, the number of at least one set of images captured is N, and the value of N is an integer greater than or equal to 2. That is, different image capture tracks are used to capture images of the first environment, and then the final map is obtained based on at least two sets of images. This is conducive to eliminating errors in the single observation process and making the final map more accurately reflect the real environment, that is, obtaining a more accurate map. For a more intuitive understanding of this solution, please refer to Figure 10 , Figure 10 A schematic diagram of two stacked map information provided in an embodiment of the present application, Figure 10 The map information obtained based on each of the two sets of images is visualized in Figure 10 The lane lines and road lines are shown in Figure 10 As shown, if there are differences between the lane lines and road lines in the two sets of map information, then obtaining the final map or mapping based on at least two sets of images is conducive to eliminating the error caused by the influence of the acquisition perspective on a single observation result (i.e., a map information obtained based on a set of images). It should be understood that Figure 10The examples in the specification are only for the convenience of understanding the scheme and do not limit the scheme.
[0150] In addition, for the line-shaped road element (for example, lane line, road line, etc.), since the line-shaped road element is generally long, it is difficult to directly adjust the final position information of the line-shaped road element to perfectly match the actual position information of the line-shaped road element at the granularity of the entire line. In the embodiment of the application, the entire line is divided into a plurality of line segments, and then fusion is performed at the granularity of the line segment, which is conducive to making the line-shaped road element more consistent, thereby being conducive to obtaining more accurate positions, that is, being conducive to obtaining a more accurate map.
[0151] In the embodiment of the application, after at least one group of images of the first environment are obtained, first description information of elements in the first environment is obtained according to each group of images of the first environment. The first description information indicates the category and position of the elements in the first environment. For the line element in the first environment, the first description information also indicates the shape of the line element. That is, the first description information includes key information related to composition obtained from the images. Then, a map corresponding to the elements in the first environment is generated according to the first description information obtained from each group of images of the first environment. Since the data amount of the first description information obtained from the at least one group of images of the first environment is much smaller than the data amount of the at least one group of images, the method of generating the map based on the obtained first description information can greatly reduce the computer resources consumed in the mapping process. In addition, after the computer resources consumed in the mapping process are reduced, the mapping method provided by the application can be deployed on the terminal side, so that the terminal side can more timely realize mapping of the elements in the environment, and a more accurate map can be obtained in time.
[0152] In Figures 1 to 10 Based on the corresponding embodiment, in order to better implement the above-mentioned scheme of the embodiment of the application, the related equipment for implementing the above-mentioned scheme is also provided. For details, please refer to Figure 11 , Figure 11 A structural diagram of the mapping device provided by the embodiment of the application is shown in FIG. 11. The mapping device 1100 includes: an acquisition module 1101 configured to acquire at least one group of images of a first environment. Each group of images in the at least one group of images includes at least one image of the first environment. The acquisition module 1101 is also configured to acquire first description information of elements in the first environment according to the at least one image of the first environment included in each group of images. The first description information indicates the category and position of the elements in the first environment. The first environment includes a line element, and the first description information of the line element further indicates the shape of the line element. A determination module 1102 is configured to determine a map corresponding to the elements in the first environment according to the first description information.
[0153] Optionally, the first environment is a road environment, the elements in the first environment include road elements, and the map corresponding to the elements in the first environment is a map containing road elements.
[0154] Optionally, the road element includes a first road element and a second road element, the first road element is a line-shaped road element, and the second road element is an element in the road environment and is different from the first road element, wherein the first description information of the first road element includes a function expression and a category of the first road element, the function expression indicates the position and shape of the first road element, and the first description information of the second road element includes the coordinates of a point and the category of the second road element, and the coordinates of the point indicate the position of the second road element.
[0155] Optionally, the acquisition module 1101 is specifically configured to: obtain M subsets based on at least one image of the first environment included in each group of images, each subset in the M subsets including at least one image of the first environment, where M is an integer greater than or equal to 2; obtain M groups of description information based on the M subsets, wherein each group of description information in the M groups of description information includes at least one first description information, the M groups of description information include at least two first description information of a target element, the element in the first environment includes the target element, and the at least two first description information of the target element come from different groups in the M groups of description information;
[0156] Determination module 1102 is specifically used to: generate second description information of the target element based on at least two first description information of the target element, the second description information indicating the category and position of the target element, and when the target element is a line element, the second description information also indicates the shape of the target element; based on the second description information of the target element, determine the map corresponding to the element in the first environment.
[0157] Optionally, the number of at least one group is N groups, where N is an integer greater than or equal to 2, and the N groups of images of the first environment are obtained by performing at least two image capture operations on the first environment, and there are different trajectories in at least two image capture trajectories corresponding to the at least two image capture operations.
[0158] Optionally, the first environment is a road environment, and the elements in the first environment include road elements. The determination module 1102 is specifically used to: determine a map information corresponding to the road element based on the first description information corresponding to each group of images in the N groups of images, wherein the N groups of images correspond one-to-one to the N map information, the first map information is any one of the N map information, and the first map information corresponds to the first image acquisition track in at least two image acquisition tracks; based on the first image acquisition track, divide the map space indicated by the first map information into multiple spatial areas; according to the first map information, obtain third description information of multiple instances in the multiple spatial areas, where one instance represents a road element in a spatial area, or one The instance represents a partial line segment of a line-shaped road element within a spatial area, wherein the third description information indicates the category and position of the instance. When the instance represents a partial line segment of a line-shaped road element within a spatial area, the third description information also indicates the shape of the instance; based on at least two third description information corresponding to the first instance, the final position information corresponding to the first instance is determined, the multiple instances within the multiple spatial areas include the first instance, and each of the at least two third description information corresponding to the first instance is derived from one map information of at least two map information, wherein the final position information corresponding to the multiple instances within the multiple spatial areas is used to obtain a map corresponding to the elements in the first environment.
[0159] It should be noted that the information interaction, execution process, etc. between the modules / units in the mapping device 1100 are the same as those in the present application. Figures 1 to 10 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0160] Next, we introduce a device provided by an embodiment of the present application. In the case where the device is specifically an execution device, please refer to Figure 12 , Figure 12 A structural diagram of a device provided in an embodiment of the present application, specifically, the device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203 and a memory 1204 (wherein the number of processors 1203 in the device 1200 can be one or more, Figure 12 (taking one processor as an example), the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 may be connected via a bus or other means.
[0161] The memory 1204 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1203. A portion of the memory 1204 may also include non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0162] Processor 1203 controls the operation of the device. In specific applications, the various components of the device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0163] The methods disclosed in the above embodiments of the present application can be applied to or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1203. The above processor 1203 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1203 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1204, and processor 1203 reads the information in memory 1204 and, in conjunction with its hardware, completes the steps of the above method.
[0164] Receiver 1201 can be used to receive input digital or character information and generate signal input related to device settings and function control. Transmitter 1202 can be used to output digital or character information through the first interface. Transmitter 1202 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1202 can also include a display device such as a display screen.
[0165] In the embodiment of the present application, the processor 1203 is used to execute Figures 1 to 10 The method executed by the execution device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 12031 in the processor 1203 executes the above steps is the same as that in the present application. Figures 1 to 10 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 10 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0166] The present application also provides a vehicle. Figure 13 , Figure 13 This is a schematic diagram of the structure of a vehicle provided in an embodiment of the present application. Vehicle 100 is configured for a fully or partially autonomous driving mode. For example, vehicle 100 can control itself while in autonomous driving mode and, through human operation, determine the current state of the vehicle and its surrounding environment, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the likelihood that the other vehicle will perform the possible behavior, and control vehicle 100 based on the determined information. While vehicle 100 is in autonomous driving mode, vehicle 100 can also be set to operate without human interaction.
[0167] The vehicle 100 may include various subsystems, such as a travel system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power source 110, a computer system 112, and a user interface 116. Alternatively, the vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and component of the vehicle 100 may be interconnected via wired or wireless connections.
[0168] Travel system 102 may include components that provide powered movement for vehicle 100. In one embodiment, travel system 102 may include engine 118, power source 119, transmission 120, and wheels / tires 121.
[0169] The engine 118 may be an internal combustion engine, an electric motor, an air compression engine, or a combination of other types of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air compression engine. The engine 118 converts the energy source 119 into mechanical energy. Examples of the energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. The energy source 119 may also provide energy for other systems of the vehicle 100. The transmission 120 may transmit the mechanical power from the engine 118 to the wheels 121. The transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 120 may also include other devices, such as a clutch. The drive shaft may include one or more shafts that can be coupled to one or more wheels 121.
[0170] Sensor system 104 may include several sensors that sense information about the environment surrounding vehicle 100. For example, sensor system 104 may include a positioning system 122 (the positioning system may be a global positioning system (GPS), a BeiDou system, or other positioning systems), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. Sensor system 104 may also include sensors for internal systems of monitored vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors may be used to detect objects and their corresponding characteristics (position, shape, direction, speed, etc.). This detection and recognition is a key function for the safe operation of autonomous vehicle 100.
[0171] Among them, the positioning system 122 can be used to estimate the geographic location of the vehicle 100. The IMU 124 is used to sense the position and orientation changes of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. The radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, and can specifically be a millimeter wave radar or a laser radar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of objects. The laser rangefinder 128 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. The camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.
[0172] Control system 106 controls the operation of vehicle 100 and its components. Control system 106 may include various components, including a steering system 132 , a throttle 134 , a brake unit 136 , a computer vision system 140 , a lane control system 142 , and an obstacle avoidance system 144 .
[0173] The steering system 132 is operable to adjust the direction of travel of the vehicle 100. For example, in one embodiment, it may be a steering wheel system. The throttle 134 is used to control the operating speed of the engine 118 and, in turn, the speed of the vehicle 100. The brake unit 136 is used to control the deceleration of the vehicle 100. The brake unit 136 may use friction to slow the wheels 121. In other embodiments, the brake unit 136 may convert the kinetic energy of the wheels 121 into electrical current. The brake unit 136 may also take other forms to slow the rotation speed of the wheels 121 to control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. These objects and / or features may include traffic signs, road boundaries, and obstacles. The computer vision system 140 may use object recognition algorithms, structure from motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 can be used to map the environment, track objects, estimate their speed, and so on. The route control system 142 is used to determine the route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to determine the route and speed for the vehicle 100 by combining data from the obstacle avoidance system 144, GPS 122, and one or more predetermined maps. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise navigate obstacles in the environment of the vehicle 100. The aforementioned obstacles can specifically be represented by actual obstacles and virtual moving objects that may collide with the vehicle 100. In one embodiment, the control system 106 may include additional or alternative components other than those shown and described. Alternatively, some of the components shown above may be reduced.
[0174] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users via peripheral devices 108. Peripheral devices 108 may include a wireless communication system 146, an onboard computer 148, a microphone 150, and / or a speaker 152. In some embodiments, peripheral devices 108 provide a means for the user of vehicle 100 to interact with user interface 116. For example, onboard computer 148 may provide information to the user of vehicle 100. User interface 116 may also operate onboard computer 148 to receive user input. Onboard computer 148 may be operated via a touchscreen. In other cases, peripheral devices 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio (e.g., voice commands or other audio input) from the user of vehicle 100. Similarly, speaker 152 may output audio to the user of vehicle 100. Wireless communication system 146 may wirelessly communicate with one or more devices directly or via a communication network. For example, the wireless communication system 146 may utilize 3G cellular communications, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communications, such as LTE. Or 5G cellular communications. The wireless communication system 146 may utilize wireless local area network (WLAN) communications. In some embodiments, the wireless communication system 146 may utilize infrared links, Bluetooth, or ZigBee to communicate directly with devices. Other wireless protocols, such as various vehicle communication systems, may include one or more dedicated short range communications (DSRC) devices, which may include public and / or private data communications between vehicles and / or roadside stations.
[0175] Power source 110 can provide power to various components of vehicle 100. In one embodiment, power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more battery packs of such batteries can be configured as a power source to provide power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 can be implemented together, such as in some all-electric vehicles.
[0176] Some or all of the functions of the vehicle 100 are controlled by a computer system 112. The computer system 112 may include at least one processor 113 that executes instructions 115 stored in a non-transitory computer-readable medium such as a memory 114. The computer system 112 may also be a plurality of computing devices that control individual components or subsystems of the vehicle 100 in a distributed manner. The processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor 113 may be a dedicated device such as an application specific integrated circuit (ASIC) or other hardware-based processor. Although Figure 13 The processor, memory, and other components of the computer system 112 are functionally illustrated as being in the same block, but one of ordinary skill in the art will appreciate that the processor, or memory, may actually include multiple processors, or memories, that are not stored in the same physical housing. For example, the memory 114 may be a hard drive or other storage medium that is located in a housing different from the computer system 112. Thus, references to the processor 113 or memory 114 will be understood to include references to a collection of processors or memories that may or may not operate in parallel. Rather than using a single processor to perform the steps described herein, some components, such as the steering assembly and the deceleration assembly, may each have their own processor that performs only calculations related to the functionality of the component.
[0177] In various aspects described herein, the processor 113 may be located remotely from the vehicle 100 and in wireless communication with the vehicle 100. In other aspects, some of the processes described herein are performed on the processor 113 disposed within the vehicle 100 while others are performed by the remote processor 113, including taking the necessary steps to perform a single maneuver.
[0178] In some embodiments, memory 114 may contain instructions 115 (e.g., program logic) that are executable by processor 113 to perform various functions of vehicle 100, including those described above. Memory 114 may also contain additional instructions, including instructions for sending data to, receiving data from, interacting with, and / or controlling one or more of travel system 102, sensor system 104, control system 106, and peripherals 108. In addition to instructions 115, memory 114 may also store data such as road maps, route information, the vehicle's location, direction, speed, and other such vehicle data, as well as other information. This information may be used by vehicle 100 and computer system 112 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is provided for providing information to or receiving information from a user of vehicle 100. Optionally, user interface 116 may include one or more input / output devices within the set of peripherals 108, such as wireless communication system 146, onboard computer 148, microphone 150, and speaker 152.
[0179] Computer system 112 may control functions of vehicle 100 based on input received from various subsystems (e.g., travel system 102, sensor system 104, and control system 106) and from user interface 116. For example, computer system 112 may utilize input from control system 106 to control steering system 132 to avoid obstacles detected by sensor system 104 and obstacle avoidance system 144. In some embodiments, computer system 112 may be operable to provide control over many aspects of vehicle 100 and its subsystems.
[0180] Alternatively, one or more of the above components may be installed or associated separately from the vehicle 100. For example, the memory 114 may be partially or completely separate from the vehicle 100. The above components may be communicatively coupled together in a wired and / or wireless manner.
[0181] Optionally, the above components are just an example. In actual applications, the components in the above modules may be added or deleted according to actual needs. Figure 13 This should not be construed as limiting the embodiments of the present application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine an adjustment to its current speed. The objects can be other vehicles, traffic control devices, or other types of objects. In some examples, each identified object can be considered independently, and based on the object's respective characteristics, such as its current speed, acceleration, distance from the vehicle, etc., the speed of the vehicle to be adjusted can be determined.
[0182] Optionally, the vehicle 100 or a computing device associated with the vehicle 100 may be Figure 13 The computer system 112, computer vision system 140, and memory 114 can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, the behavior of each identified object depends on the behavior of each other, so the behavior of all identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, the vehicle 100 can determine what stable state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. Other factors can also be considered in determining the speed of the vehicle 100, such as the lateral position of the vehicle 100 on the road it is traveling on, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the vehicle's speed, the computing device can also provide instructions to modify the steering angle of the vehicle 100 to ensure that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).
[0183] In the embodiment of the present application, the processor 113 in the vehicle 100 is used to execute Figures 1 to 10 The method executed by the vehicle in the corresponding embodiment. It should be noted that the specific manner in which the processor 113 executes the above steps is the same as that in the present application. Figures 1 to 10 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 10 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0184] The present application also provides a computer-readable storage medium in which a program is stored. When the program is run on a computer, the computer executes the above-mentioned Figures 1 to 10 The illustrated embodiments describe the steps performed by the device in the method.
[0185] The present application also provides a computer program product including a program that, when executed on a computer, enables the computer to execute the aforementioned Figures 1 to 10 The illustrated embodiments describe the steps performed by the device in the method.
[0186] The present application also provides a circuit system in an embodiment, wherein the circuit system includes a processing circuit, wherein the processing circuit is configured to perform the aforementioned Figures 1 to 10 The illustrated embodiment describes the method.
[0187] The execution device or drawing device provided in the embodiment of the present application can be specifically a chip, which includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip to execute the above Figures 1 to 10 The method described in the embodiment shown. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), etc.
[0188] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.
[0189] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CLUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0191] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0192] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).
Claims
1. A mapping method, characterized in that: The method comprises: Acquire at least one set of images of a first environment, each set of images in the at least one set of images including at least one image of the first environment; acquiring, based on at least one image of the first environment included in each group of images, first description information of an element in the first environment, wherein the first description information indicates a category and a position of the element in the first environment, the first environment includes a line element, and the first description information of the line element indicates a shape of the line element; A map corresponding to the element in the first environment is determined according to the first description information.
2. The method according to claim 1, characterized in that The first environment is a road environment, the elements in the first environment include road elements, and the map corresponding to the elements in the first environment is a map containing the road elements.
3. The method according to claim 2, characterized in that The road elements include a first road element and a second road element, the first road element is a line-shaped road element, and the second road element is an element in the road environment and is different from the first road element, wherein the first description information of the first road element includes a function expression and a category of the first road element, the function expression indicates the position and shape of the first road element, and the first description information of the second road element includes the coordinates of a point and the category of the second road element, and the coordinates of the point indicate the position of the second road element.
4. The method according to any one of claims 1 to 3, characterized in that The acquiring, according to the at least one image of the first environment included in each group of images, first description information of an element in the first environment includes: Obtaining M subsets based on at least one image of the first environment included in each group of images, each subset of the M subsets including at least one image of the first environment, where M is an integer greater than or equal to 2; Acquire M groups of description information according to the M subsets, wherein each group of description information in the M groups of description information includes at least one first description information, the M groups of description information include at least two first description information of a target element, the element in the first environment includes the target element, and the at least two first description information of the target element come from different groups in the M groups of description information; The determining, according to the first description information, a map corresponding to the element in the first environment includes: generating second description information of the target element according to the at least two first description information of the target element, where the second description information indicates the category and position of the target element, and when the target element is a line element, the second description information further indicates the shape of the target element; Based on the second description information of the target element, a map corresponding to the element in the first environment is determined.
5. The method according to any one of claims 1 to 3, characterized in that The number of the at least one group is N groups, where N is an integer greater than or equal to 2. The N groups of images of the first environment are obtained by performing at least two image acquisitions on the first environment, and there are different trajectories in at least two image acquisition trajectories corresponding to the at least two image acquisitions.
6. The method according to claim 5, characterized in that The first environment is a road environment, and the elements in the first environment include road elements. Determining, based on the first description information, a map corresponding to the elements in the first environment includes: Determining, based on first description information corresponding to each of the N groups of images, a piece of map information corresponding to the road element, wherein the N groups of images correspond one-to-one to N pieces of map information, the first map information being any one of the N pieces of map information, and the first map information corresponding to a first image acquisition track of the at least two image acquisition tracks; dividing the map space indicated by the first map information into a plurality of spatial regions based on the first image acquisition trajectory; Acquire, based on the first map information, third description information of a plurality of instances within the plurality of spatial regions, wherein one instance represents a road element within the spatial region, or one instance represents a partial line segment of a line-shaped road element within the spatial region, wherein the third description information indicates a category and a location of the instance. If the instance represents a partial line segment of a line-shaped road element within the spatial region, the third description information also indicates a shape of the instance; Based on at least two third description information corresponding to a first instance, final position information corresponding to the first instance is determined, the multiple instances within the multiple spatial areas include the first instance, each third description information of the at least two third description information corresponding to the first instance is derived from one of the at least two map information, wherein the final position information corresponding to the multiple instances within the multiple spatial areas is used to obtain the map corresponding to the elements in the first environment.
7. A mapping device, characterized in that: The device comprises: an acquisition module, configured to acquire at least one set of images of a first environment, wherein each set of images in the at least one set of images includes at least one image of the first environment; The acquisition module is further configured to acquire first description information of an element in the first environment based on at least one image of the first environment included in each group of images, wherein the first description information indicates a category and a position of the element in the first environment, and the first environment includes a line element, and the first description information of the line element indicates a shape of the line element; A determination module is used to determine a map corresponding to the element in the first environment according to the first description information.
8. The device according to claim 7, characterized in that The first environment is a road environment, the elements in the first environment include road elements, and the map corresponding to the elements in the first environment is a map containing the road elements.
9. The device according to claim 8, characterized in that The road elements include a first road element and a second road element, the first road element is a line-shaped road element, and the second road element is an element in the road environment and is different from the first road element, wherein the first description information of the first road element includes a function expression and a category of the first road element, the function expression indicates the position and shape of the first road element, and the first description information of the second road element includes the coordinates of a point and the category of the second road element, and the coordinates of the point indicate the position of the second road element.
10. The device according to any one of claims 7 to 9, characterized in that The acquisition module is specifically used to: Obtaining M subsets based on at least one image of the first environment included in each group of images, each subset of the M subsets including at least one image of the first environment, where M is an integer greater than or equal to 2; Acquire M groups of description information according to the M subsets, wherein each group of description information in the M groups of description information includes at least one first description information, the M groups of description information include at least two first description information of a target element, the element in the first environment includes the target element, and the at least two first description information of the target element come from different groups in the M groups of description information; The determining module is specifically configured to: generating second description information of the target element according to the at least two first description information of the target element, where the second description information indicates the category and position of the target element, and when the target element is a line element, the second description information further indicates the shape of the target element; Based on the second description information of the target element, a map corresponding to the element in the first environment is determined.
11. The device according to any one of claims 7 to 9, characterized in that The number of the at least one group is N groups, where N is an integer greater than or equal to 2. The N groups of images of the first environment are obtained by performing at least two image acquisitions on the first environment, and there are different trajectories in at least two image acquisition trajectories corresponding to the at least two image acquisitions.
12. The device according to claim 11, characterized in that The first environment is a road environment, and the elements in the first environment include road elements. The determining module is specifically configured to: Determining, based on first description information corresponding to each of the N groups of images, a piece of map information corresponding to the road element, wherein the N groups of images correspond one-to-one to N pieces of map information, the first map information being any one of the N pieces of map information, and the first map information corresponding to a first image acquisition track of the at least two image acquisition tracks; dividing the map space indicated by the first map information into a plurality of spatial regions based on the first image acquisition trajectory; Acquire, based on the first map information, third description information of a plurality of instances within the plurality of spatial regions, wherein one instance represents a road element within the spatial region, or one instance represents a partial line segment of a line-shaped road element within the spatial region, wherein the third description information indicates a category and a location of the instance. If the instance represents a partial line segment of a line-shaped road element within the spatial region, the third description information also indicates a shape of the instance; Based on at least two third description information corresponding to a first instance, final position information corresponding to the first instance is determined, the multiple instances within the multiple spatial areas include the first instance, each third description information of the at least two third description information corresponding to the first instance is derived from one of the at least two map information, wherein the final position information corresponding to the multiple instances within the multiple spatial areas is used to obtain the map corresponding to the elements in the first environment.
13. A device, characterized in that The method comprises a processor coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
14. A vehicle, characterized in that: The method comprises a processor coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and when the program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
16. A computer program product, characterized in that The computer program product comprises a program, which, when run on a computer, causes the computer to perform the method according to any one of claims 1 to 6.