Map updating method and apparatus, model training method and apparatus, electronic device, and medium

By encoding and feature extraction of historical map data and collected image data, determining the location and category information of the target instance and updating the map data, the problem of coverage scale and update time of autonomous driving maps is solved, and efficient and automated map updates are achieved.

WO2025118538A1PCT designated stage expired Publication Date: 2025-06-12BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
PCT/CN2024/099546
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-06-17
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The coverage scale and update timeliness of autonomous driving maps have become key issues that restrict the development of the autonomous driving field.

Method used

A map update method and a training method for map update model are provided. By determining the associated historical map data and time-based acquisition image data for the same location information, encoding and feature extraction are performed, location information and category information of the target instance are determined, and map data is then updated.

Benefits of technology

It realizes high-fidelity map reconstruction and vectorized modeling and generation of map elements, improves the efficiency and automation capabilities of map updates, and solves the problems of coverage scale and update timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099546_12062025_PF_FP_ABST
    Figure CN2024099546_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a map updating method and apparatus, a model training method and apparatus, an electronic device, and a medium, relating to the technical field of artificial intelligence, and in particular to the fields such as automatic driving, intelligent traffic, computer vision, and image processing. A specific implementation solution is: for a same piece of position information, determining associated historical map data and M pieces of collected image data based on a time sequence, wherein M is an integer greater than or equal to 1; respectively encoding the historical map data and the M pieces of collected image data to obtain a historical map feature and M collected image features; determining position information and category information of a target instance on the basis of the historical map feature and the M collected image features, wherein the target instance represents lane information; and updating the map data on the basis of the position information and the category information of the target instance.
Need to check novelty before this filing date? Find Prior Art

Description

Map updating method, model training method, device, electronic device and medium

[0001] This application claims priority to Chinese patent application No. 202311662055.4, filed on December 6, 2023, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of autonomous driving, intelligent transportation, computer vision, image processing, etc. More specifically, the present disclosure provides a map update method, a map update model training method, a map update device, a map update model training device, an electronic device, a storage medium, and a computer program product. Background Art

[0003] Autonomous driving maps can provide autonomous vehicles with rich road element data and lane-level path planning, thus safeguarding autonomous driving. However, the coverage scale and update timeliness of autonomous driving maps have become key issues hindering the development of this field.

[0004] Summary of the Invention

[0005] The present disclosure provides a map updating method, a training method for a map updating model, a map updating device, a training device for a map updating model, an electronic device, a storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, a map updating method is provided, comprising: determining, for identical location information, associated historical map data and M time-series collected image data, where M is an integer greater than or equal to 1; encoding the historical map data and the M collected image data, respectively, to obtain historical map features and M collected image features; determining, based on the historical map features and the M collected image features, location information and category information of a target instance, where the target instance represents lane information; and updating the map data based on the location information and category information of the target instance.

[0007] According to another aspect of the present disclosure, a method for training a map update model is provided, comprising: determining training samples, the training samples comprising: a map prior, N time-series-based sample images, and a label; the map prior and the N sample images are all associated with the same location information, the label represents whether a real instance exists in the N sample images and represents the real location information and real category of the real instance, the real instance represents lane information, and N is an integer greater than or equal to 1; encoding the map prior and the N sample images respectively to obtain map features and N sample image features; determining output location information and output category of at least one output instance based on the map features and the N sample image features; and training the map update model based on the real location information, the output location information, the real category, and the output category.

[0008] According to another aspect of the present disclosure, a map updating device is provided, comprising: a first data determination module, a first encoding module, a target instance determination module, and an update module. The first data determination module is used to determine, for the same location information, associated historical map data and M acquired image data based on a time series, where M is an integer greater than or equal to 1. The first encoding module is used to encode the historical map data and the M acquired image data, respectively, to obtain historical map features and M acquired image features. The target instance determination module is used to determine the location information and category information of the target instance based on the historical map features and the M acquired image features, where the target instance represents lane information. The update module is used to update the map data based on the location information and category information of the target instance.

[0009] According to another aspect of the present disclosure, a training device for a map update model is provided, comprising: a sample determination module, a second encoding module, a second data determination module, and a training module. The sample determination module is used to determine training samples, and the training samples include: a map prior, N sample images based on time series, and a label; the map prior and the N sample images are associated with the same location information, the label represents whether there is a real instance in the N sample images, and represents the real location information and real category of the real instance, the real instance represents lane information, and N is an integer greater than or equal to 1. The second encoding module is used to encode the map prior and the N sample images respectively to obtain map features and N sample image features. The second data determination module is used to determine the output location information and output category of at least one output instance based on the map features and the N sample image features. The training module is used to train the map update model based on the real location information, output location information, real category, and output category.

[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.

[0012] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided in the present disclosure when executed by a processor.

[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0015] FIG1 is a schematic diagram of an application scenario of a map update model training method, a map update method, and an apparatus according to an embodiment of the present disclosure;

[0016] FIG2 is a schematic flow chart of a method for training a map update model according to an embodiment of the present disclosure;

[0017] FIG3A is a schematic diagram of a lane line representation method according to an embodiment of the present disclosure;

[0018] FIG3B is a schematic diagram of a lane group representation method according to an embodiment of the present disclosure;

[0019] FIG4 is a schematic diagram of a training method for a map update model according to an embodiment of the present disclosure;

[0020] FIG5 is a schematic diagram of a map updating method according to an embodiment of the present disclosure;

[0021] FIG6 is a schematic flow chart of a map updating method according to an embodiment of the present disclosure;

[0022] FIG7 is a schematic structural block diagram of a training device for a map update model according to an embodiment of the present disclosure;

[0023] FIG8 is a schematic structural block diagram of a map updating device according to an embodiment of the present disclosure; and

[0024] FIG9 is a structural block diagram of an electronic device for implementing the map update model training method and / or map update method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In some embodiments, a labor-intensive mode can be used to update the map. In this mode, the entire process is divided into stages and elements, and trained operators discover and modify changed map elements.

[0027] The above-mentioned labor-intensive model has high costs, low efficiency, and uncertain production quality due to the level of the operators.

[0028] In other embodiments, maps can be updated based on a general image recognition algorithm combined with human interactive interpretation. This broadly encompasses two approaches. The first involves the model directly providing the modified results, typically through a multi-stage generation model based on image segmentation combined with post-processing. A segmentation model is used to obtain the pixel locations of line and surface features within the image. Post-processing strategies then extract vectorized line and surface information, which is then provided to operators to assist in interactive map generation. The second approach involves a change detection model directly providing a determination of whether a point of interest has changed, which is then provided to operators for modification assistance.

[0029] The first approach, combining the aforementioned general image recognition algorithm with human interaction, is limited by multi-stage error accumulation and a funneling effect on automation capabilities. For example, after the segmentation model outputs the mask, a series of operations such as fitting, line extraction, and thinning are required to extract lane lines, resulting in low recognition accuracy and recall. In the second approach, the information provided by the change detection model is crude and insufficient for achieving a highly automated, end-to-end model of all elements.

[0030] The disclosed embodiments aim to propose a map update model training method and a map update method. In some embodiments, this method uses existing maps as a priori information and performs training based on image data captured by a camera. By constructing an end-to-end generation framework and training paradigm, it achieves vectorized modeling and generation of map elements, thus avoiding the funnel effect. Furthermore, this method fully integrates existing maps and images, enabling high-fidelity map reconstruction while conveniently capturing changes in map elements, achieving end-to-end map updates.

[0031] The technical solutions provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] FIG1 is a schematic diagram of an application scenario of a map update model training method, a map update method, and an apparatus according to an embodiment of the present disclosure.

[0033] As shown in FIG1 , the application scenario 100 of this embodiment may include an electronic device 110 , which may be any electronic device with processing capabilities, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, a server, and the like.

[0034] The electronic device 110 may, for example, process the input captured image 140 and historical map data 150 to obtain a target instance 170 , and then use the target instance 170 to update the historical map data 150 .

[0035] According to an embodiment of the present disclosure, as shown in Figure 1, the application scenario 100 may further include a server 120. The electronic device 110 may be communicatively connected to the server 120 via a network, which may include a wireless or wired communication link.

[0036] For example, server 120 can be used to train a map update model 160 and, in response to a model acquisition request sent by electronic device 110, send the trained map update model 160 to electronic device 110 to facilitate map updates by electronic device 110. In one embodiment, electronic device 110 can also send captured images 140 to server 120 via a network, and the server can update the map based on the trained map update model 160.

[0037] According to an embodiment of the present disclosure, as shown in FIG1 , the application scenario 100 may further include a database 130 . The database 130 may maintain a large number of training samples, each of which may include a sample image, a map prior, and a label. The server 120 may access the database 130 and extract some training samples from the database 130 to train the map update model 160 .

[0038] When training the map update model 160, a loss function may be used to determine the loss of the map update model based on the output instances and labels, and the model training may be completed by minimizing the model loss.

[0039] It should be noted that the map update model training method provided in the present disclosure can be executed by the server 120, and the map update method provided in the present disclosure can be executed by the electronic device 110 or the server 120. Accordingly, the map update model training device provided in the present disclosure can be set in the server 120, and the map update device provided in the present disclosure can be set in the electronic device 110 or the server 120.

[0040] It should be understood that the number and types of electronic devices, servers, and databases in Figure 1 are merely illustrative. Depending on implementation requirements, any number and type of electronic devices, servers, and databases may be used.

[0041] FIG2 is a schematic flowchart of a method for training a map update model according to an embodiment of the present disclosure.

[0042] As shown in FIG. 2 , the map update model training method 200 may include operations S210 to S240 .

[0043] In operation S210 , a training sample is determined, where the training sample includes: a map prior, N sample images based on a time sequence, and a label, where N is an integer greater than or equal to 1.

[0044] In operation S220 , the map prior and the N sample images are encoded respectively to obtain map features and N sample image features.

[0045] In operation S230 , output position information and an output category of at least one output instance are determined based on the map feature and the N sample image features.

[0046] In operation S240 , a map update model is trained based on the real position information, the output position information, the real category, and the output category.

[0047] For example, multiple training samples are pre-constructed, and the training samples can be constructed based on manual annotation. Each training sample can include a map prior and at least one sample image. The map prior represents the map information of a certain geographical area, and the at least one sample image is collected using a camera located in the geographical area, so that the map prior and the N sample images are all associated with the same location information. The sample image may include at least one real instance or may not include a real instance. The real instance represents lane information. Lane information may include line elements and surface elements. Line elements may include lane lines. Lane lines may include single solid lines, double yellow lines, etc. Surface elements may include lane groups, zebra crossings, etc. For example, if a sample image includes two lane lines, then the sample image includes two real instances.

[0048] Each training sample can also include a label that indicates whether a real instance exists in the N images. Furthermore, if a real instance exists in the sample image, the label also indicates the real location information and real category of the real instance. Multiple points can be used to represent a real instance, so the real location information can include the coordinates of multiple points, with each point corresponding to a piece of real location information. Real categories can include a single solid white line, a dashed white line, a double yellow line, and so on.

[0049] For example, a map prior and N sample images are input into a map update model to be trained. The map update model then performs encoding, feature fusion, decoding, and other operations. The map update model then outputs a prediction result, which can include the output location information and output category of the output instance. The output location information is similar to the actual location information and can include the coordinates of only one point or multiple points. The output category can refer to the actual category.

[0050] Next, a predetermined loss function can be used to calculate a total loss function value based on the difference between the actual location information and the output location information, as well as the difference between the actual category and the output category. The network weights in the map update model can be adjusted using a backpropagation algorithm to complete training of the map update model. The predetermined loss function may include a cross-entropy loss function, etc., which is not limited in this disclosure.

[0051] The disclosed embodiments train a map update model based on sample images and map priors, enabling end-to-end vectorized map feature generation. This method transforms conventional automated map update methods by encoding and learning map priors to ensure the fidelity of map reconstruction. Furthermore, it directly uses serialized sample images as input, reducing process steps and improving update efficiency.

[0052] Before training the map update model, it is necessary to pre-build training samples. This embodiment adopts a solution that directly converts the results of manual work into a training data structure. The process of automatically building training samples will be described in detail below.

[0053] First, we introduce how to obtain sample images in training samples. For example, we can obtain sample images, location information, and equipment parameters based on self-collecting vehicles, that is, Where N represents the total number of sample images based on the time series. Represents the i-th frame sample image based on time sequence. i Indicates the camera's position information when collecting sample images. iRepresents the parameters of the camera used to capture sample images. It should be noted that a single sample image can be a surround view image. In this case, a single sample image can include K images of the same location from different perspectives. For example, six images from different perspectives are acquired at the first moment, constituting the first sample image frame. Subsequently, six images from different perspectives are acquired at the second moment, constituting the second sample image frame. In other embodiments, a single sample image can also include only one image from a single perspective. The above describes the method for obtaining sample images in training samples.

[0054] Then, based on the above position information Pos i , extract the location information Pos from the pre-built database i The associated map data, the map data in the database is the result of pre-marking by production workers (such as lane lines, lane groups, etc.). It should be noted that the map data obtained from the database includes two versions of map data, the new version of the map data is consistent with the sample image, and the old version of the map data is inconsistent with the sample image. For example, if there are 4 lane lines in the sample image, the new version of the map data also needs to correspond to the 4 lane lines, while the old version of the map data may only include 3 lane lines. The training sample includes a sample image, a map prior, and a label, wherein the sample image is acquired through a camera, the new version of the map data is used to determine the label, and the old version of the map data is used to determine the map prior in the training sample.

[0055] Next, the process of determining labels based on the new version of map data will be described.

[0056] For the new version of map data, we can first extract the manual work results of the lane lines based on the geometry (such as the coordinates of the multiple points of the lane lines), style (such as straight line, curve, white, yellow) and other category information. Through the above operation, we can get some original vectorized point sets and the style and color information corresponding to each original vectorized point set. Each point set represents the location information of an instance, and the style and color information corresponding to the point set is the category of the instance. In this way, we can get the original vectorized point set. and categories Among them, N ori Indicates the number of real instances in a single frame sample image, represents the true category of a single instance, and c is the total number of actual categories.

[0057] Get the original vectorized point set Afterwards, for each point set P i ∈P ori Perform fitting and interpolation to obtain a uniform fixed number of points where N pIndicates the number of points. If the number of instance pixels is less than N p , then the actual number of pixels can be used. Considering the consistency of model building, this embodiment fixes the number of instances of a single sample image to N ins , the number N ins is a preset value, which can be larger than the number of instances of a regular image, such as N ins Take 50. For instances with number less than N ins The image can be filled with the point set P by padding ori , where the filling part is represented by This results in a point set representation of a single image And the corresponding categories The label of each instance includes the location information expressed by the above point set and the true category corresponding to the point set, so the label of each instance (Ground True, GT) can be expressed as Y i =(P i , L i ). It should be noted that the above point set expression In the dataset, part of the point set is obtained by filling, and the other part of the point set exists before filling. The label needs to indicate whether the point set is obtained by filling, so that the loss function value can be accurately calculated later.

[0058] The above describes how to obtain labels in training samples. Next, we will explain the process of determining map priors based on old versions of map data.

[0059] For the old version of the map data, a similar method is used to process the new version of the map data to obtain the point sets and categories in the old version of the map data. Next, the point sets in the old version of the map data need to be intuitively represented and processed into a mode that is convenient for the map update model to receive. One way to process is to map the point set to a mask That is, the point set is mapped to a mask with a background of predetermined pixel value with a width of h according to its category. Its value range is 0, 1..., c, for example, 1 represents a single solid line, 2 represents a double yellow line, H represents a double yellow line, and H represents a double yellow line. gt and W gt Representing the height and width of the map prior, respectively, the predetermined pixel value can be 0. For example, a sample image includes three lane lines, each of which is a real instance, corresponding to a class, and represented by a set of 50 points. For each real instance, the 50 points of the real instance are mapped to a mask with a black background. The width h of each point is a preset value. This results in three lane lines on the black background mask, thus obtaining the map prior.

[0060] The above describes how to obtain sample images, map priors, and labels in training samples. Next, we will explain how to represent lane information.

[0061] In this embodiment, lane information includes line elements and surface elements. Line elements may include lane lines, and surface elements may include lane groups. Each single lane information may be represented by multiple points.

[0062] Figure 3A is a schematic diagram of a lane line representation method according to an embodiment of the present disclosure, and Figure 3B is a schematic diagram of a lane group representation method according to an embodiment of the present disclosure. The point with a black background in the figure represents the starting point, and the arrow represents the direction.

[0063] As shown in Figure 3A , the point set corresponding to the lane line includes multiple points arranged linearly. Since the lane line does not need to distinguish directions, the start and end points do not need to be consistent. The two equivalent arrangements 311 and 312 in Figure 3A can both represent lane lines.

[0064] As shown in Figure 3B, the point set corresponding to a lane group includes multiple points, and the order of the points does not affect the instance representation. Therefore, a lane group consisting of k points can be represented using 2×k equivalent permutations. For example, in Figure 3B, three points are used to represent a lane group. There are six possible representations (321–326) for a lane group, and any of these representations can be used to represent the lane group.

[0065] Next, we will introduce the structure and training algorithm of the map update model.

[0066] In this embodiment, the map update model may include an image encoder E bev , map prior encoder E map , confidence map network and decoder, and multiple task modules can be connected to the output of the decoder. The image encoder can use a BEV encoder.

[0067] First, obtain pre-built training samples, which include map priors And N sample images based on time series Map prior Input map prior encoder E map , map prior encoder E map Map it to the feature space and get the map feature F map . Input the sample image into the image encoder E bev , image encoder E bev Map it to the feature space to obtain the sample image features

[0068] Next, N second confidence maps corresponding to the N sample image features can be determined based on the N sample image features. Input the confidence map network to obtain N second confidence maps The sample image features correspond one-to-one to the second confidence maps and have the same size. Each sample image feature includes multiple sub-features, and each sub-feature can represent a pixel point in the sample image or multiple points in an area. Accordingly, each second confidence map includes multiple confidences. The multiple confidences in the second confidence map represent the weights of the multiple sub-features in the corresponding sample image features, for example, the first confidence represents the weight of the first sub-feature.

[0069] Next, output position information and output category of at least one output instance may be determined based on the N sample image features, the N second confidence maps, and the map features.

[0070] In one embodiment, N sample image features and N second confidence maps may be fused to obtain a second intermediate fusion feature F sequence The fusion process may include: performing a dot multiplication process on the corresponding sample image features and the second confidence map, and then adding them frame by frame to obtain the fused second intermediate fusion feature F sequence For example, a sample image is collected at multiple moments, and each sample image corresponds to a sample image feature. The confidence map network is used to learn the importance of each sub-feature in each sample image feature. For example, the points near the lane line in the sample image are more important. If the importance of a point is high, the confidence map network outputs a larger weight for the point. Then, based on the confidence map network, a second confidence map is output, and based on the second confidence map, multiple sample image features are fused into a second intermediate fusion feature F. sequence .

[0071] Then we can use the second intermediate fusion feature F sequence and map feature F map , determine the second target fusion feature F all =Cat(F sequence , F map ), for example, the second intermediate fusion feature F sequence and map feature F map Splicing along the channel dimension to obtain the second target fusion feature F all =Cat(F sequence , F map ). For example, the intermediate fusion feature F sequence The size is C1*H*W, the map feature F map The size of the second target fusion feature F after splicing is C2*H*W all=Cat(F sequence , F map ) is (C1+C2)*H*w. Where C represents the number of channels, H represents the height, and W represents the width.

[0072] Then the second target feature F can be fused all =Cat(F sequence , F map ) is decoded to obtain the output position information and output category of at least one output instance. For example, the second target fusion feature F all =Cat(F sequence , F map ) is input into the decoder to obtain the output position information and output category confidence of each output instance. The decoder can include multiple network layers (transformer layer), which learns instance-level and point-level context information and combines the multi-task learning framework of semantic segmentation and instance segmentation to obtain the output result. in Indicates the output location information of the output instance. Indicates multiple output categories, It represents the confidence of each output category, and the output category with the maximum confidence can be used as the output category determined by the model.

[0073] After obtaining the output location information, output category, and confidence of the output category for each output instance, the total loss function value can be calculated based on the difference between the output location information and the true location information, as well as the difference between the output category and the true category. The parameters of the map update model can be adjusted by minimizing the total loss function value.

[0074] This embodiment considers the consistency of multi-frame, multi-angle sample images and simultaneously learns point-level confidence maps, enabling adaptive fusion of multiple frames to improve generation quality. Furthermore, by using temporal information to improve visual feature quality, it can effectively alleviate issues such as occlusion and irregular truncation. By learning intuitive visual representations and manually defined rule representations, it avoids multi-stage processes that require separate links and elements. This improves end-to-end automation capabilities across all elements and addresses the funnel effect.

[0075] In addition, by fusing multiple sample images based on the second confidence map, sub-features with large weights can have a greater impact on the second intermediate fusion feature, thereby improving the representation accuracy of the second target fusion feature and further improving the model training effect.

[0076] It should be noted that in other embodiments, the confidence map network can be omitted and multiple sample image features can be directly fused, for example, corresponding sub-features can be fused to obtain fused features, and then the fused features are fused again with the map features and then input into the decoder together.

[0077] It should be noted that during the training process, the output of the map update model includes multiple output instances. Therefore, the model output information and labels can be matched to obtain a matching relationship, and the total loss function value can be accurately calculated based on the matching relationship to ensure the model iteration effect.

[0078] In one example, the matching relationship includes a first matching relationship, where the first matching relationship represents a matching relationship between a real instance and an output instance, that is, which real instance the output instance matches.

[0079] The first matching relationship can be determined by arranging at least one real instance to obtain a real instance sequence. Also, arranging and combining at least one output instance to obtain multiple candidate instance sequences. Subsequently, for each candidate output instance sequence, a first difference evaluation value for the candidate instance sequence is determined based on the output position information and output category of each output instance in the candidate output instance sequence, the real position information and real category of each real instance in the real instance sequence, and the order of the real instance sequence and the order of each candidate instance sequence. In this way, multiple candidate instance sequences correspond to multiple first difference evaluation values, and the first matching relationship is then determined based on the minimum first difference evaluation value.

[0080] For example, use Represents a set of candidate instance sequences, which includes multiple candidate instance sequences, each candidate instance sequence includes multiple output instances arranged in sequence, and the instance arrangement order of multiple candidate instance sequences is different from each other. The i-th real instance in the real instance sequence and the i-th output instance in the candidate instance sequence form an instance pair, so that the sum of the differences between the multiple instance pairs is determined as the first difference evaluation value. The candidate instance matching sequence corresponding to the minimum first difference evaluation value It can be expressed as:

[0081] Among them, cost ins Represents the matching loss term between instances, which includes at least one of the category loss function and the geometric loss function. The following formula is based on cost ins Including category loss function and geometric loss function as examples.

[0082] Among them, cost clsRepresents the category loss function, and focal loss can be used as the category loss function. cost geo represents a geometric loss function, which is used to measure the geometric correlation between the point set of the output instance and the point set of the real instance. The Hungarian matching algorithm can be used as the geometric loss function. This embodiment does not limit the category loss function and the geometric loss function.

[0083] This embodiment determines the first matching relationship, so that the total loss function value can be calculated based on the first matching relationship, thereby improving the model training effect and ensuring the model reasoning accuracy.

[0084] In another example, the real location information of a single real instance includes multiple real sub-location information. Correspondingly, the output location information of a single output instance includes multiple output sub-location information. It is understood that both the real instance and the output instance are represented by multiple points, each real sub-location information can represent the coordinates of a point in the real instance, and each output sub-location information can represent the coordinates of a point in the output instance. The matching relationship includes a second matching relationship, which represents the matching relationship between the real sub-location information and the output sub-location information, i.e., which point in the output instance matches the point in the real instance.

[0085] The second matching relationship can be determined by: for an output instance and a real instance that satisfy the first matching relationship, the multiple real sub-location information of the real instance is arranged to obtain a real point sequence. The multiple output sub-location information of the output instance is also arranged and combined to obtain multiple candidate point sequences. A second difference evaluation value is then determined for each candidate point sequence based on the multiple real sub-location information, the multiple output sub-location information, the order of the real point sequence, and the order of the candidate point sequences. In this manner, multiple candidate point sequences correspond to multiple second difference evaluation values. The second matching relationship is then determined based on the minimum second difference evaluation value.

[0086] For example, use Represents a set of candidate point sequences, which includes multiple candidate point sequences. Each candidate point sequence includes multiple points in an output instance, and the order of points in multiple candidate point sequences is different from each other. The i-th point in the real point sequence and the i-th point in the candidate point sequence form a point pair, so that the sum of the differences between the multiple point pairs is determined as the second difference evaluation value. The candidate point sequence corresponding to the minimum second difference evaluation value It can be expressed as:

[0087] Where cost pointThe term "matching loss" represents the matching loss between the points in the output instance and the points in the real instance, that is, the matching loss between the output sub-location information and the real sub-location information. Manhattan distance can be used as a function of the matching loss. This embodiment does not limit the function of the matching loss.

[0088] This embodiment determines the second matching relationship, so that the total loss function value can be subsequently calculated based on the second matching relationship, thereby improving the model training effect and ensuring the model reasoning accuracy.

[0089] The above describes the first matching relationship at the instance level and the second matching relationship at the point level within the instance. In practical applications, the matching relationship includes at least one of the first and second matching relationships. After obtaining the matching relationship, the total loss function value can be determined based on the matching relationship, the true location information, the output location information, the true category, and the output category. The map update model is then trained based on the total loss function value.

[0090] It should be noted that in other embodiments, other methods may be used to determine the above-mentioned matching relationship. For example, the first matching relationship may be determined based on the relative positional relationship between each output instance and the relative positional relationship between each real instance. Alternatively, based on the output position information of the output instance and the real position information of the real instance, multiple output instances and multiple real instances may be numbered from top to bottom and from left to right, and the output instance and real instance with the same number may be determined to satisfy the first matching relationship.

[0091] The loss function used in the model training process is explained below.

[0092] In this embodiment, multiple predetermined loss function values ​​can be calculated, and then the total loss function value can be calculated based on the multiple predetermined loss functions. The multiple predetermined loss function values ​​can include point-level geometric loss function values, instance-level geometric loss function values, category loss function values, binary classification loss function values, semantic segmentation loss function values ​​and instance segmentation loss function values. Each predetermined loss function is explained below.

[0093] For the point-level geometric loss function, the output sub-position information and the true sub-position information that satisfy the second matching relationship can be formed into a point pair, and the sum of the distances between each point pair is calculated to determine the point-level geometric loss function value.

[0094] For example, the point-level geometric loss function is as follows:

[0095] Where dist represents the distance between points, and Manhattan distance can be used. ins and N p Represent the number of instances and the number of points within a single instance, Li Indicates the instance category.

[0096] For the instance-level geometric loss function value, the instance-level geometric loss function value can be determined based on the difference between the first line segment and the second line segment in the line segment pair; wherein, the first line segment has two output sub-position information as endpoints, and the second line segment has two real sub-position information as endpoints; in the same line segment pair, the two endpoints of the first line segment and the two endpoints of the second line segment satisfy the second matching relationship.

[0097] For example, the instance-level geometric loss function is as follows:

[0098] Where cosine represents similarity, for example, cosine similarity is used, link represents a line segment connecting two points, and the line connecting two adjacent points can be used as a line segment. For example, if the real instance includes 50 real sub-location information, 49 line segments can be determined. The output instance is similar and will not be repeated here. ins and N p Represent the number of instances and the number of points within a single instance, L i Indicates the instance category.

[0099] For the category loss function, the function value can be calculated in the following way: for the output instance and the real instance that satisfy the first matching relationship, the category loss function value is determined according to the difference between the output category of the output instance and the real category of the real instance.

[0100] For example, the category loss function is as follows:

[0101] in Represents loss, for example, using focal loss, and L i Represent the output category and the true category, respectively. The category probability can be used as a confidence level. For categories specified by confidence levels above a threshold, this level is used as the output category. Furthermore, this embodiment incorporates the instance geometry positional relationship into the category constraint for task alignment, allowing the confidence level to simultaneously represent both geometry and category.

[0102] Furthermore, due to the complex nature of road conditions, such as surface wear, occlusion, and sudden changes in illumination, the map update model can easily output redundant and inconsistent feature vectors. Therefore, to mitigate this issue from a model perspective, this embodiment introduces a binary classification loss function as a constraint.

[0103] The binary classification loss function can be calculated as follows: for each output instance, the binary classification loss function value is determined based on whether there is a real instance that satisfies the first matching relationship with the output instance. In other words, predicted instances that do not match real instances are classified as invalid, and predicted instances that do match real instances are classified as valid. The valid and invalid classes have different binary classification loss function values.

[0104] For example, the binary classification loss function is as follows:

[0105] in, represents the binary cross entropy loss function, L bce(i) Represents the true value, which is 0 or 1, where L is the output instance that matches the real instance bce(i) =1, L of the output instance that does not match the real instance bce(i) is 0.

[0106] This embodiment also introduces semantic segmentation tasks and instance segmentation tasks, both of which implement feature-level constraints by determining the segmentation masks of map elements, thereby improving the model prediction accuracy and recall.

[0107] For the semantic segmentation loss function, the output position information of at least one output instance and the real position information of at least one real instance can be processed based on semantic segmentation to obtain a first output mask and a first real mask respectively; and the semantic segmentation loss function value is determined according to the difference between the first output mask and the first real mask.

[0108] For the instance segmentation loss function, the output position information of at least one output instance and the real position information of at least one real instance can be processed based on the instance segmentation to obtain a second output mask and a second real mask respectively; according to the difference between the second output mask and the second real mask, the instance segmentation loss function value is determined.

[0109] For example, the semantic segmentation loss function and instance segmentation loss function are as follows:

[0110] in and M represent the output mask and the true mask respectively, the subscript sem represents semantic segmentation, and the subscript ins represents instance segmentation. represents the binary cross entropy loss, represents the damage function (Dice loss).

[0111] After obtaining the above-mentioned point-level geometric loss function value, instance-level geometric loss function value, category loss function value, binary classification loss function value, semantic segmentation loss function value and instance segmentation loss function value, the weighted sum of the above-mentioned predetermined loss function values ​​can be calculated as the total loss function.

[0112] For example, the total loss function is as follows:

[0113] Where λ represents the weight coefficient.

[0114] This embodiment uses the above six predetermined loss functions to calculate the total loss function, thereby constraining the training of the map update model from multiple dimensions, so that the output results of the map update model are more accurate and have stronger consistency.

[0115] It should be noted that the above introduces 6 predetermined loss functions. In practical applications, you can select 1 to 6 of them at will, and calculate the weighted sum of the selected predetermined loss function values ​​as the total loss function value.

[0116] It should be noted that when calculating the above six predetermined loss functions, functions such as Manhattan distance, cosine similarity, focus loss, binary cross entropy loss function and damage function are used. In actual applications, these functions can be adjusted according to actual needs. This embodiment does not limit the functions used.

[0117] It should be noted that the difference between semantic segmentation and instance segmentation is that, for example, if there are three lane lines and the first and second lane lines have the same category, in semantic segmentation, the first and second lane lines are in the same mask, while in instance segmentation, the first and second lane lines are in two different masks.

[0118] Next, with reference to Figure 4, we will introduce the structure and training algorithm of the map update model.

[0119] FIG4 is a schematic diagram of a method for training a map update model according to an embodiment of the present disclosure.

[0120] As shown in Figure 4, the map update model may include an image encoder 407, a map prior encoder 409, a confidence map network 408 and a decoder 410. Multiple task modules may also be provided at the output end of the decoder 410. For example, the multiple task modules may include a regression module 411, a classification module 412 and a segmentation module 413. The BEV encoder may be used as the image encoder 407.

[0121] First, a pre-built training sample is obtained. The training sample includes a map prior 405 and N sample images 401 based on a time sequence. The map prior 405 is input into a map prior encoder 409, which outputs map features 406. The N sample images 401 are input into an image encoder 407, which outputs N sample image features 402.

[0122] Next, N sample image features 402 are input into the confidence map network 408 to generate N second confidence maps 403. The sample image features 402 and the second confidence maps 403 have a one-to-one correspondence and are of the same size. A dot product is then performed on the corresponding sample image features 402 and the second confidence maps 403. After the dot product, these are summed frame by frame to generate a fused second intermediate fused feature 404.

[0123] Then, the second intermediate fusion feature 404 and the map feature 406 are concatenated along the channel dimension to obtain the second target fusion feature.

[0124] The second target fusion feature is then input into the decoder 410 to obtain the output position information 407 of each output instance and the output category and the confidence of each output category.

[0125] The regression module 411 can determine the point set corresponding to each output instance based on the output position information 407 in order to calculate the point-level geometric loss function value and the instance-level geometric loss function value. The classification module 412 is used to perform category classification in order to calculate the category loss function value. The classification module 412 is also used to perform valid and invalid binary classification in order to calculate the binary classification loss function value. The segmentation module 413 is used to perform semantic segmentation processing and instance segmentation processing in order to calculate the semantic segmentation loss function value and the instance segmentation loss function value. Through the above-mentioned multiple predetermined loss function values, the total loss function value can be calculated, and then the parameters of the map update model can be updated. For example, the parameters of the image encoder 407, the map prior encoder 409, the confidence map network 408 and the decoder 410 are updated.

[0126] FIG5 is a schematic diagram of a map updating method according to an embodiment of the present disclosure.

[0127] As shown in Figure 5, the map update method includes the following operations: For the same location information, associated historical map data 502 and M time-series captured image data 501 are determined, where M is an integer greater than or equal to 1 and can be the same as or different from N described above. The historical map data 502 and the M captured image data 501 are then input into a trained map update model 503 to obtain the location information and category of the target instance. The historical map data is then updated based on the location information and category of the target instance, thereby obtaining updated map data 504.

[0128] In this embodiment, the map update model 503 may be trained using the training method described above. The data processing process of the map update model 503 may include encoding the historical map data 502 and the M collected image data 501 to obtain historical map features and the M collected image features, and then determining the location information and category information of the target instance based on the historical map features and the M collected image features.

[0129] The map updating method is described in detail below with reference to FIG6 .

[0130] FIG6 is a schematic flowchart of a map updating method according to an embodiment of the present disclosure.

[0131] As shown in FIG. 6 , the map updating method 600 may include operations S610 to S640 .

[0132] In operation S610 , for the same location information, associated historical map data and M time-series collected image data are determined, where M is an integer greater than or equal to 1.

[0133] In operation S620, the historical map data and the M collected image data are respectively encoded to obtain historical map features and M collected image features.

[0134] In operation S630 , location information and category information of a target instance are determined based on the historical map features and the M collected image features, where the target instance represents lane information.

[0135] In operation S640 , map data is updated based on the location information and category information of the target instance.

[0136] For example, a trained map update model includes an image encoder, a map prior encoder, and a decoder. Historical map data and M collected image data can be input into the image encoder and the map prior encoder, respectively, to obtain historical map features and M collected image features. These historical map features and M collected image features can then be fused and fed into the decoder to obtain the location and category information of the target instance.

[0137] It should be noted that the number of target instances can be one or more, and the location information of each target instance includes the coordinates of multiple points. The category information may include a single solid line, a double yellow line, etc. The location information can be mapped to the historical map data, and each point can be displayed as an indication of the category of the target instance. Multiple points can also be fitted into lines to achieve the effect of updating the map.

[0138] This disclosed embodiment uses existing historical map data as a prior and captured images as input, achieving end-to-end geographic feature generation based on a map update model. This includes, but is not limited to, lines (lane lines), surfaces (zebra crossings), and manually defined rules (lane groups). Using existing historical map data as a prior allows for high-fidelity map reconstruction. By directly generating vectorized target results from captured images as input, it can fully learn the visual representation of complex geographic scenes, avoiding a series of pre- and post-processing operations, thereby improving the final recognition effect.

[0139] According to another embodiment of the present disclosure, the map update model further includes a confidence map network. Accordingly, the process of determining the location information and category information of the target instance based on the historical map features and the M acquired image features may include: determining M first confidence maps corresponding to the M acquired image features respectively based on the M acquired image features. For example, the M acquired image features are input into the confidence map network, and the confidence map network outputs M first confidence maps. It should be noted that, similar to the training process, each acquired image feature includes multiple sub-features, and each first confidence map includes multiple confidences. The multiple confidences in the first confidence map represent the weights of the multiple sub-features in the corresponding acquired image features. Subsequently, the location information and category information of the target instance may be determined based on the M acquired image features, the M first confidence maps, and the historical map features.

[0140] In one example, M acquired image features and M first confidence maps can be fused to obtain a first intermediate fusion feature. For example, the corresponding image features and the second confidence map are subjected to dot multiplication, and after the dot multiplication, they are added frame by frame to obtain the fused first intermediate fusion feature. Next, the first target fusion feature can be determined based on the first intermediate fusion feature and the historical map feature. For example, the first intermediate fusion feature and the historical map feature can be spliced ​​along the channel dimension to obtain the first target fusion feature. The first target fusion feature is then decoded to obtain the location information and category information of the target instance. For example, the first target fusion feature can be input into a decoder to obtain the output location information and the confidence of the output category of each target instance.

[0141] This embodiment determines the weights of sub-features through a confidence map network and fuses multiple sample images based on the second confidence map, so that sub-features with large weights can have a greater impact on the first intermediate fusion feature, thereby improving the representation accuracy of the first target fusion feature and further improving the model training effect.

[0142] FIG7 is a schematic structural block diagram of a training device for a map update model according to an embodiment of the present disclosure.

[0143] As shown in FIG. 7 , the map updating apparatus 700 may include a first data determining module 710 , a first encoding module 720 , a target instance determining module 730 and an updating module 740 .

[0144] The first data determination module 710 is used to determine, for the same location information, associated historical map data and M time-series collected image data, where M is an integer greater than or equal to 1.

[0145] The first encoding module 720 is used to encode the historical map data and the M collected image data respectively to obtain historical map features and M collected image features.

[0146] The target instance determination module 730 is used to determine the location information and category information of the target instance based on the historical map features and the M collected image features. The target instance represents lane information.

[0147] The updating module 740 is used to update the map data according to the location information and category information of the target instance.

[0148] In this embodiment, the target instance determination module includes: a first determination submodule and a second determination submodule. The first determination submodule is configured to determine, based on M acquired image features, M first confidence maps corresponding to the M acquired image features, respectively; each acquired image feature includes multiple subfeatures, each first confidence map includes multiple confidence levels, and the multiple confidence levels in the first confidence maps represent the weights of the multiple subfeatures in the corresponding acquired image feature. The second determination submodule is configured to determine the location information and category information of the target instance based on the M acquired image features, the M first confidence maps, and the historical map features.

[0149] In this embodiment, the second determination submodule includes: a first fusion unit, a first determination unit, and a first decoding unit. The first fusion unit is configured to fuse the M acquired image features with the M first confidence maps to obtain a first intermediate fused feature. The first determination unit is configured to determine a first target fused feature based on the first intermediate fused feature and the historical map features. The first decoding unit is configured to decode the first target fused feature to obtain location information and category information of the target instance.

[0150] FIG8 is a schematic structural block diagram of a map updating apparatus according to an embodiment of the present disclosure.

[0151] As shown in FIG. 8 , the map update model training device 800 may include a sample determination module 810 , a second encoding module 820 , a second data determination module 830 and a training module 840 .

[0152] The sample determination module 810 is used to determine training samples, which include: map priors, N sample images based on time series, and labels; the map priors and the N sample images are all associated with the same location information, the labels represent whether there are real instances in the N sample images, and represent the real location information and real category of the real instances, and the real instances represent lane information, where N is an integer greater than or equal to 1.

[0153] The second encoding module 820 is used to encode the map prior and the N sample images respectively to obtain map features and N sample image features.

[0154] The second data determination module 830 is configured to determine output location information and an output category of at least one output instance according to the map features and the N sample image features.

[0155] The training module 840 is used to train the map update model based on the real position information, the output position information, the real category and the output category.

[0156] In this embodiment, the second data determination module includes: a third determination submodule and a fourth determination submodule. The third determination submodule is configured to determine, based on N sample image features, N second confidence maps corresponding to the N sample image features, respectively; each sample image feature includes multiple subfeatures, each second confidence map includes multiple confidence levels, and the multiple confidence levels in the second confidence maps represent the weights of the multiple subfeatures in the corresponding sample image feature. The fourth determination submodule is configured to determine the output location information and output category of at least one output instance based on the N sample image features, the N second confidence maps, and the map features.

[0157] In this embodiment, the fourth determination submodule includes: a second fusion unit, a second determination unit, and a second decoding unit. The second fusion unit is configured to fuse the N sample image features and the N second confidence maps to obtain a second intermediate fused feature. The second determination unit is configured to determine a second target fused feature based on the second intermediate fused feature and the map features. The second decoding unit is configured to decode the second target fused feature to obtain output location information and output category of at least one output instance.

[0158] In this embodiment, the training module includes: a matching submodule, a total loss determination submodule, and a training submodule. The matching submodule is used to determine a matching relationship, which includes at least one of a first matching relationship and a second matching relationship. The first matching relationship represents the matching relationship between the real instance and the output instance. The real location information of the same real instance includes multiple real sub-location information, and the output location information of the same output instance includes multiple output sub-location information. The second matching relationship represents the matching relationship between the real sub-location information and the output sub-location information. The total loss determination submodule is used to determine the total loss function value based on the matching relationship, the real location information, the output location information, the real category, and the output category. The training submodule is used to train the map update model based on the total loss function value.

[0159] In this embodiment, the matching submodule includes: a first arrangement unit, a second arrangement unit, a first difference determination unit and a first relationship determination unit. The first arrangement unit is used to arrange at least one real instance to obtain a real instance sequence. The second arrangement unit is used to arrange and combine at least one output instance to obtain multiple candidate instance sequences. The first difference determination unit is used to determine, for each candidate output instance sequence, a first difference evaluation value for the candidate instance sequence based on the output position information and output category of each output instance in the candidate output instance sequence, the real position information and real category of each real instance in the real instance sequence, the order of the real instance sequence and the order of each candidate instance sequence. The first relationship determination unit is used to determine a first matching relationship based on the minimum first difference evaluation value among multiple first difference evaluation values ​​for multiple candidate instance sequences.

[0160] In this embodiment, the matching submodule includes: a second difference determination unit and a second relationship determination unit. The second difference determination unit is configured to, for an output instance and a real instance that satisfy the first matching relationship, arrange multiple real sub-location information of the real instance to obtain a real point sequence; arrange and combine multiple output sub-location information of the output instance to obtain multiple candidate point sequences; and determine multiple second difference evaluation values ​​for the multiple candidate point sequences based on the order of the multiple real sub-location information, the multiple output sub-location information, the real point sequence, and the candidate point sequence. The second relationship determination unit is configured to determine a second matching relationship based on the minimum second difference evaluation value among the multiple second difference evaluation values.

[0161] In this embodiment, the total loss determination submodule includes: a total loss determination unit, which is used to determine the total loss function value according to a predetermined loss function value; wherein the process of determining the predetermined loss function value includes at least one of the following: determining the point-level geometric loss function value according to the distance between the output sub-position information and the real sub-position information that satisfies the second matching relationship; determining the instance-level geometric loss function value according to the difference between the first line segment and the second line segment in the line segment pair; wherein the first line segment has two output sub-position information as endpoints, and the second line segment has two real sub-position information as endpoints; in the same line segment pair, the two endpoints of the first line segment and the two endpoints of the second line segment satisfy the second matching relationship; for the output instance and the real instance that satisfy the first matching relationship, The value of the category loss function is determined based on the difference between the category and the real category of the real instance; for each output instance, the value of the binary classification loss function is determined according to whether there is a real instance that satisfies the first matching relationship with the output instance; based on semantic segmentation, the output position information of at least one output instance and the real position information of at least one real instance are processed to obtain a first output mask and a first real mask respectively; according to the difference between the first output mask and the first real mask, the value of the semantic segmentation loss function is determined; and based on instance segmentation, the output position information of at least one output instance and the real position information of at least one real instance are processed to obtain a second output mask and a second real mask respectively; according to the difference between the second output mask and the second real mask, the value of the instance segmentation loss function is determined.

[0162] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0163] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0164] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0165] FIG9 is a block diagram of an electronic device used to implement the map update model training method and / or map update method of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0166] As shown in Figure 9, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0167] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0168] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the above-mentioned method. For example, in some embodiments, the above-mentioned method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the above-mentioned method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the above-mentioned method in any other appropriate manner (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0173] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0174] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0175] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0176] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0177] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0178] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A map updating method, comprising: For the same location information, determine the associated historical map data and M collected image data based on the time series, where M is an integer greater than or equal to 1; Encoding the historical map data and the M collected image data respectively to obtain historical map features and M collected image features; Determine location information and category information of a target instance according to the historical map features and the M collected image features, wherein the target instance represents lane information; as well as The map data is updated according to the location information and category information of the target instance.

2. The method according to claim 1, wherein: Determining the location information and category information of the target instance according to the historical map features and the M collected image features includes: According to the M acquired image features, M first confidence maps corresponding to the M acquired image features are determined respectively; each acquired image feature includes a plurality of sub-features, each first confidence map includes a plurality of confidences, and the plurality of confidences in the first confidence map represent weights of the plurality of sub-features in the corresponding acquired image feature; and The location information and category information of the target instance are determined according to the M collected image features, the M first confidence maps and the historical map features.

3. The method according to claim 2, wherein: The determining the location information and category information of the target instance according to the M collected image features, the M first confidence maps, and the historical map features includes: Fusing the M acquired image features and the M first confidence maps to obtain a first intermediate fused feature; determining a first target fusion feature according to the first intermediate fusion feature and the historical map feature; and The first target fusion feature is decoded to obtain location information and category information of the target instance.

4. A method for training a map update model, comprising: Determine a training sample, wherein the training sample includes: a map prior, N sample images based on a time series, and a label; the map prior and the N sample images are associated with the same position information, the label represents whether a real instance exists in the N sample images and represents the real position information and the real category of the real instance, the real instance represents lane information, and N is an integer greater than or equal to 1; Encoding the map prior and the N sample images respectively to obtain map features and N sample image features; Determining output location information and output category of at least one output instance according to the map feature and the N sample image features; and The map update model is trained according to the real position information, the output position information, the real category and the output category.

5. The method according to claim 4, wherein: The determining, according to the map feature and the N sample image features, output location information and output category of at least one output instance comprises: According to the N sample image features, determine N second confidence maps corresponding to the N sample image features respectively; each sample image feature includes a plurality of sub-features, each second confidence map includes a plurality of confidences, and the plurality of confidences in the second confidence map represent weights of the plurality of sub-features in the corresponding sample image feature; and The output position information and output category of the at least one output instance are determined according to the N sample image features, the N second confidence maps, and the map features.

6. The method according to claim 5, wherein: Determining the output position information and output category of the at least one output instance according to the N sample image features, the N second confidence maps, and the map features comprises: Fusing the N sample image features and the N second confidence maps to obtain a second intermediate fused feature; determining a second target fusion feature according to the second intermediate fusion feature and the map feature; and The second target fusion feature is decoded to obtain output position information and output category of the at least one output instance.

7. The method according to claim 4, wherein: The training of the map update model according to the real position information, the output position information, the real category and the output category comprises: Determine a matching relationship, the matching relationship including at least one of a first matching relationship and a second matching relationship; the first matching relationship represents a matching relationship between a real instance and an output instance; the real location information of the same real instance includes a plurality of real sub-location information, the output location information of the same output instance includes a plurality of output sub-location information, and the second matching relationship represents a matching relationship between the real sub-location information and the output sub-location information; Determine a total loss function value according to the matching relationship, the real position information, the output position information, the real category and the output category; and According to the total loss function value, the map is trained to update the model.

8. The method according to claim 7, wherein: Determining the matching relationship includes: Arranging the at least one real instance to obtain a real instance sequence; Arrange and combine the at least one output instance to obtain multiple candidate instance sequences; For each candidate output instance sequence, determining a first difference evaluation value for the candidate instance sequence according to output position information and output category of each output instance in the candidate output instance sequence, real position information and real category of each real instance in the real instance sequence, and the order of the real instance sequence and the order of each candidate instance sequence; and The first matching relationship is determined according to a minimum first difference evaluation value among a plurality of first difference evaluation values ​​for the plurality of candidate instance sequences.

9. The method according to claim 7, wherein: Determining the matching relationship includes: For the output instance and the real instance satisfying the first matching relationship, Arranging multiple real sub-position information of the real instance to obtain a real point sequence; Arrange and combine multiple output sub-position information of the output instance to obtain multiple candidate point sequences; Determining a plurality of second difference evaluation values ​​for the plurality of candidate point sequences according to the plurality of real sub-position information, the plurality of output sub-position information, the order of the real point sequences, and the order of the candidate point sequences; and The second matching relationship is determined according to the minimum second difference evaluation value among the plurality of second difference evaluation values.

10. The method according to claim 7, wherein: The determining of the total loss function value according to the matching relationship, the real position information, the output position information, the real category and the output category comprises: The total loss function value is determined according to a predetermined loss function value; wherein the process of determining the predetermined loss function value includes at least one of the following: Determining a point-level geometric loss function value according to a distance between the output sub-position information and the real sub-position information that satisfies the second matching relationship; Determine the instance-level geometric loss function value according to the difference between the first line segment and the second line segment in the line segment pair; wherein the first line segment has two output sub-position information as endpoints, and the second line segment has two real sub-position information as endpoints; in the same line segment pair, the two endpoints of the first line segment and the two endpoints of the second line segment satisfy the second matching relationship; For the output instance and the real instance satisfying the first matching relationship, determining a category loss function value according to a difference between an output category of the output instance and a real category of the real instance; For each output instance, according to whether there is a real instance that satisfies the first matching relationship with the output instance, Determine the binary classification loss function value; Processing the output position information of the at least one output instance and the real position information of the at least one real instance based on semantic segmentation to obtain a first output mask and a first real mask respectively; determining a semantic segmentation loss function value according to a difference between the first output mask and the first real mask; and Based on instance segmentation, the output position information of the at least one output instance and the real position information of the at least one real instance are processed to obtain a second output mask and a second real mask respectively; and according to the difference between the second output mask and the second real mask, the instance segmentation loss function value is determined.

11. A map updating device, comprising: A first data determination module is used to determine, for the same location information, associated historical map data and M collected image data based on a time series, where M is an integer greater than or equal to 1; A first encoding module, used to encode the historical map data and the M collected image data respectively to obtain historical map features and M collected image features; A target instance determination module, used to determine the location information and category information of a target instance according to the historical map features and the M collected image features, wherein the target instance represents lane information; as well as An updating module is used to update the map data according to the location information and category information of the target instance.

12. The device according to claim 11, wherein The target instance determination module includes: A first determination submodule, configured to determine, based on the M acquired image features, M first confidence maps respectively corresponding to the M acquired image features; each acquired image feature includes a plurality of sub-features, each first confidence map includes a plurality of confidences, and the plurality of confidences in the first confidence map represent weights of the plurality of sub-features in the corresponding acquired image feature; and The second determination submodule is used to determine the location information and category information of the target instance according to the M collected image features, the M first confidence maps and the historical map features.

13. The device according to claim 12, wherein: The second determining submodule includes: A first fusion unit, used for fusing the M collected image features and the M first confidence maps to obtain a first intermediate fusion feature; a first determining unit, configured to determine a first target fusion feature according to the first intermediate fusion feature and the historical map feature; and The first decoding unit is used to decode the first target fusion feature to obtain the location information and category information of the target instance.

14. A training device for a map update model, comprising: A sample determination module is used to determine a training sample, wherein the training sample includes: a map prior, N sample images based on a time series, and a label; the map prior and the N sample images are associated with the same position information, the label represents whether a real instance exists in the N sample images and represents the real position information and the real category of the real instance, the real instance represents lane information, and N is an integer greater than or equal to 1; A second encoding module is used to encode the map prior and the N sample images respectively to obtain map features and N sample image features; A second data determination module, configured to determine output location information and an output category of at least one output instance according to the map feature and the N sample image features; and A training module is used to train the map update model according to the real position information, the output position information, the real category and the output category.

15. The device according to claim 14, wherein: The second data determination module includes: a third determination submodule, configured to determine, based on the N sample image features, N second confidence maps corresponding to the N sample image features respectively; each sample image feature includes a plurality of sub-features, each second confidence map includes a plurality of confidences, and the plurality of confidences in the second confidence map represent weights of the plurality of sub-features in the corresponding sample image feature; and The fourth determination submodule is used to determine the output position information and output category of the at least one output instance according to the N sample image features, the N second confidence maps and the map features.

16. The device according to claim 15, wherein: The fourth determining submodule includes: A second fusion unit, used for fusing the N sample image features and the N second confidence maps to obtain a second intermediate fusion feature; a second determining unit, configured to determine a second target fusion feature according to the second intermediate fusion feature and the map feature; and The second decoding unit is used to decode the second target fusion feature to obtain the output position information and output category of the at least one output instance.

17. The device according to claim 14, wherein: The training module includes: A matching submodule, used to determine a matching relationship, wherein the matching relationship includes at least one of a first matching relationship and a second matching relationship; the first matching relationship represents a matching relationship between a real instance and an output instance; the real location information of the same real instance includes a plurality of real sub-location information, the output location information of the same output instance includes a plurality of output sub-location information, and the second matching relationship represents a matching relationship between the real sub-location information and the output sub-location information; a total loss determination submodule, configured to determine a total loss function value according to the matching relationship, the real position information, the output position information, the real category, and the output category; and The training submodule is used to train the map update model according to the total loss function value.

18. The device according to claim 17, wherein: The matching submodule includes: A first arrangement unit, configured to arrange the at least one real instance to obtain a real instance sequence; A second arrangement unit, configured to arrange and combine the at least one output instance to obtain a plurality of candidate instance sequences; a first difference determination unit, configured to determine, for each candidate output instance sequence, a first difference evaluation value for the candidate instance sequence according to output position information and output category of each output instance in the candidate output instance sequence, real position information and real category of each real instance in the real instance sequence, and an order of the real instance sequence and an order of each candidate instance sequence; and The first relationship determining unit is configured to determine the first matching relationship according to a minimum first difference evaluation value among a plurality of first difference evaluation values ​​for the plurality of candidate instance sequences.

19. The device according to claim 17, wherein: The matching submodule includes: A second difference determination unit is configured to determine, for each output instance and a real instance that satisfy the first matching relationship, Arranging multiple real sub-position information of the real instance to obtain a real point sequence; Arrange and combine multiple output sub-position information of the output instance to obtain multiple candidate point sequences; Determining a plurality of second difference evaluation values ​​for the plurality of candidate point sequences according to the plurality of real sub-position information, the plurality of output sub-position information, the order of the real point sequence, and the order of the candidate point sequence; The second relationship determining unit is configured to determine the second matching relationship according to the minimum second difference evaluation value among the plurality of second difference evaluation values.

20. The device according to claim 17, wherein: The total loss determination submodule includes: A total loss determination unit is used to determine the total loss function value according to a predetermined loss function value; wherein the process of determining the predetermined loss function value includes at least two of the following: Determining a point-level geometric loss function value according to a distance between the output sub-position information and the real sub-position information that satisfies the second matching relationship; Determine the instance-level geometric loss function value according to the difference between the first line segment and the second line segment in the line segment pair; wherein the first line segment has two output sub-position information as endpoints, and the second line segment has two real sub-position information as endpoints; in the same line segment pair, the two endpoints of the first line segment and the two endpoints of the second line segment satisfy the second matching relationship; For the output instance and the real instance satisfying the first matching relationship, determining a category loss function value according to a difference between an output category of the output instance and a real category of the real instance; For each output instance, determining a binary classification loss function value according to whether there is a real instance that satisfies a first matching relationship with the output instance; Processing the output position information of the at least one output instance and the real position information of the at least one real instance based on semantic segmentation to obtain a first output mask and a first real mask respectively; determining a semantic segmentation loss function value according to a difference between the first output mask and the first real mask; and Based on instance segmentation, the output position information of the at least one output instance and the real position information of the at least one real instance are processed to obtain a second output mask and a second real mask respectively; and according to the difference between the second output mask and the second real mask, the instance segmentation loss function value is determined.

21. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

23. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Map generation method and device, aircraft and storage medium

    CN110799983A

  • Detection model training method and device, high-precision map updating method and device, medium and product

    CN113762397A

  • Image processing method and device, map data updating method and device and storage medium

    CN113837155A

  • Map data sample generation method and device, map data sample training method and device, map updating method and device and medium

    CN116401330A

  • Map generation method and device, computer equipment and storage medium

    CN116772821A

Cited By

  • SPR response region identification method based on image semantic segmentation and time sequence alignment

    CN120894544A

  • Map data updating method and device, equipment and storage medium

    CN122045323A

  • Road structure reconstruction method and device based on vehicle-mounted panoramic image, and medium

    CN122312953A

  • Road structure reconstruction method and device based on vehicle-mounted panoramic image and medium

    CN122312953B