Map correction method and device

By combining a training target localization model and a multimodal large language model, the problems of long cycle, high investment and insufficient accuracy in electronic map production and correction are solved, and fast and accurate map correction and semantic enhancement are achieved.

CN121725352APending Publication Date: 2026-03-24XINXING JIHUA (BEIJING) INTELLIGENT EQUIP TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing electronic map production and correction technologies suffer from long cycles, high costs, and slow timeliness, making it difficult to meet the need for rapid real-time understanding of specific areas. Furthermore, they lack accuracy in multi-target recognition in complex environments, often resulting in common-sense errors such as discontinuous roads and overlapping buildings and roads.

Method used

An initial map localization model trained on labeled map image samples of multiple target categories is used for target localization. A target multimodal large language model is then combined for correction prediction. Based on the input of the semantic map and the initial map, error localization and correction suggestions are generated. The semantic map is then corrected based on the correction suggestions to achieve accurate target map recognition and semantic enhancement.

Benefits of technology

It enables accurate identification of ground buildings and facilities and semantic information completion, improves the data quality and compliance of electronic maps, meets the need for rapid real-time correction, and reduces the investment in manual annotation and survey.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725352A_ABST
    Figure CN121725352A_ABST
Patent Text Reader

Abstract

The invention provides a map correction method and device, and the method comprises the steps: inputting an initial map into a trained target positioning model for target positioning, and obtaining a semantic map; the target positioning model is obtained by training a map image sample marked with multiple types of targets; inputting the semantic map and the initial map into a target multi-modal large language model for correction prediction to obtain error location and correction suggestions corresponding to the semantic map; wherein the target multi-modal large language model is obtained by training based on map error samples, and the map error samples comprise the map image samples, semantic map samples corresponding to the map image samples and correction labels corresponding to the semantic map samples; and correcting the error location in the semantic map based on the correction suggestion to obtain a target map corresponding to the initial map. According to the invention, semantic enhancement and compliance test of the electronic map can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing and the technical field of electronic map making and correction, and in particular to a map correction method and device. BACKGROUND

[0002] In the current technical field of electronic map making and correction, the existing technology has the characteristics of long cycle, high investment, and slow timeliness. In addition, on the basis of geographic data, a large amount of manual field annotation investigation and attribute information acquisition are required. In the event of an emergency, it is difficult to meet the demand of quickly and timely grasping the specific area. Some existing related artificial intelligence and electronic map applications are often limited to the recognition and detection of a certain type of target, such as buildings, roads, and bridges. And the accuracy of multi-target recognition in complex environments is still insufficient, and common sense errors such as discontinuous roads and overlapping buildings and roads cannot be avoided. It is impossible to form an electronic map with rich content and accurate details. SUMMARY

[0003] To solve at least one of the above problems, the present application provides a map correction method and device.

[0004] The present application provides a map correction method, comprising: inputting an initial map into a trained target positioning model for target positioning to obtain a semantic map; wherein the target positioning model is trained based on a map image sample labeled with multiple categories of targets; inputting the semantic map and the initial map into a target multi-modal large language model for correction prediction to obtain error positioning and correction suggestions corresponding to the semantic map; wherein the target multi-modal large language model is trained based on a map error sample, and the map error sample includes the map image sample, a semantic map sample corresponding to the map image sample, and a correction label corresponding to the semantic map sample; correcting the error positioning in the semantic map based on the correction suggestions to obtain a target map corresponding to the initial map.

[0005] According to the map correction method provided by the present application, before the semantic map and the initial map are input into the target multi-modal large language model for correction prediction to obtain error positioning and correction suggestions corresponding to the semantic map, the method further comprises: for each map image sample labeled with multiple categories of targets, inputting the map image sample into the target positioning model for target positioning to obtain a semantic map sample corresponding to the map image sample; construct a correction prompt corresponding to the map image sample based on the map image sample, the positioning label corresponding to the map image sample, and the semantic map sample corresponding to the map image sample; the positioning label represents the correct labeling of the multi-category target on the map image sample; input the correction prompt into an initial multi-modal large language model to obtain a correction evaluation corresponding to the semantic map sample; error correction is performed on the correction evaluation to obtain a correction label corresponding to the semantic map sample; determine the map error sample corresponding to the map image sample based on the map image sample, the semantic map sample, and the correction label; train the initial multi-modal large language model based on each map error sample to obtain the target multi-modal large language model.

[0006] According to the map correction method provided by the application, the initial multi-modal large language model is trained based on each map error sample to obtain the target multi-modal large language model, which comprises: for each map error sample, input the map image sample and the semantic map sample in the map error sample into the initial multi-modal large language model for correction prediction to obtain the predicted correction content of the semantic map sample; calculate a first loss value based on the predicted correction content and the correction label in the map error sample; based on the first loss value, fine-tune the initial multi-modal large language model using low-rank adaptation technology; continue to train the fine-tuned initial multi-modal large language model until the training stopping condition is reached to obtain the target multi-modal large language model.

[0007] According to the map correction method provided by the application, before the initial map is input into the trained target positioning model for target positioning to obtain the semantic map, it further comprises: construct a plurality of map image samples labeled with multi-category targets; divide a plurality of map image samples into a training set and a validation set; based on the training set, train the initial positioning model for multiple rounds to obtain a candidate positioning model corresponding to each round of training, and based on the validation set, validate each candidate positioning model to obtain a validation result of each candidate positioning model; based on each validation result, determine the target positioning model from each candidate positioning model.

[0008] According to the map correction method provided by the application, the plurality of map image samples labeled with multiple categories of targets are constructed, comprising: Obtaining a plurality of map images; For each of the map images, selecting an unlabeled target from the map image, and labeling the target according to the category of the target; If the labeling of the map image is not completed, continue to perform the steps of selecting an unlabeled target from the map image and labeling the target according to the category of the target; If the labeling of the map image is completed, the labeled map image is determined as the map image sample labeled with multiple categories of targets.

[0009] According to the map correction method provided by the application, the labeling of the target according to the category of the target comprises: If the category of the target is a first category, the target is labeled with a box; the first category represents that the edge box of the target is a regular shape; If the category of the target is a second category, the target is labeled with an image mask; the second category represents that the edge box of the target is an irregular shape.

[0010] According to the map correction method provided by the application, the plurality of map image samples labeled with multiple categories of targets are constructed, comprising: Based on the size of the training set and the data processing performance of the initial positioning model, the number of batches is determined; The initial positioning model is used as a specified positioning model; Based on the training set and the number of batches, the specified positioning model is trained to obtain a candidate positioning model corresponding to the current round of training; If the training is not completed for a set number of rounds, the candidate positioning model corresponding to the current round of training is used as a specified positioning model, and the step of training the specified positioning model based on the training set and the number of batches is continued.

[0011] According to the map correction method provided by the application, the plurality of map image samples labeled with multiple categories of targets are constructed, comprising: According to the number of batches, the training set is divided into at least one sub-training set; For any of the sub-training sets, each of the map image samples in the training set is input into the specified positioning model for target positioning to obtain a predicted positioning of each of the map image samples. calculate a second loss value based on the predicted positions of each of the map image samples and the position labels of each of the map image samples; adjust model parameters of the specified positioning model based on the second loss value; if the current round of training does not traverse all of the sub-training sets, for any of the sub-training sets that the current round of training does not traverse, continue to perform the steps of inputting each of the map image samples in the training set into the specified positioning model for target positioning and thereafter; if the current round of training traverses all of the sub-training sets, determine the specified positioning model that has completed training as the candidate positioning model corresponding to the current round of training.

[0012] According to the map correction method provided by the application, the target positioning of each of the map image samples in the training set into the specified positioning model obtains the predicted position of each of the map image samples, which comprises: for any of the map image samples in the training set, for the target belonging to the first category in the map image sample, the target positioning of the map image sample is performed by the specified positioning model using the target detection task, and for the target belonging to the second category in the map image sample, the target positioning of the map image sample is performed by the specified positioning model using the image segmentation task, thereby obtaining the predicted position of the map image sample.

[0013] The application further provides a map correction device, which comprises: a target positioning module configured to input an initial map into a trained target positioning model for target positioning to obtain a semantic map; wherein the target positioning model is trained based on a map image sample labeled with multiple categories of targets; a correction prediction module configured to input the semantic map and the initial map into a target multi-modal large language model for correction prediction to obtain error positioning and correction suggestions corresponding to the semantic map; wherein the target multi-modal large language model is trained based on a map error sample, and the map error sample comprises the map image sample, a semantic map sample corresponding to the map image sample, and a correction label corresponding to the semantic map sample; a correction module configured to correct the error positioning in the semantic map based on the correction suggestions to obtain a target map corresponding to the initial map.

[0014] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the map correction method according to any of the above when executing the computer program.

[0015] The application also provides a non-transitory computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the map correction method according to any one of the above.

[0016] The application also provides a computer program product comprising a computer program which, when executed by a processor, implements the map correction method according to any one of the above.

[0017] First, some related content involved in the application is briefly described.

[0018] Map image is a concept combining the characteristics of maps and images, usually refers to the visualization image of the earth's surface or a specific area obtained by remote sensing technology (such as satellite, aerial photography, etc.), and has the positioning, scale and other attributes of the map after processing. It not only retains the real visual information of the image, but also integrates the spatial reference and geographical feature identification function of the map.

[0019] Deep learning is a method of machine learning that simulates the multi-layer artificial neural network structure of the human brain, allowing computers to automatically learn features and rules from data and achieve complex task processing. Its "depth" is reflected in the network containing multiple hidden layers, which can extract abstract features of data layer by layer (for example, from image pixel values, gradually learn edges, textures, components, and finally recognize complete targets), without manual design of features. It is widely used in image recognition, speech processing, natural language understanding and other fields.

[0020] CV (Computer vision, computer vision) target detection is one of the core tasks in the field of electronic map making and correction technology, aiming to automatically identify and locate specific categories of targets (such as people, cars, animals, traffic signs, etc.) from images or videos through algorithms. Its core output usually includes: the category of the target (such as "car") and the location of the target in the image (commonly represented by bounding box coordinates, i.e. the coordinates of the upper left corner and the lower right corner of the rectangular box).

[0021] Instance segmentation is an advanced task in computer vision that further distinguishes different individuals of the same category and outlines the complete contour of each individual (rather than just positioning with a bounding box) based on target detection. Simply put, it not only answers "what is in the picture (category)" and "where is it (location)", but also clearly defines "the specific shape and range of each target". For example, in a picture containing multiple cats, instance segmentation will label the category of each cat ("cat") and mark the hair, limbs, etc. of each cat with a pixel-level mask (mask), clearly distinguishing different individuals.

[0022] MLLM (Multimodal Large Language Model) is an important development in the field of artificial intelligence. It is based on a large language model and integrates the ability to process multiple non-text modal information. Its main capabilities include cross-modal unified representation, establishing connections between different modalities such as text, vision, and hearing, and implementing tasks such as text-to-image and visual question answering.

[0023] LoRA (Low-Rank Adaptation) technology, also known as LoRA large model fine-tuning technology, is a technique used to reduce the number of parameters and computational resources in large model fine-tuning. Lora technology can freeze the weights of a large model and insert a small, low-rank adaptation layer in the key layers of the large model. The dimension of the small, low-rank adaptation layer is much smaller than that of the original layer of the large model. The training parameter quantity is much smaller than that of the large model. During inference, the adaptation layer and the large model layer results are superimposed without additional inference cost. By adjusting only about 0.1% of the parameters, similar performance to full fine-tuning can be achieved.

[0024] The map correction method and device provided by the application corrects the semantic map based on the target positioning model and the target multimodal large language model. The target positioning model is trained based on map image samples labeled with multiple categories of targets. The semantic map and the initial map are input into the target multimodal large language model for correction prediction to obtain error positioning and correction suggestions corresponding to the semantic map. The target multimodal large language model is trained based on map error samples, which include the map image samples, the semantic map samples corresponding to the map image samples, and the correction labels corresponding to the semantic map samples. The error positioning in the semantic map is corrected based on the correction suggestions to obtain the target map corresponding to the initial map. The target positioning model realizes accurate identification and completion of semantic information of ground building facilities, and the multimodal large language model automatically detects and corrects the identification results, thereby achieving semantic enhancement and compliance testing of electronic map data. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0026] Figure 1 is one of the flowcharts of the map correction method provided by the application.

[0027] Figure 2 This is the second flowchart of the map correction method provided by the present invention.

[0028] Figure 3 This is a schematic diagram of the process for constructing map image samples provided by the present invention.

[0029] Figure 4 This is a flowchart illustrating the training target localization model provided by the present invention.

[0030] Figure 5 This is a schematic diagram of the process for constructing map error samples provided by the present invention.

[0031] Figure 6 This is a flowchart illustrating the process of fine-tuning the target multimodal large language model provided by the present invention.

[0032] Figure 7 This is a schematic diagram of the map correction device provided by the present invention.

[0033] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0035] The following is combined Figures 1-8 The map correction method and apparatus of the present invention are described.

[0036] Figure 1 This is one of the flowcharts illustrating the map correction method provided by the present invention, such as... Figure 1 As shown, the method includes: Step 101: Input the initial map into the trained target localization model to perform target localization and obtain a semantic map; wherein, the target localization model is trained based on map image samples labeled with multiple categories of targets; Step 102: Input the semantic map and the initial map into the target multimodal large language model for correction prediction to obtain the error localization and correction suggestions corresponding to the semantic map; wherein, the target multimodal large language model is trained based on map error samples, and the map error samples include the map image samples, the semantic map samples corresponding to the map image samples, and the correction labels corresponding to the semantic map samples; Step 103: Correct the erroneous localization in the semantic map based on the correction suggestions to obtain the target map corresponding to the initial map.

[0037] Specifically, the target localization model can be a CV target localization model. The target multimodal large language model can be a fine-tuned multimodal large language model. The initial map can be a map image.

[0038] Specifically, multi-category targets refer to objects of different categories, such as buildings, lakes, roads, etc.

[0039] Specifically, a semantic map refers to the initial map after target labeling. Modification suggestions refer to methods and / or strategies for correcting erroneous localizations, including the location of the error, the cause of the error, and the correction method.

[0040] In practical applications, the CV target localization model can be cascaded with a fine-tuned multimodal large language model: the original map image to be corrected (initial map) is first processed by the CV target localization model to obtain a predicted semantic map; then the semantic map and the original map image are used as inputs to the target multimodal large language model, which obtains the incorrect localization and modification suggestions for the semantic map.

[0041] It should be noted that the target localization model can be obtained by training the initial localization model on map image samples with multiple labeled multi-category targets. Similarly, the target multimodal large language model can be obtained by training the initial multimodal large language model on multiple map error samples containing map image samples, semantic map samples, and correction labels, where the correction labels include erroneous localization labels and modification suggestion labels.

[0042] For example, see Figure 2 , Figure 2 This is the second flowchart of the map correction method provided by the present invention: The CV target localization model (target localization model) and the multimodal large language model are deployed together. The map image first passes through the CV target localization model to obtain a rough electronic map label (semantic map); then the map image, electronic map label and prompt are simultaneously provided to the target multimodal large language model (i.e., the multimodal large language model). The target multimodal large language model provides error localization and correction suggestions (modification suggestions) for the electronic map label, and modifies the electronic map label based on the error localization and correction suggestions to obtain the modified electronic map.

[0043] The map correction method provided by this invention involves inputting an initial map into a trained target localization model for target localization to obtain a semantic map. The target localization model is trained based on map image samples labeled with multiple target categories. The semantic map and the initial map are then input into a target multimodal large language model for correction prediction, yielding erroneous localizations and correction suggestions corresponding to the semantic map. The target multimodal large language model is trained based on map error samples, which include the map image samples, corresponding semantic map samples, and correction labels. Based on the correction suggestions, the erroneous localizations in the semantic map are corrected to obtain the target map corresponding to the initial map. This method achieves accurate identification and semantic information completion of ground buildings and facilities through the target localization model, while automatically detecting and correcting the identification results using a multimodal large language model, thereby realizing semantic enhancement and compliance verification of electronic map data.

[0044] Optionally, before inputting the initial map into the trained target localization model for target localization to obtain the semantic map, the method further includes: Construct multiple map image samples labeled with multiple categories of targets; The multiple map image samples are divided into a training set and a validation set; The initial localization model is trained multiple times based on the training set to obtain candidate localization models corresponding to each training round. Each candidate localization model is then verified based on the verification set to obtain the verification results of each candidate localization model. Based on the verification results, the target localization model is determined from the candidate localization models.

[0045] Specifically, the initial localization model refers to a deep learning CV model that employs a multi-task approach of object detection and instance segmentation, such as YOLO-v8 or YOLOP.

[0046] In practical applications, a map image dataset can be constructed first, which contains multiple map image samples, each labeled with multiple categories of objects. To facilitate model training, the map image dataset can be divided into a training set and a validation set.

[0047] Then, the target detection and localization capabilities of the initial localization model are trained using the training dataset. To ensure the robustness of the target localization model, the initial localization model can be trained in multiple rounds.

[0048] Furthermore, the accuracy of target detection and localization (such as image detection and segmentation) of the candidate localization models corresponding to each round of training is evaluated using a validation set. The set of candidate localization models that performs best on the validation set (best validation results) is selected as the target localization model. For example, by observing the evaluation metrics on the validation set, the best-performing model is selected to prevent overfitting. The evaluation metrics may include at least one of accuracy, recall, and F1 score.

[0049] Optionally, the construction of multiple map image samples labeled with multiple categories of targets includes: Acquire multiple map images; For each map image, select unlabeled targets from the map image and label the targets according to their categories; If the map image is not fully labeled, continue with the steps of selecting unlabeled targets from the map image and labeling the targets according to their categories; Once the map image annotation is completed, the annotated map image is identified as a map image sample annotated with multiple categories of targets.

[0050] In practical applications, multiple map images can be obtained from different scenarios, such as urban scenes, rural scenes, school scenes, and farmland scenes, by collecting satellite imagery and / or aerial images acquired by drones. Furthermore, to ensure the robustness of training, the number of acquired map images should be greater than 100.

[0051] Then, for each map image, annotation tools such as Labelme or Labelbox are used to annotate the key targets that need to be focused on when constructing the electronic map, thus obtaining map image samples with annotations for multiple categories of targets. Different annotation methods can be used for different categories of targets. By traversing each map image and obtaining multiple map image samples, a map image dataset is obtained.

[0052] Alternatively, publicly available map image datasets can be used, but it is necessary to ensure that the annotation formats and methods of each map image dataset are consistent.

[0053] Optionally, labeling the target according to its category includes: When the target belongs to the first category, the target is labeled with a box; the first category indicates that the bounding box of the target has a regular shape. When the target is classified as the second category, the target is labeled using an image mask; the second category indicates that the bounding box of the target is an irregular shape.

[0054] Specifically, the two main categories of targets that need to be focused on when constructing the electronic map are then labeled. One category consists of relatively regular buildings, which are labeled with boxes. The other category consists of irregular roads, open spaces, and water bodies, which are labeled with image masks.

[0055] For example, see Figure 3 , Figure 3 This is a flowchart illustrating the process of constructing map image samples provided by the present invention: For each map image, it is determined whether the labeling is complete. If the labeling is complete, the process ends and a map image sample is obtained. If the labeling is not complete, an unlabeled target is selected from the map image (target selection). It is determined whether the target belongs to the first category. If it belongs to the first category, a box label is made. If it does not belong to the first category, an instance polygon label is used (image mask labeling). After labeling, it is determined whether the labeling is complete again until the process ends.

[0056] In this embodiment of the invention, different annotation methods are used for different categories of targets, which can improve the accuracy of annotation and thus improve the efficiency of model training.

[0057] Optionally, the step of training the initial localization model multiple times based on the training set to obtain a candidate localization model corresponding to each training round includes: The batch size is determined based on the size of the training set and the data processing performance of the initial localization model. Use the initial positioning model as the specified positioning model; Based on the training set and the batch size, the specified localization model is trained to obtain the candidate localization model corresponding to the current training round. If the set number of training rounds has not been completed, the candidate localization model corresponding to the current training round is taken as the designated localization model, and the step of training the designated localization model based on the training set and the batch number continues.

[0058] Specifically, data processing performance refers to the hardware conditions of the initial localization model, such as the performance of the GPU (Graphics Processing Unit).

[0059] Specifically, the batch size refers to the number of training batches used in model training. The appropriate training process parameter, i.e., the batch size, can be selected based on the size of the training set and hardware conditions.

[0060] Specifically, training the training set for 10 to 20 rounds usually yields good training results; therefore, the number of rounds can be set to 10-20.

[0061] In practical applications, the number of batches is first determined based on the size of the training set and the data processing performance of the initial localization model. Then, based on the training set and the number of batches, the initial localization model is trained in batches to obtain the candidate localization model obtained in the first round of training. Further, based on the training set and the number of batches, the candidate localization model obtained in the first round of training is trained in batches to obtain the candidate localization model obtained in the second round of training. This process continues iteratively, based on the candidate localization model obtained in the previous round of training, until the set number of training rounds is reached, thus completing the multi-round training.

[0062] Furthermore, during the training of the target localization model, the optimization method used can be either SGD (Stochastic Gradient Descent) or Adam (Adaptive Moment Estimation) gradient descent. The initial learning rate for training the target localization model can be set to less than 1e-3.

[0063] Optionally, training the specified localization model based on the training set and the batch size to obtain the candidate localization model corresponding to the current training round includes: The training set is divided into at least one sub-training set according to the batch quantity; For any of the aforementioned sub-training sets, each of the aforementioned map image samples in the training set is input into the specified positioning model for target positioning, thereby obtaining the predicted positioning of each of the aforementioned map image samples; Based on the predicted location of each map image sample and the location label of each map image sample, a second loss value is calculated; The model parameters of the specified localization model are adjusted based on the second loss value; If the current training round has not traversed all the aforementioned sub-training sets, for any of the aforementioned sub-training sets that have not been traversed in the current training round, the steps of inputting each of the aforementioned map image samples in the training set into the specified positioning model for target positioning and thereafter shall continue to be executed. If all the sub-training sets have been traversed in the current training round, the specified localization model that has been trained is determined as the candidate localization model corresponding to the current training round.

[0064] Specifically, the second loss value can be based on a loss value determined by any loss function. The loss function can be a loss function based on distance metrics, such as mean squared error loss function, L2 loss function, L1 loss function, etc., or a loss function based on probability distribution metrics, such as KL divergence function, cross-entropy loss, softmax loss function, etc.

[0065] Specifically, location tags represent the correct labeling of multiple categories of targets on map image samples.

[0066] For example, the batch training process can be as follows: The training set is divided into batches of a specified number of sub-training sets. The first training set is then input into a designated localization model for target localization, obtaining the predicted localization of each map image sample in the first training set. A second loss value is calculated based on the predicted localization and localization labels of each map image sample in the first training set. Then, based on the second loss value, the model parameters of the designated localization model are optimized using an optimizer. Next, the second training set is input into the first batch of adjusted designated localization models for further optimization, and so on, until the last training set is input into the previous batch of adjusted designated localization models for optimization, resulting in the candidate localization model for the current round. This improves training efficiency.

[0067] See Figure 4 , Figure 4 This is a flowchart illustrating the training target localization model provided by the present invention: First, it is determined whether the set number of training rounds has been completed; if yes, the training ends; if no, it is necessary to perform batch localization model prediction, calculate loss values ​​with labels, and optimize model parameters using an optimizer to complete the current round of training. Then, the parameters are frozen and validated using a validation set. Next, it is determined whether the validation result is better than the previous rounds. If it is better than the previous rounds, the model parameters are saved, and the process of determining whether the set number of training rounds has been completed continues. If it is not better than the previous rounds, the process of determining whether the set number of training rounds has been completed continues.

[0068] Optionally, the step of inputting each of the map image samples in the training set into the designated positioning model for target positioning to obtain the predicted positioning of each of the map image samples includes: For any map image sample in the training set, for targets belonging to the first category in the map image sample, target localization is performed using the specified localization model with a target detection task; for targets belonging to the second category in the map image sample, target localization is performed using the specified localization model with an image segmentation task, thereby obtaining the predicted localization of the map image sample.

[0069] Specifically, the localization model employs a multi-task model of object detection and instance segmentation to uniformly detect and localize two types of targets in the map image dataset. For regular building targets (belonging to the first category), the object detection task is used to obtain predicted localization; for irregular targets such as roads (belonging to the second category), the image segmentation task is used to maximize the intersection-union ratio between the predicted mask and the labeled mask to obtain predicted localization.

[0070] In this embodiment of the invention, different tasks are used for target localization for different types of targets, which can improve the accuracy and reliability of prediction.

[0071] Optionally, before inputting the semantic map and the initial map into the target multimodal large language model for correction prediction to obtain the error localization and correction suggestions corresponding to the semantic map, the method further includes: For each map image sample labeled with multiple target categories, the map image sample is input into the target localization model for target localization, thereby obtaining the semantic map sample corresponding to the map image sample; Based on the map image sample, the location tags corresponding to the map image sample, and the semantic map sample corresponding to the map image sample, correction prompt words corresponding to the map image sample are constructed; the location tags represent the correct labeling of the multi-category targets on the map image sample; The correction prompts are input into the initial multimodal large language model to obtain the correction evaluation corresponding to the semantic map sample; Error correction is performed on the correction evaluation to obtain the correction label corresponding to the semantic map sample; Based on the map image sample, the semantic map sample, and the correction label, determine the map error sample corresponding to the map image sample; Based on the map error samples, the initial multimodal large language model is trained to obtain the target multimodal large language model.

[0072] Specifically, the corrective evaluation includes the mislocation of the predicted samples and corrective recommendations.

[0073] Specifically, the initial multimodal large language model adopted open-source models such as LLaVA-v1.5-13B or Qwen2-VL-7B.

[0074] In practical applications, the trained target localization model is used to predict the predicted label for each map image sample, i.e., the semantic map sample. These predicted labels, map image samples, and their accompanying localization labels (real labels) form a triple of map image—real label—predicted label. Each data point in this triple is then combined with a prompt template from a specific building code text to obtain a corrected prompt.

[0075] The quality of the prompt template affects the quality of the output of the multimodal large language model. For example, the prompt template can be as follows: "You are an AI (Artificial Intelligence) visual assistant analyzing whether the annotations on some map images are unreasonable. You will get 3 images: the first image is the real map image, the second image is the real annotation of this image, where different colors represent different objects in the map (the correspondence needs to be explained in detail), and the third image is the incorrect annotation of the map image, where different colors represent different objects in the map (the correspondence needs to be explained in detail).

[0076] Your task is to use the first real image, the second real-labeled image, and the third incorrectly labeled image: Describe where the annotations for the third image are incorrect.

[0077] Describe the content and visual details of the error area.

[0078] Describe how to correct the error area to make it correct; Do not mention images 2 and 3 in your answer; always assume you are observing image 1. When dealing with problem 1, use natural language to describe the absolute location of the error, and do not give vague or uncertain locations: Absolute location: The position of the error label relative to the entire image, such as left, right, top, bottom, etc. (can be expanded as needed); When addressing problems 2 and 3, describe the objects in the error area using natural language, considering the following environmental details: Continuity: Observe whether the road is continuous; usually, the road will not be interrupted at will.

[0079] ...(Other building codes that need to be considered can be added)."

[0080] Furthermore, the pre-trained multimodal large language model (initial multimodal large language model) outputs errors in the predicted labels and suggestions for correction, i.e., correction evaluation. Then, the correction evaluation can be corrected, such as through manual review, to obtain corrected labels. Thus, a set of map error ternary data, i.e., map error samples, can be constructed, containing map imagery, predicted labels, and correct / incorrect reasoning (corrected labels).

[0081] By iterating through each map image sample, we can obtain the map error sample corresponding to each map image sample. Then, we can train the initial multimodal large language model based on each map error sample to obtain the target multimodal large language model.

[0082] For example, see Figure 5 , Figure 5This is a schematic diagram of the process for constructing map error samples provided by the present invention: First, a trained CV model (target localization model) is used to generate predicted labels for each map image. Then, the predicted labels are combined with the original map image dataset (each map image sample labeled with multiple categories of targets) to form a triplet, which is map image—real label—predicted label. This triplet is combined with a prompt word template set according to building codes to form a correction prompt word, which is input into the initial multimodal large language model. The initial multimodal large language model generates map error correction (correction evaluation). Then, manual verification is performed to obtain the combined map error dataset as the map error dataset (map error sample corresponding to each map image sample) forming the map image—predicted label—error correction dataset.

[0083] In this embodiment of the invention, data processing is performed based on the map image samples used to train the target localization model, the trained target localization model, and the initial multimodal large language model to construct map error samples. This not only reduces the amount of data acquisition and further verifies the localization capability of the target localization model, but also makes the map error samples more consistent with the processing format of the multimodal large language model, which is beneficial to improving the robustness of training the multimodal large language model.

[0084] Optionally, training the initial multimodal large language model based on each of the map error samples to obtain the target multimodal large language model includes: For each map error sample, the map image sample and the semantic map sample in the map error sample are input into the initial multimodal large language model for correction and prediction, so as to obtain the predicted correction content of the semantic map sample; Calculate the first loss value based on the predicted correction content and the correction label in the map error sample; Based on the first loss value, the initial multimodal large language model is fine-tuned using low-rank adaptation techniques; Continue training the fine-tuned initial multimodal large language model until the training stopping condition is met, and obtain the target multimodal large language model.

[0085] Specifically, the first loss value can be based on a loss value determined by any loss function. The loss function can be a loss function based on distance metrics, such as mean squared error loss function, L2 loss function, L1 loss function, etc., or a loss function based on probability distribution metrics, such as KL divergence function, cross-entropy loss, softmax loss function, etc.

[0086] Specifically, the prediction correction can include prediction error location and prediction modification suggestions.

[0087] In practical applications, LoRA technology can be used to fine-tune some parameters of a multimodal large model based on a ternary dataset of map errors (map error samples corresponding to each map image sample). This improves the model's ability to identify and correct label prediction errors without relying on real labels, but only based on map images and CV models.

[0088] Specifically, map image samples and predicted labels (semantic map samples) from map error samples can be used as inputs to the initial multimodal language model. LoRA technology is then used to fine-tune the initial multimodal language model until the training stopping condition is met, resulting in the target multimodal language model. The input prompts for the initial multimodal language model are relatively simple, such as: "Please analyze the errors in the incorrectly labeled image relative to the real map image and provide a method for correction." The cross-entropy function can be used as the loss function between the output error prediction (predicted correction content) and the corrected label of the initial multimodal language model to calculate the first loss value.

[0089] Furthermore, during fine-tuning, an iterative training approach can be adopted, with 10-20 iterations and a learning rate less than 1e-3. The average loss on the validation dataset is then observed to determine the optimal fine-tuning parameters. Similarly, each map error sample can be divided into a training set and a validation set. The target multimodal large language model is trained based on the training set and validated based on the validation set.

[0090] For example, see Figure 6 , Figure 6 This is a flowchart illustrating the process of fine-tuning the target multimodal large language model provided by the present invention: First, it is determined whether the set number of training rounds has been completed; if yes, the training ends; if no, it is necessary to extract map image samples and semantic map samples in batches, combine the prompt template, generate predicted and corrected content using the multimodal large language model, calculate the loss and optimize the model parameters to complete the current round of training, then freeze the parameters, validate them using the validation set, and then determine whether the validation result is better than the previous rounds. If it is better than the previous rounds, the model parameters are saved, and the process of determining whether the set number of training rounds has been completed continues; if it is not better than the previous rounds, the process of determining whether the set number of training rounds has been completed continues.

[0091] The overall process of this invention includes the collection and creation of map image datasets, the training of a CV target localization model, the creation of a map error dataset, and finally the fine-tuning and inference application of a multimodal large model. The two models cooperate with each other: the CV target localization model generates coarse labels for the electronic map, constituting the map error dataset for fine-tuning the multimodal large model; the evaluation and modification suggestions of the labels generated by the multimodal large model are used to identify and correct erroneous labels on the electronic map.

[0092] This invention utilizes classic deep learning CV object detection + instance segmentation technology, multimodal large language and large model technology, and LoRA large model fine-tuning technology to achieve accurate identification of ground buildings and facilities and complete semantic information. At the same time, it combines building code knowledge base to automatically detect and correct the identification results, thereby realizing semantic enhancement and compliance verification of electronic map data.

[0093] The map correction device provided by the present invention will be described below. The map correction device described below and the map correction method described above can be referred to in correspondence.

[0094] Figure 7 This is a schematic diagram of the map correction device provided by the present invention, as shown below. Figure 7 As shown, the device includes: The target localization module 701 is configured to input an initial map into a trained target localization model to perform target localization and obtain a semantic map; wherein, the target localization model is trained based on map image samples labeled with multiple categories of targets; The correction prediction module 702 is configured to input the semantic map and the initial map into the target multimodal large language model for correction prediction, and obtain the error localization and correction suggestions corresponding to the semantic map; wherein, the target multimodal large language model is trained based on map error samples, and the map error samples include the map image samples, the semantic map samples corresponding to the map image samples, and the correction labels corresponding to the semantic map samples. The correction module 703 is configured to correct the erroneous location in the semantic map based on the correction suggestion, and obtain the target map corresponding to the initial map.

[0095] The map correction device provided by this invention obtains a semantic map by inputting an initial map into a trained target localization model for target localization. The target localization model is trained based on map image samples labeled with multiple target categories. The semantic map and the initial map are then input into a target multimodal large language model for correction prediction, yielding erroneous localizations and correction suggestions corresponding to the semantic map. The target multimodal large language model is trained based on map error samples, which include the map image samples, corresponding semantic map samples, and correction labels. Based on the correction suggestions, the erroneous localizations in the semantic map are corrected to obtain the target map corresponding to the initial map. The target localization model enables accurate identification and semantic information completion of ground buildings and facilities, while the multimodal large language model facilitates automatic detection and correction of the identification results, thereby achieving semantic enhancement and compliance verification of electronic map data.

[0096] Optionally, the device further includes a first training module configured to: For each map image sample labeled with multiple target categories, the map image sample is input into the target localization model for target localization, thereby obtaining the semantic map sample corresponding to the map image sample; Based on the map image sample, the location tags corresponding to the map image sample, and the semantic map sample corresponding to the map image sample, correction prompt words corresponding to the map image sample are constructed; the location tags represent the correct labeling of the multi-category targets on the map image sample; The correction prompts are input into the initial multimodal large language model to obtain the correction evaluation corresponding to the semantic map sample; Error correction is performed on the correction evaluation to obtain the correction label corresponding to the semantic map sample; Based on the map image sample, the semantic map sample, and the correction label, determine the map error sample corresponding to the map image sample; Based on the map error samples, the initial multimodal large language model is trained to obtain the target multimodal large language model.

[0097] Optionally, the first training module is specifically configured as follows: For each map error sample, the map image sample and the semantic map sample in the map error sample are input into the initial multimodal large language model for correction and prediction, so as to obtain the predicted correction content of the semantic map sample; Calculate the first loss value based on the predicted correction content and the correction label in the map error sample; Based on the first loss value, the initial multimodal large language model is fine-tuned using low-rank adaptation techniques; Continue training the fine-tuned initial multimodal large language model until the training stopping condition is met, and obtain the target multimodal large language model.

[0098] Optionally, the device further includes a second training module configured to: Construct multiple map image samples labeled with multiple categories of targets; The multiple map image samples are divided into a training set and a validation set; The initial localization model is trained multiple times based on the training set to obtain candidate localization models corresponding to each training round. Each candidate localization model is then verified based on the verification set to obtain the verification results of each candidate localization model. Based on the verification results, the target localization model is determined from the candidate localization models.

[0099] Optionally, the second training module is specifically configured as follows: Acquire multiple map images; For each map image, select unlabeled targets from the map image and label the targets according to their categories; If the map image is not fully labeled, continue with the steps of selecting unlabeled targets from the map image and labeling the targets according to their categories; Once the map image annotation is completed, the annotated map image is identified as a map image sample annotated with multiple categories of targets.

[0100] Optionally, the second training module is specifically configured as follows: When the target belongs to the first category, the target is labeled with a box; the first category indicates that the bounding box of the target has a regular shape. When the target is classified as the second category, the target is labeled using an image mask; the second category indicates that the bounding box of the target is an irregular shape.

[0101] Optionally, the second training module is specifically configured as follows: The batch size is determined based on the size of the training set and the data processing performance of the initial localization model. Use the initial positioning model as the specified positioning model; Based on the training set and the batch size, the specified localization model is trained to obtain the candidate localization model corresponding to the current training round. If the set number of training rounds has not been completed, the candidate localization model corresponding to the current training round is taken as the designated localization model, and the step of training the designated localization model based on the training set and the batch number continues.

[0102] Optionally, the second training module is specifically configured as follows: The training set is divided into at least one sub-training set according to the batch quantity; For any of the aforementioned sub-training sets, each of the aforementioned map image samples in the training set is input into the specified positioning model for target positioning, thereby obtaining the predicted positioning of each of the aforementioned map image samples; Based on the predicted location of each map image sample and the location label of each map image sample, a second loss value is calculated; The model parameters of the specified localization model are adjusted based on the second loss value; If the current training round has not traversed all the aforementioned sub-training sets, for any of the aforementioned sub-training sets that have not been traversed in the current training round, the steps of inputting each of the aforementioned map image samples in the training set into the specified positioning model for target positioning and thereafter shall continue to be executed. If all the sub-training sets have been traversed in the current training round, the specified localization model that has been trained is determined as the candidate localization model corresponding to the current training round.

[0103] Optionally, the second training module is specifically configured as follows: For any map image sample in the training set, for targets belonging to the first category in the map image sample, target localization is performed using the specified localization model with a target detection task; for targets belonging to the second category in the map image sample, target localization is performed using the specified localization model with an image segmentation task, thereby obtaining the predicted localization of the map image sample.

[0104] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a map correction method. This method includes: inputting an initial map into a trained target localization model for target localization to obtain a semantic map; wherein the target localization model is trained based on map image samples labeled with multiple target categories; inputting the semantic map and the initial map into a target multimodal large language model for correction prediction to obtain error localization and correction suggestions corresponding to the semantic map; wherein the target multimodal large language model is trained based on map error samples, the map error samples including the map image samples, semantic map samples corresponding to the map image samples, and correction labels corresponding to the semantic map samples; and correcting the error localization in the semantic map based on the correction suggestions to obtain a target map corresponding to the initial map.

[0105] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the map correction method provided by the above methods. The method includes: inputting an initial map into a trained target localization model for target localization to obtain a semantic map; wherein the target localization model is trained based on map image samples labeled with multiple categories of targets; inputting the semantic map and the initial map into a target multimodal large language model for correction prediction to obtain error localization and correction suggestions corresponding to the semantic map; wherein the target multimodal large language model is trained based on map error samples, the map error samples including the map image samples, semantic map samples corresponding to the map image samples, and correction labels corresponding to the semantic map samples; and correcting the error localization in the semantic map based on the correction suggestions to obtain a target map corresponding to the initial map.

[0107] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the map correction method provided by the above methods. The method includes: inputting an initial map into a trained target localization model for target localization to obtain a semantic map; wherein the target localization model is trained based on map image samples labeled with multiple categories of targets; inputting the semantic map and the initial map into a target multimodal large language model for correction prediction to obtain erroneous localization and correction suggestions corresponding to the semantic map; wherein the target multimodal large language model is trained based on map error samples, the map error samples including the map image samples, semantic map samples corresponding to the map image samples, and correction labels corresponding to the semantic map samples; and correcting the erroneous localization in the semantic map based on the correction suggestions to obtain a target map corresponding to the initial map.

[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A map correction method, characterized in that, include: The initial map is input into a trained target localization model for target localization to obtain a semantic map; wherein, the target localization model is trained based on map image samples labeled with multiple categories of targets; The semantic map and the initial map are input into the target multimodal large language model for correction prediction to obtain the error localization and correction suggestions corresponding to the semantic map; wherein, the target multimodal large language model is trained based on map error samples, and the map error samples include the map image samples, the semantic map samples corresponding to the map image samples, and the correction labels corresponding to the semantic map samples. Based on the proposed corrections, the erroneous localizations in the semantic map are corrected to obtain the target map corresponding to the initial map.

2. The map correction method according to claim 1, characterized in that, Before inputting the semantic map and the initial map into the target multimodal large language model for correction prediction to obtain the error localization and correction suggestions corresponding to the semantic map, the method further includes: For each map image sample labeled with multiple target categories, the map image sample is input into the target localization model for target localization, thereby obtaining the semantic map sample corresponding to the map image sample; Based on the map image sample, the location tags corresponding to the map image sample, and the semantic map sample corresponding to the map image sample, correction prompt words corresponding to the map image sample are constructed; the location tags represent the correct labeling of the multi-category targets on the map image sample; The correction prompts are input into the initial multimodal large language model to obtain the correction evaluation corresponding to the semantic map sample; Error correction is performed on the correction evaluation to obtain the correction label corresponding to the semantic map sample; Based on the map image sample, the semantic map sample, and the correction label, determine the map error sample corresponding to the map image sample; Based on the map error samples, the initial multimodal large language model is trained to obtain the target multimodal large language model.

3. The map correction method according to claim 2, characterized in that, The step of training the initial multimodal large language model based on each of the map error samples to obtain the target multimodal large language model includes: For each map error sample, the map image sample and the semantic map sample in the map error sample are input into the initial multimodal large language model for correction and prediction, so as to obtain the predicted correction content of the semantic map sample; Calculate the first loss value based on the predicted correction content and the correction label in the map error sample; Based on the first loss value, the initial multimodal large language model is fine-tuned using low-rank adaptation techniques; Continue training the fine-tuned initial multimodal large language model until the training stopping condition is met, and obtain the target multimodal large language model.

4. The map correction method according to claim 1, characterized in that, Before inputting the initial map into the trained target localization model for target localization to obtain the semantic map, the process also includes: Construct multiple map image samples labeled with multiple categories of targets; The multiple map image samples are divided into a training set and a validation set; The initial localization model is trained multiple times based on the training set to obtain candidate localization models corresponding to each training round. Each candidate localization model is then verified based on the verification set to obtain the verification results of each candidate localization model. Based on the verification results, the target localization model is determined from the candidate localization models.

5. The map correction method according to claim 4, characterized in that, The construction of multiple map image samples labeled with multiple categories of targets includes: Acquire multiple map images; For each map image, select unlabeled targets from the map image and label the targets according to their categories; If the map image is not fully labeled, continue with the steps of selecting unlabeled targets from the map image and labeling the targets according to their categories; Once the map image annotation is completed, the annotated map image is identified as a map image sample annotated with multiple categories of targets.

6. The map correction method according to claim 5, characterized in that, The step of labeling the target according to its category includes: When the target belongs to the first category, the target is labeled with a box; the first category indicates that the bounding box of the target has a regular shape. When the target is classified as the second category, the target is labeled using an image mask; the second category indicates that the bounding box of the target is an irregular shape.

7. The map correction method according to claim 4, characterized in that, The step of training the initial localization model multiple times based on the training set to obtain a candidate localization model for each training round includes: The batch size is determined based on the size of the training set and the data processing performance of the initial localization model. Use the initial positioning model as the specified positioning model; Based on the training set and the batch size, the specified localization model is trained to obtain the candidate localization model corresponding to the current training round. If the set number of training rounds has not been completed, the candidate localization model corresponding to the current training round is taken as the designated localization model, and the step of training the designated localization model based on the training set and the batch number continues.

8. The map correction method according to claim 7, characterized in that, The step of training the specified localization model based on the training set and the batch size to obtain the candidate localization model corresponding to the current training round includes: The training set is divided into at least one sub-training set according to the batch quantity; For any of the aforementioned sub-training sets, each of the aforementioned map image samples in the training set is input into the specified positioning model for target positioning, thereby obtaining the predicted positioning of each of the aforementioned map image samples; Based on the predicted location of each map image sample and the location label of each map image sample, a second loss value is calculated; The model parameters of the specified localization model are adjusted based on the second loss value; If the current training round has not traversed all the aforementioned sub-training sets, for any of the aforementioned sub-training sets that have not been traversed in the current training round, the steps of inputting each of the aforementioned map image samples in the training set into the specified positioning model for target positioning and thereafter shall continue to be executed. If all the sub-training sets have been traversed in the current training round, the specified localization model that has been trained is determined as the candidate localization model corresponding to the current training round.

9. The map correction method according to claim 8, characterized in that, The step of inputting each of the map image samples in the training set into the designated positioning model for target positioning, and obtaining the predicted positioning of each of the map image samples, includes: For any map image sample in the training set, for targets belonging to the first category in the map image sample, target localization is performed using the specified localization model with a target detection task; for targets belonging to the second category in the map image sample, target localization is performed using the specified localization model with an image segmentation task, thereby obtaining the predicted localization of the map image sample.

10. A map correction device, characterized in that, include: The target localization module is configured to input an initial map into a trained target localization model to perform target localization and obtain a semantic map; wherein, the target localization model is trained based on map image samples labeled with multiple categories of targets; The correction prediction module is configured to input the semantic map and the initial map into the target multimodal large language model for correction prediction, and obtain the error localization and correction suggestions corresponding to the semantic map; wherein, the target multimodal large language model is trained based on map error samples, and the map error samples include the map image samples, the semantic map samples corresponding to the map image samples, and the correction labels corresponding to the semantic map samples. The correction module is configured to correct the erroneous localization in the semantic map based on the correction suggestions, thereby obtaining the target map corresponding to the initial map.