Method for making high-quality map multi-modal data set based on POI (Point of Interest) data
Through the automated multimodal dataset production method based on POI data, the problem of inefficient construction of map texts is solved, and high-quality multimodal data is efficiently generated, which improves the map recognition accuracy and generalization ability of large models.
Patent Information
- Application Number
- CN202510410973.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, manual construction of map text expression is inefficient, time-consuming and labor-intensive, and it is difficult to meet the needs of large-scale data for large-scale data in large-scale models, and it is easy to introduce labeling deviations, affecting the generalization ability and recognition accuracy of the model.
Using a high-quality map multimodal dataset production method based on POI data, we automatically traverse the map range to generate map pictures and expression texts. By setting the map range and scaling ratio, calculating the field of view, obtaining POI data and generating expression texts, and using computer automation processing to generate efficient and accurate multimodal datasets.
It greatly improves the data generation efficiency, generates a large amount of high-quality map multimodal data, meets the training needs of large models, improves the generalization ability and recognition accuracy of the model, and avoids the labeling deviation introduced by manual subjective judgment.
Smart Images

Figure CN120277166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of map data set production, and in particular to a method for producing a high-quality map multimodal data set based on POI data. Background Art
[0002] With the development of artificial intelligence technology, large models represented by DeepSeek are accelerating their integration and application in various industries. How to enable large models to automatically recognize maps has become a key challenge for current large model technology. Automatic recognition of maps can not only improve the ability of large models in geographic information processing, but can also be widely used in smart cities, autonomous driving, logistics optimization and other fields. However, there are many challenges in achieving this goal.
[0003] First, map data is usually highly complex and multi-scale, containing a wealth of geographic elements, such as roads, buildings, and waters, and the relationships between these elements are intricate. Second, different map sources and mapping standards may lead to differences in data formats and expressions, which increases the difficulty of model learning. In addition, the information in the map often needs to be understood in context, such as the connectivity of roads, the functional division of regions, etc., which places higher demands on the reasoning ability of the model.
[0004] To solve these problems, researchers began to explore ways to combine deep learning with traditional geographic information systems (GIS). By leveraging the powerful representation capabilities of large models, features in maps can be better extracted, and more efficient recognition algorithms can be built in combination with domain knowledge. For example, large models such as DeepSeek can learn a large number of geographic data patterns through pre-training, so that they can quickly adapt to and complete recognition tasks when faced with new maps. However, how to quickly generate massive maps and corresponding text expressions for large model training is key to the recognition effect of large models. Existing methods basically rely on manual construction methods, that is, professionals manually write the expression text of each map for each map. This method is not only inefficient, time-consuming and labor-intensive, but also difficult to meet the large-scale data requirements of large model training. The process of manually constructing map text expressions is not only limited by the experience and subjective judgment of professionals, but also easily introduces annotation bias, thereby affecting the generalization ability and recognition accuracy of the model. Summary of the invention
[0005] The purpose of the present invention is to overcome the above-mentioned technical deficiencies and provide a method for producing a high-quality map multimodal dataset based on POI data, so as to solve the technical problems in the prior art that the manual construction of map text expression is inefficient, time-consuming and labor-intensive, and difficult to meet the requirements of large-scale data for large-scale model training.
[0006] To achieve the above technical objectives, in a first aspect, the technical solution of the present invention provides a method for producing a high-quality map multi-modal dataset based on POI data, including the steps:
[0007] Set the map range and map scale, calculate the range of each map view at the current map scale, and count the number of the map views.
[0008] Read the map view information and display the current map view.
[0009] Obtain the POI data within the current map view range, and display the spatial position distribution of the POI with type icons in the current map view.
[0010] Select the POI type, and generate an expression text expressing the POI situation of the map view based on the type, quantity, and distribution of the POI.
[0011] Traverse all the map views in the map range to generate the map pictures and expression texts of the map range.
[0012] Compared with the prior art, the beneficial effects of the present invention include:
[0013] Compared with the manual construction method, the method for producing a multi-modal dataset based on POI data can automatically traverse all the map views in the map range to generate map pictures and expression texts, greatly improving the data generation efficiency, being able to produce a large amount of map multi-modal data in a short time, and meeting the requirements of large model training for large-scale data. For example, it may take several days or even weeks for manual work to write expression texts for a small number of maps, while this method can process a large number of map views in a short time with the help of computer automated processing, quickly expanding the scale of the training dataset. Due to getting rid of the limitation of manual subjective judgment, the expression texts generated according to established rules and algorithms can be more objective and accurate, avoiding the annotation deviation introduced by the differences in experience and subjective factors of professionals, enabling the large model to learn based on more accurate and unbiased data, helping to improve the generalization ability and recognition accuracy of the model, and ensuring the effect of the large model in map automatic recognition applications.
[0014] According to some embodiments of the present invention, setting the map range includes the steps: The input format of the map range is in bbox format, represented by the upper left corner coordinates and lower right corner coordinates of a rectangle, and the unit is in longitude and latitude.
[0015] According to some embodiments of the present invention, reading the map view information includes the steps:
[0016] Set the tile service address of the map, and obtain the tile map based on the tile service address.
[0017] According to some embodiments of the present invention, obtaining POI data within the current map field of view and displaying the POIs as map pictures with icons in the current map field of view includes the steps of:
[0018] Dividing the current map field of view range into nine regions of a nine-square grid, counting the number of POIs falling in each grid, and displaying them on the current map field of view in icon form according to the coordinates of the POIs.
[0019] According to some embodiments of the present invention, generating an expression text expressing the situation of the POIs in the map field of view includes the steps of:
[0020] Obtaining the administrative region where the current map range is located through the map field of view and generating a map location text based on the administrative region;
[0021] Generating an expression text of the situation of the POIs in the current map field of view based on the selected POI types and quantities within the map field of view;
[0022] Generating a distribution description text of the POIs based on the number of POIs in each grid of the nine-square grid of the map field of view;
[0023] Generating a synthetic text based on the map location text, the expression text of the situation of the POIs, and the distribution description text.
[0024] According to some embodiments of the present invention, the POI types include:
[0025] Shops, restaurants, schools, hospitals, banks, gas stations, scenic spots, bus stops, etc., and each type is represented by a different icon.
[0026] In a second aspect, the present invention provides a high-quality map multi-modal dataset production system based on POI data, including:
[0027] A map selection module for setting the map range and map zoom ratio, calculating the range of each map field of view at the current map scale, and counting the number of the map fields of view;
[0028] A field of view acquisition module for reading the map field of view information and displaying the current map field of view;
[0029] A picture display module for obtaining POI data within the current map field of view and displaying the POIs as map pictures with icons in the current map field of view;
[0030] A text generation module, configured to select POI types, generate an expression text for expressing the POI situation of the map view based on the types, quantities, and distribution of the POIs, traverse all map views within the map range, and generate the map images and expression texts for the map range.
[0031] In a third aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method for making a high-quality map multi-modal dataset based on POI data as described in any one of the first aspects.
[0032] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where the abstract drawing should be exactly the same as one of the drawings in the specification drawings:
[0034] Figure 1 It is a flowchart of the method for making a high-quality map multi-modal dataset based on POI data provided by an embodiment of the present invention;
[0035] Figure 2 It is a flowchart of the method for making a high-quality map multi-modal dataset based on POI data provided by an embodiment of the present invention;
[0036] Figure 3 It is a flowchart of the method for making a high-quality map multi-modal dataset based on POI data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0038] It should be noted that although functional module division is performed in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the system or the sequence in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0039] In recent years, with the rapid development of generative artificial intelligence technology, methods based on large language models and multimodal learning have provided new ideas for solving this problem. For example, by combining computer vision technology and natural language processing technology, an end-to-end system can be designed that can extract key geographical features from raw map images and automatically generate text descriptions that conform to semantic logic. Specifically, such technologies can use deep learning models to segment and extract features from map images, identify core elements such as roads, buildings, and water areas, and infer the relationships between these elements by combining context information. On this basis, by introducing pre-trained large language models (such as DeepSeek), the extracted structured geographical information can be transformed into natural language descriptions to form a text expression corresponding to the map. In addition, technologies such as reinforcement learning or generative adversarial networks (GANs) are also introduced. For example, by designing a reward mechanism, the model can be guided to generate text descriptions that are more in line with the actual geographical scenario; or generative adversarial networks can be used to generate realistic map images and their corresponding text expressions to expand the scale of the training dataset. However, these methods are limited by the performance of the model, and the accuracy of map recognition and text generation is limited, so the quantitative relationships on the map cannot be accurately expressed.
[0040] To solve this bottleneck problem, a technical solution that can automatically generate maps and their corresponding text expressions is urgently needed. This method needs to have the following characteristics: First, it can efficiently process diverse map data, including map information from different sources and at different scales; Second, it can automatically generate high-quality, semantically rich text descriptions without manual intervention, accurately reflecting the geographical features and their relationships in the map; Finally, the generated data should have high scalability to support the training needs of large models in different application scenarios.
[0041] Therefore, inventing a technology that can automatically generate high-precision map expression texts not only helps to improve the performance of large models in map recognition training tasks but also will greatly promote the application of artificial intelligence technology in the field of geographical information. This will provide strong technical support for scenarios such as smart city construction, autonomous driving navigation, and logistics path optimization, and also lay a foundation for more complex and intelligent geographical information processing systems in the future.
[0042] Refer to Figures 1 to 3 , Figure 1 is a flowchart of a method for making a high-quality map multimodal dataset based on POI data provided by an embodiment of the present invention; Figure 2 is a flowchart of a method for making a high-quality map multimodal dataset based on POI data provided by an embodiment of the present invention; Figure 3Flowchart of a method for creating a high-quality map multi-modal dataset based on POI data provided by an embodiment of the present invention. The method for creating a high-quality map multi-modal dataset based on POI data includes but is not limited to the following steps:
[0043] Step S110: Set the map range and map scale, calculate the range of each map view at the current map scale, and count the number of map views.
[0044] Step S120: Read the map view information and display the current map view.
[0045] Step S130: Obtain the POI data within the current map view range, and display the spatial position distribution of the POIs with corresponding type icons in the current map view.
[0046] Step S140: Select the POI data in the view, and generate an expression text representing the POI situation of the map view based on the type, quantity, and distribution of the POIs.
[0047] Step S150: Traverse all the map views within the map range to generate map pictures and expression texts for the map range.
[0048] In one embodiment, the method for creating a high-quality map multi-modal dataset based on POI data includes the steps of: setting the map range and map scale, calculating the range of each map view at the current map scale, and counting the number of map views; reading the map view information and displaying the current map view; obtaining the POI data within the current map view range and displaying the POIs as map pictures with corresponding type icons in the current map view; selecting the POI type and generating an expression text representing the POI situation of the map view based on the type, quantity, and distribution of the POIs; traversing all the map views within the map range to generate map pictures and expression texts for the map range.
[0049] In dealing with the complexity and multi-scale characteristics of map data, by setting the map range and zoom ratio, calculating the view range of each map and counting the quantity, the present invention can refine the complex map data with multi-scale characteristics in a hierarchical and regional manner. For example, for a large-scale map covering the whole city to local blocks, the map views at different zoom ratios can focus on geographical elements at different scales, from the macroscopic urban block division to the microscopic street details, making the subsequent extraction of POI data and generation of expression texts more organized, helping the large model to gradually understand the complex relationships between elements at different scales, and reducing the difficulty of directly processing the whole complex map. The step of selecting POI types can screen out specific types of points of interest from a rich variety of geographical elements, avoiding the large model having to face a vast and complex array of all geographical elements at the beginning, but focusing on the parts more relevant to the application scenario. For example, when studying the application of smart city transportation, key POIs such as transportation hubs can be focused on, helping the large model to more efficiently learn the characteristics and relationships of key elements in the map and cope with the complexity challenges.
[0050] In solving the differences in different map sources and cartographic standards, the present invention constructs a relatively unified processing flow for map data. Regardless of the map source and the cartographic standard followed, as long as the POI data within the corresponding map view can be obtained, map images and expression texts are generated according to the established steps. For example, for maps from different surveying and mapping institutions, using different projection methods but covering the same area, they can all be transformed into standardized multi-modal datasets under this processing flow, reducing the interference caused by differences in data formats and expression methods to the learning of the large model, enabling the large model to learn and recognize based on standardized inputs.
[0051] In enhancing the model's understanding and reasoning ability of map context, by generating expression texts based on the type, quantity and distribution of POIs, the present invention can integrate and sort out the information related to POIs in the map, reflecting the interrelationships and overall layout of the originally isolated POI icons in the text description, providing strong support for the large model to understand context-related information such as road connectivity and regional function division. For example, when describing the map view of a commercial area, the quantity and distribution of POIs such as shopping malls, restaurants and parking lots will be reflected in the text, and the large model can infer context-related information such as the function and traffic flow of the area based on this, thus enhancing its reasoning ability.
[0052] In solving the problems of the training data generation efficiency and scale of large models, compared with the manual construction method, this method for producing multi-modal data sets based on POI data can automatically traverse all the map views within the map range to generate map pictures and descriptive texts, greatly improving the data generation efficiency and being able to produce a large amount of map multi-modal data in a short time to meet the needs of large model training for large-scale data. For example, it may take days or even weeks for humans to write descriptive texts for a small number of maps, while this method can process a vast amount of map views in a short time with the help of computer automation, quickly expanding the scale of the training data set. Due to getting rid of the limitations of human subjective judgment, the descriptive texts generated according to established rules and algorithms can be more objective and accurate, avoiding annotation biases introduced by differences in the experience of professionals and subjective factors, enabling large models to learn based on more accurate and unbiased data, helping to improve the generalization ability and recognition accuracy of the models, and ensuring the effect of large models in map automatic recognition applications.
[0053] POI types include: stores, restaurants, schools, hospitals, banks, gas stations, scenic spots, bus stops, etc.
[0054] In one embodiment, the method for producing a high-quality map multi-modal data set based on POI data includes the steps of: setting the map range and map zoom ratio, calculating the range of each map view at the current map ratio, and counting the number of map views; reading the map view information and displaying the current map view; obtaining the POI data within the current map view range and displaying the POIs as icons in the current map view as a map picture; selecting the POI type and generating a descriptive text expressing the POI situation of the map view based on the type, quantity, and distribution of the POIs; traversing all the map views within the map range to generate map pictures and descriptive texts for the map range. Setting the map range includes the steps of: the input format of the map range is in bbox format, represented by the upper left corner coordinates and lower right corner coordinates of a rectangle, and the unit is in longitude and latitude. For example, BBOX=-74.0479,40.6829,-73.9067,40.8782 defines a rectangular area from the lower left corner (-74.0479,40.6829) to the upper right corner (-73.9067,40.8782).
[0055] The bbox format is a relatively common standard format for representing map ranges in the field of geographic information and is widely recognized and applied in many software, platforms related to geographic data, and various data processing processes. This means that the map ranges set in this format can be easily integrated and interacted with other geographic data resources without compatibility issues due to format differences. For example, when docking a multi-modal dataset of a produced map with an external geographic information database, if this standard bbox format is used to define the range, the data fusion process will be smoother. This embodiment provides a unified specification for setting map ranges, enabling different researchers and development teams to operate based on the same format standard when dealing with map-related tasks, which is convenient for communication and collaboration. For example, in a smart city project involving multiple units, if all parties define the map range according to the bbox format, they can accurately and clearly divide the data areas they are responsible for, avoiding confusion caused by unclear range definition.
[0056] Using latitude and longitude as units can achieve very precise geographical location positioning and accurately frame the desired map range. Every location on the earth's surface can be represented by a unique latitude and longitude coordinate. Therefore, by using this coordinate form to determine the four corner points of a rectangular area, a specific geographical area can be accurately covered. Whether it is a small-scale block, scenic area, or a large-scale city, province, etc., it can be accurately set according to actual needs, which lays an accurate foundation for subsequent operations such as accurately obtaining POI data within the corresponding range.
[0057] Furthermore, reading map view information includes the steps of: setting the tile service address of the map and obtaining the tile map based on the tile service address. When setting the tile service address of the map, the map service address can be a publicly available map on the network or a map published by the unit itself.
[0058] Furthermore, obtaining POI data within the current map view range and displaying the POIs as icons of corresponding types in the current map view as a map picture includes the steps of: dividing the current map view range into nine regions in a nine-grid pattern, counting the number of POIs falling in each grid, and displaying them in the current map view in icon form according to the coordinates of the POIs.
[0059] By dividing the map view range into nine regions of a nine-square grid, the distribution of POIs in different local spaces can be shown in more detail. Compared with directly displaying POI icons across the entire map view, the nine-square grid division allows users to immediately see the differences in the density of POIs in each small region. For example, when analyzing the distribution of commercial POIs in different city blocks, it is possible to intuitively distinguish which nine-square grid regions have a large number of shopping malls, stores, etc., and which have relatively few, thus better grasping the uneven distribution characteristics of POIs in space and making the visualization effect of the spatial distribution more prominent. It is convenient to compare different grid regions, which helps to quickly lock in the key areas of concern. For instance, when studying POIs related to transportation hubs, by comparing the number of transportation stations in each nine-square grid, it is easy to identify the busy transportation areas with concentrated POIs and relatively sparse areas, and then deeply analyze the connections with surrounding areas, traffic flow, etc. in the key areas, providing an intuitive visual reference for subsequent applications such as transportation planning and resource allocation.
[0060] Counting the number of POIs falling into each grid adds a refined data statistical dimension. This regional statistical method can provide richer information than simply counting the total number of POIs in the entire map view. It not only allows us to understand the overall scale of POIs but also know the number of POIs in different local spaces, which helps with more in-depth data analysis. For example, analyzing the differences in the development levels of different regions (measured by the number of different types of POIs), the rationality of resource allocation, etc., providing stronger support for data-based decision-making. Based on the statistics and display of the nine-square grid, it helps to discover the laws and patterns of POI distribution. For example, observing that the types and numbers of POIs in some adjacent grids show similarities or gradual change patterns may imply the coherence of regional functions or the influence of geographical factors on POI layout, etc. This has positive significance for studying the formation mechanism of POIs in geographical space and for predictive modeling, etc., and can help uncover the spatial correlation information hidden behind the data.
[0061] When POIs are displayed as icons according to coordinates on the map view, due to the division of the nine-square grid, it avoids the visual chaos and difficulty in distinguishing caused by a large number of POI icons concentrated in one place, making the entire map image more regular and orderly, reducing the visual complexity, and enabling users to browse and understand the POI distribution information conveyed by the map image more efficiently. Especially when dealing with map views with a large number of POIs and a relatively complex distribution, this advantage is more obvious, improving the efficiency and accuracy of information transmission.
[0062] When it is necessary to search for specific types of POIs or pay attention to the POI situation within a certain area, with the help of the nine-square grid partition, it is possible to quickly locate the corresponding grid area, and then search for the POIs corresponding to the icons within that area. Compared with searching on the entire unpartitioned map, it greatly saves time and effort and improves the convenience of information acquisition.
[0063] Furthermore, generate an expression text representing the POI situation in the map view, including the steps of: obtaining the administrative region where the current map range is located through the map view and generating a map location text based on the administrative region; generating an expression text of the POI situation in the current map view based on the selected POI types and quantities within the map view; generating a distribution description text of the POIs based on the POI quantities in each grid of the nine-square grid in the map view; and generating a synthesized text based on the map location text, the expression text of the POI situation, and the distribution description text.
[0064] Specific example: (1) Obtain the administrative region where the current map is located through this view range, such as *** District (County), *** Street (Town), *** City, *** Province, and generate a map location text, which is that this map is in ***** place; (2) Use the types and the queried quantities of POIs when querying in step 3 to generate the feature text of this map, for example: there are 10 tourist attractions on the map; (3) Generate a distribution description of the POI points based on the POI quantities in each grid of the nine-square grid in the map view. For example, if the quantity in the middle grid is the largest, then generate that the tourist attractions are mainly distributed in the middle area of the map; (4) Combine the generated map texts to form a descriptive text.
[0065] By obtaining the administrative region where the current map range is located and generating a map location text, it is possible to endow the POI information in the entire map view with a clear geographical background, enabling people to clearly know the large regional scope where these POIs are located. For example, clearly indicating that it is in a certain district of a certain city in a certain province is very helpful for subsequent users to understand the distribution of POIs at the macroscopic level and their relationship with the surrounding areas. Whether it is for urban planning analysis or studying the relationship between regional economic development and POI layout, more in-depth discussions can be carried out based on accurate regional positioning. Generating the corresponding expression text based on the selected POI types and quantities within the map view can accurately extract the key feature information of the POIs in that area. For example, specifically stating the quantities of different types of POIs such as how many restaurants and how many shopping malls, directly focusing on these elements closely related to the application scenarios, enabling users to quickly understand the core composition of the POIs in this map view and providing very targeted data support in application scenarios such as commercial site selection and tourism resource recommendation that require attention to specific types of POIs.
[0066] Generate distribution description text based on the number of POIs in each cell of the nine - grid of the map view, which depicts in detail the spatial distribution characteristics and patterns of POIs. This helps to reveal the spatial features such as the density differences and concentration trends of POIs between different local areas. For example, it can describe that the number of POIs in a certain nine - grid area is large and concentrated, while the adjacent grids are relatively sparse. For analyzing issues such as functional zoning, traffic flow guidance, and resource allocation rationality in the geographical space, it can provide detailed and valuable basis for spatial distribution information.
[0067] Generate a synthetic text based on the map location text, the expression text of POI situation, and the distribution description text, which organically integrates key information from multiple aspects to form a complete, comprehensive, and well - structured description content. Such a synthetic text enables users to obtain information about various levels of the map view at one time, from macroscopic regional positioning to mesoscopic POI composition and then to microscopic POI distribution. Whether for manual viewing and analysis or as input data for machine - learning models, it can more efficiently understand the rich information contained therein, improving the utilization value of information and the convenience of application. Since it covers information in different dimensions, the generated synthetic text can be applied to a variety of different application scenarios and requirements. Whether it is scientific research projects related to geographical information, various business applications in smart city construction, or ordinary people querying and understanding the POI distribution in a geographical area, useful information can be obtained from this comprehensive text, improving the generality and adaptability of the text among different fields and user groups, and endowing it with a wider application value.
[0068] In one embodiment, a high - quality map multi - modal dataset production system based on POI data includes: a map selection module for setting the map range and map zoom ratio, calculating the range of each map view at the current map scale, and counting the number of map views; a view acquisition module for reading map view information and displaying the current map view; a picture display module for obtaining POI data within the current map view range and displaying the POIs as icons on the current map view as a map picture; a text generation module for selecting POI types, generating expression text expressing the POI situation of the map view based on the type, quantity, and distribution of POIs, and traversing all map views within the map range to generate map pictures and expression texts for the map range.
[0069] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0070] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0071] In addition, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, which are executed by a processor or a controller, for example, executed by a processor in the above terminal embodiment, so that the processor can execute the method for creating a high-quality map multimodal dataset based on POI data in the above embodiment.
[0072] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0073] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.
[0074] The specific embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A method for producing a high-quality map multi-modal dataset based on POI data, characterized in that, Including the steps: Set the map range and map zoom ratio, calculate the range of each map view at the current map ratio, and count the number of the map views; Read the map view information and display the current map view; Obtain the POI data within the current map view range, and display the spatial position distribution of the POI with type icons in the current map view; Select the POI data in the view, and generate an expression text expressing the POI situation of the map view based on the type, quantity, and distribution of the POI; Traverse all map views within the map range to generate the map picture and expression text of the map range; 2. The method for producing a high-quality map multi-modal data set based on POI data according to claim 1, wherein, Setting the map range includes the steps: The input format of the map range is in bbox format, represented by the upper left corner coordinates and the lower right corner coordinates of a rectangle, and the unit is longitude and latitude; 3. The method for producing a high-quality map multi-modal data set based on POI data according to claim 1, wherein Reading the map view information includes the steps: Set the tile service address of the map, and obtain the tile map based on the tile service address; 4. The method for producing a high-quality map multi-modal data set based on POI data according to claim 3, wherein, Obtain the POI data within the current map view range, and display the POI with type icons in the current map view, including the steps: Divide the current map view range into nine regions of a nine-square grid, count the number of the POI falling in each grid, and display it on the current map view in the form of icons according to the coordinates of the POI; 5. The method for producing a high-quality map multi-modal data set based on POI data according to claim 4, wherein, Generating the expression text expressing the POI situation of the map view includes the steps: Obtain the administrative region where the current map range is located through the map view, and generate a map location text based on the administrative region; Generate an expression text of the POI situation of the current map view based on the type and quantity of the POI selected within the map view; Generate a distribution description text of the POI based on the number of POI in each grid of the nine-square grid of the map view; Generate a synthesized text based on the map location text, the expression text of the POI situation, and the distribution description text; 6. The method for producing a high-quality map multi-modal data set based on POI data according to claim 1, wherein The POI types include: Stores, restaurants, schools, hospitals, banks, gas stations, scenic spots, and bus stops; 7. A high-quality map multi-modal dataset production system based on POI data, characterized in that, Including: A map selection module for setting the map range and map zoom ratio, calculating the range of each map view at the current map ratio, and counting the number of the map views; A view acquisition module for reading the map view information and displaying the current map view; A picture display module for obtaining the POI data within the current map view range, and displaying the POI as a map picture with corresponding type icons in the current map view; A text generation module for selecting the POI type, and generating an expression text expressing the POI situation of the map view based on the type, quantity, and distribution of the POI, and traversing all map views within the map range to generate the map picture and expression text of the map range; 8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method for making a high-quality map multi-modal dataset based on POI data as described in any one of claims 1 to 6.
Citation Information
Cited By
Intelligent map reading method and device for hierarchical statistical map
CN120766307A
Intelligent Map Reading Method and Device for Hierarchical Statistical Maps
CN120766307B