Geospatial information sample generation method and device, equipment and medium
By extracting structured triplets from map service data and using large language models to generate natural language texts, the problem of generating high-quality geospatial information samples in the prior art is solved, and efficient, diverse and semantically accurate data generation is achieved.
Patent Information
- Application Number
- CN202510421483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art is difficult to efficiently generate high-quality geospatial information samples, manual labeling is costly and inefficient, and method based on rule templates is poor in flexibility, data enhancement may destroy semantic integrity, and traditional models are difficult to adapt to diversified corpus.
By obtaining map service data, determining the geographical elements of the research area, calculating spatial relationships to form a structured triple, analyzing the geospatial information text set and clustering, and using a large language model to generate natural language text based on the propt template expansion.
It improves the diversity of geospatial information data generation, avoids the high cost and inefficiency problems of manual annotation, overcomes the flexibility limitations of rule templates, and ensures the semantic accuracy of generated samples.
Smart Images

Figure CN119938806A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spatiotemporal knowledge extraction, and in particular to a method, device, equipment and medium for generating geographic space information samples. Background Art
[0002] Geospatial information is used to describe the location, characteristics, distribution and mutual relationship of geographic entities in space. It deeply reflects the complex attributes, dynamic changes and internal connections of geographic elements, and provides vital knowledge support for regional planning, resource management, disaster warning and other fields. It is also an important knowledge component for the construction of spatiotemporal knowledge graphs. The sources of geospatial information are very wide, including not only basic surveying and mapping, sensors, etc., but also text data such as news reports, scientific and technological literature and social media. These text data contain rich geographic information and are easy to obtain, have a wide coverage and high update frequency. They are of great value for improving the timeliness of knowledge graphs and updating knowledge graphs.
[0003] With the advent of the digital age, the speed of text data generation is growing exponentially. Therefore, the rapid and accurate extraction and analysis of geospatial information in text data will help promote the deep utilization of geospatial data and intelligent geographic semantic understanding. Information extraction is a key technology in natural language processing, which aims to extract structured information from unstructured text, mainly including named entity recognition, relationship extraction and other technologies. Related model methods have experienced a development process from rule-based template matching, statistical learning, deep learning to pre-trained models. In particular, model algorithms represented by deep learning and pre-trained language models rely on large-scale, high-quality sample data for model training. However, the scarcity of text annotation data in the field of geospatial information makes it difficult for traditional general models that rely on a large amount of annotated data to fully capture and identify the geographic information contained in the text. In addition, the label density of geographic information is relatively sparse, which makes it difficult for the model to distinguish the boundaries of different geographic information, and thus cannot accurately locate its entities and describe relationships. Therefore, how to obtain high-quality geospatial information training samples becomes the key.
[0004] The existing methods for obtaining annotated corpora mainly include manual annotation, rule-based template generation, and data augmentation. Among them, manual annotation ensures high accuracy and meets the task requirements of specific fields by finely processing and annotating corpora. For example, some projects have built large-scale syntactic treebanks through expert annotation, and some projects have combined semantic role annotation with core reference annotation to provide rich corpus support. However, manual annotation relies on manual fine processing of each data, which is inefficient and costly. At the same time, due to subjective bias, it is difficult to unify the quality and consistency of the annotated data. The rule-based template method provides data generation through cosine setting of clear logical rules or semantic templates, for example, by designing rules and templates to convert structured data into natural language text; based on predefined rules and templates, a large amount of virtual data is generated, which is widely used in testing and data generation scenarios. However, these methods often cannot adapt to the changes in corpus structure or task requirements, resulting in inaccurate or incomplete extracted information. The fixed nature of the rules makes them lack flexibility and scalability. For each new task or data set, the rule template needs to be adjusted frequently. Data augmentation effectively expands the data scale and improves the diversity of the corpus by transforming the existing corpus (such as data replacement, synonym replacement, sentence conversion, etc.). For example, a large amount of data is obtained through data replacement methods and the robustness of text classification models is improved; back-translation is widely used in the field of machine translation to generate diverse parallel corpora. However, this process may damage the semantic integrity of the original corpus, and the quality of the generated samples is difficult to fully guarantee. In addition, data augmentation is usually based on existing data and cannot generate truly new information. Summary of the invention
[0005] The present invention provides a method, device, equipment and medium for generating geospatial information samples, the purpose of which is to improve the diversity of geospatial information data generation.
[0006] In order to achieve the above object, the present invention provides a method for generating a geospatial information sample, comprising: Step 1, obtaining map service data of the study area, and determining geographic elements within the study area based on the map service data; Step 2: Calculate the spatial relationship between geographic elements based on map service data to obtain structured triples of geographic spatial information; Step 3, analyzing the obtained geospatial information text set corresponding to the study area to form a text description structure, and clustering the text description structure to obtain multiple clustering results; Step 4, combining the structural features of the geospatial information text set to identify common features of the text description structure in multiple clustering results, and analyzing each clustering result based on the common features to obtain a spatial relationship text description form; Step 5, encode the spatial relationship text description form, and input the geospatial information structured triples and the encoded spatial relationship text description into the large language model, and the large language model expands the geospatial information structured triples based on the prompt template and the encoded spatial relationship text description to generate natural language text; Step 6: Merge the natural language text with the structured triples of geospatial information to obtain a geospatial information sample.
[0007] Specifically, before determining the geographic features within the study area based on map service data, the following is also included: The map service data is formatted and parsed, and abnormal data is removed to obtain preprocessed map service data. The abnormal data includes standard data caused by human errors and missing data.
[0008] Furthermore, the spatial relationship between geographic elements is calculated based on the map service data to obtain structured triplets of geographic spatial information, including: Build a spatial index for all geographic features in the study area through R-tree; A geographical element is randomly selected from all the geographical elements in the study area as the target geographical element, and the neighbor elements whose distance to the target geographical element is less than the preset threshold are obtained through spatial index to form the nearest neighbor candidate set; Randomly select several neighbor elements from the nearest neighbor candidate set, and calculate the spatial relationship between each neighbor element and the target geographic element, and obtain the calculation result in the form of triples; Repeatedly select the target geographic element, and calculate the spatial relationship between the target geographic element selected each time and each neighbor element, and obtain the calculation result in the form of a triple; All calculation results in triple form that meet the saving conditions are saved to obtain structured triples of geospatial information.
[0009] Furthermore, before saving all the calculation results in triple form that meet the saving conditions, including: Count the spatial relationship categories in the calculation results of all triples; Set the saving conditions according to the spatial relationship category: in, Indicates The spatial relationship of the categories, Indicates The target weight ratio, Indicates the number of categories.
[0010] Furthermore, the obtained geospatial information text set corresponding to the study area is analyzed to form a text description structure, including: According to the data characteristics of the geospatial information field, a set of geospatial information texts corresponding to the research area is obtained on public information websites; The Qwen-max model is used to perform part-of-speech tagging on the geospatial information text set, and the geographic entities, action descriptions and modification information in the geospatial information text set are identified; The grammatical structure of the sentences in the geospatial information text set after part-of-speech tagging is parsed through the dependency tree to form a text description structure.
[0011] Furthermore, the text description structure is clustered to obtain multiple clustering results, including: Use the Sentence-BERT model to vectorize the text description structure to obtain a high-dimensional text description structure; The DBSCAN model is used to perform density clustering on high-dimensional text description structures to obtain multiple clustering results, each of which contains a group of similar high-dimensional text description structures.
[0012] Specifically, the large language model expands the structured triples of geospatial information based on the prompt template and the encoded spatial relationship text description to generate the natural language text expression: in, represents natural language text, represents a large language model, Represents a structured triple of geospatial information, , They all represent geographical elements. Representing geographic features With geographical elements The spatial relationship between Represents the encoded text description of the spatial relationship.
[0013] The present invention also provides a device for generating a geographic space information sample, comprising: An acquisition module is used to acquire map service data of the study area and determine geographic elements within the study area based on the map service data; A calculation module is used to calculate the spatial relationship between geographic elements based on map service data to obtain structured triples of geographic spatial information; The first analysis module is used to analyze the acquired geographic spatial information text set corresponding to the study area to form a text description structure, and cluster the text description structure to obtain multiple clustering results; The second analysis module is used to identify common features of text description structures in multiple clustering results in combination with structural features of the geospatial information text set, and analyze each clustering result based on the common features to obtain a spatial relationship text description form; The expansion module is used to encode the spatial relationship text description form, and input the structured triples of geospatial information and the encoded spatial relationship text description into the large language model. The large language model expands the structured triples of geospatial information based on the prompt template and the encoded spatial relationship text description to generate natural language text; The merging module is used to merge the natural language text with the structured triples of geospatial information to obtain geospatial information samples.
[0014] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method for generating geographic space information samples is implemented when the processor executes the computer program.
[0015] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for generating a geographic space information sample is implemented.
[0016] The above scheme of the present invention has the following beneficial effects: The present invention determines the geographical elements in the study area based on the map service data of the study area, calculates the spatial relationship between the geographical elements, and obtains the structured triples of geographical spatial information; analyzes the obtained geographical spatial information text set corresponding to the study area, forms a text description structure and clusters it to obtain multiple clustering results; identifies the common features of the text description structure in the multiple clustering results, analyzes each clustering result, and obtains the spatial relationship text description form; encodes the spatial relationship text description form, and inputs the structured triples of geographical spatial information and the encoded spatial relationship text description into a large language model for expansion based on a prompt template to generate natural language text; The language text is merged with the structured triples of geospatial information to obtain geospatial information samples; compared with the prior art, the present invention utilizes spatial relationship calculation to extract structured geospatial triples from map service data, thereby avoiding the high cost and inefficiency of manual annotation; the acquired geospatial information text set corresponding to the study area is analyzed to obtain the spatial relationship text description form, and the structured triples of geospatial information and the encoded spatial relationship text description are input into the large language model for expansion based on the prompt template, thereby avoiding the semantic deviation that may be caused by data enhancement, overcoming the limitations of rule templates in processing diverse corpora, and improving the diversity of geospatial information data generation.
[0017] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of a flow chart of an embodiment of the present invention; Figure 2 It is a schematic diagram of the structure of a device for generating geographic space information samples according to an embodiment of the present invention; Figure 3 Schematic diagram of the structure of a terminal device in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0021] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0022] In view of the existing problems, the present invention provides a method, device, equipment and medium for generating a geographic space information sample.
[0023] like Figure 1 As shown, an embodiment of the present invention provides a method for generating a geospatial information sample, comprising: Step 1, obtaining map service data of the study area, and determining geographic elements within the study area based on the map service data; Step 2: Calculate the spatial relationship between geographic elements based on map service data to obtain structured triples of geographic spatial information; Step 3, analyzing the obtained geospatial information text set corresponding to the study area to form a text description structure, and clustering the text description structure to obtain multiple clustering results; Step 4, combining the structural features of the geospatial information text set to identify common features of the text description structure in multiple clustering results, and analyzing each clustering result based on the common features to obtain a spatial relationship text description form; Step 5, encode the spatial relationship text description form, and input the geospatial information structured triples and the encoded spatial relationship text description into the large language model, and the large language model expands the geospatial information structured triples based on the prompt template and the encoded spatial relationship text description to generate natural language text; Step 6: Merge the natural language text with the structured triples of geospatial information to obtain a geospatial information sample.
[0024] The embodiment of the present invention takes a first-tier city in the mainland as the research area, and the map service data includes point of interest (POI, Point of Interest) data, road data and building data in a public map (OSM, Open Street Map).
[0025] Most preferably, before determining the geographic features within the study area based on the map service data, the method further includes: The map service data is formatted and parsed, and abnormal data is removed to obtain preprocessed map service data. The abnormal data includes standard data caused by human errors and missing data.
[0026] Specifically, the spatial relationship between geographic elements is calculated based on the map service data to obtain structured triplets of geographic spatial information, including: All geographic features in the study area are analyzed through R-tree Build a spatial index , the expression is: All geographic features within the study area Randomly select a geographic feature as the target geographic feature , through the spatial index Get the target geographic feature The neighbor elements whose distance is less than the preset threshold constitute the nearest neighbor candidate set , the expression is: in, represents the number of neighbor elements, represents the k-nearest neighbor algorithm, , Indicates Neighbor features whose distance is less than the preset threshold; From the nearest neighbor candidate set Randomly select Neighborhood features , and calculate the relationship between each neighbor feature and the target geographic feature The spatial relationship between them is used to obtain the calculation result in the form of a triple. The calculation expression is: in, Indicates Neighbor features whose distance is less than the preset threshold. Indicates the target geographic element and the The spatial relationship between neighboring features whose distance is less than the preset threshold. ; Repeatedly select target geographic features , and calculate the target geographic elements selected each time The spatial relationship between each neighbor element includes: adjacent relationship (such as location A+action+location B), orientation relationship (such as location A+location word+location B), intersection relationship (such as location A+and+location B+intersect / intersect / intersect), measurement relationship (such as location A+distance+location B+additional information), inclusion relationship (such as location A+located+location B+within), and separation relationship (such as location A+action+quantity+distance unit+location B), and the calculation result is obtained in the form of triples; All calculation results in triple form that meet the saving conditions are saved to obtain structured triples of geospatial information.
[0027] Most preferably, before saving all the calculation results in triplet form that meet the saving condition, the method further includes: Count all the spatial relationship categories in the calculation results of triple form ; In order to ensure the balance of the number of different spatial relationship categories, each time the calculation results in triple form are saved, the saving conditions need to be set according to the spatial relationship category: in, Indicates The spatial relationship of the categories, Indicates The target weight ratio, Indicates the number of categories.
[0028] Specifically, the obtained geospatial information text set corresponding to the study area is analyzed to form a text description structure, including: Since the text data in the field of geospatial information has complex characteristics of multi-scale, multi-granularity and multi-dimensionality, it faces the problems of scarce annotated data and sparse label density. According to the data characteristics in the field of geospatial information, the embodiment of the present invention efficiently obtains the geospatial information text set corresponding to the research area on the public information website through the web crawler tool. The public information website is various travel websites, and the geospatial information text set can be travel guides shared by multiple travelers. The Qwen-max model is used to perform part-of-speech tagging on the geospatial information text set, and the geographic entities (such as place names), action descriptions (such as verbs) and modifying information (such as time, distance, etc.) in the geospatial information text set are identified, which can be expressed as: and ,in, Represents a sequence of words, Indicates the corresponding part-of-speech tag; Through the dependency tree The grammatical structure of the sentences in the geospatial information text set after part-of-speech tagging is parsed to form a text description structure, where: Represents a word node, It represents dependency relationship edge. The grammatical structure includes subject-predicate relationship and object relationship.
[0029] Specifically, we use web crawler tools to efficiently obtain geospatial information text sets corresponding to the study area on public information websites, including: Efficiently obtain text information related to the research area through web crawler tools on public information websites; Regular expressions are used to remove special characters and escape characters from text information related to the study area to ensure the standardization of the extracted text; By using the sentence segmentation method, the long paragraphs of text information related to the research area are broken down into independent sentences to obtain a set of candidate sentences that are convenient for subsequent processing; Combined with the powerful data processing and reasoning capabilities of the large language model, the screening rules are established with the help of domain experts’ experience, and multiple rounds of screening are performed on the candidate sentence set. The large language model used in the embodiment of the present invention is the Qwen-max model. Suppose the sentence set after each round of screening is , the screening logic can be described as: in, Represents the set of sentences after the next round of screening, Represents a filtering operation. Indicates the filtering rules;
[0030] After multiple rounds of screening, a high-quality geospatial information text set rich in geospatial descriptions was finally constructed.
[0031] The embodiment of the present invention takes "A city museum is located at the junction of a river and a river" as an example, performs part-of-speech tagging and dependency syntactic analysis, identifies the core components in the sentence and their syntactic relations, and then combines the idea of frame semantics to annotate each component in the sentence into a label that conforms to the semantic role, such as labeling the geographic entity as Location, labeling "located" as V, and labeling the junction as Junction. The text description structure is: .
[0032] Specifically, the text description structure is clustered to obtain multiple clustering results, including: Use the Sentence-BERT model to vectorize the text description structure to obtain a high-dimensional text description structure; The DBSCAN model is used to perform density clustering on high-dimensional text description structures to obtain multiple clustering results, each of which contains a group of similar high-dimensional text description structures.
[0033] It should be noted that the Sentence-BERT model is a sentence-level embedding model that can efficiently generate semantically rich sentence vectors and support fast text similarity calculation and clustering analysis; the DBSCAN model can effectively process high-dimensional text vector data, automatically identify noise points, and automatically determine the number of clusters based on density; after combining the Sentence-BERT model for text vectorization, neighborhood density clustering is performed through the DBSCAN model, which can effectively identify text description forms with similar structures.
[0034] Therefore, the text description structure First, the Sentence-BERT model is applied for vectorization to convert each text description into a high-dimensional vector representation to obtain a high-dimensional text description structure; then, the DBSCAN model is used for density clustering to divide the high-dimensional text description structure into several clusters. , each cluster Contains a group of similar text description structures, and each cluster corresponds to each clustering result one by one; the DBSCAN model determines the core points of the cluster through the following conditions: in, Represents the text description structure With radius The neighborhood within the range of Indicates the density threshold. When a text describes a structure When this condition is met, it is regarded as a core point and expanded into a cluster. This time, the area radius is set is 0.3, the minimum number of neighbor points is 10.
[0035] Finally, through in-depth analysis of all clustering results and combining the structural characteristics of the geospatial information text set, the common characteristics of the text description structure contained in different clustering results can be identified; through in-depth analysis of the center point or representative text of each cluster center, the corresponding spatial relationship text description form can be extracted. , for example, adjacent relations (such as place A + action + place B), directional relations (such as place A + directional word + place B), intersection relations (such as place A + and + place B + intersect / intersect / intersect), measurement relations (such as place A + distance + place B + additional information), inclusion relations (such as place A + located + place B + within), and separation relations (such as place A + action + quantity + distance unit + place B).
[0036] It should be noted that, in order to further generate semantically rich natural language text, the embodiment of the present invention adopts a hybrid method combining Chain-of-Thought and Zero-Shot prompting to construct a series of highly general prompt templates. According to the logical deduction steps of the prompt template, the large language model is guided to gradually expand the semantics of the triples, while ensuring that the generated sentences are natural, fluent and rich in information.
[0037] Specifically, the large language model expands the structured triples of geospatial information based on the prompt template and the encoded spatial relationship text description, and generates the expression of natural language text as follows: in, represents natural language text, represents a large language model, Represents a structured triple of geospatial information, , They all represent geographical elements. Representing geographic features With geographical elements The spatial relationship between Represents the encoded text description of the spatial relationship.
[0038] The embodiment of the present invention takes "a certain commercial plaza includes a certain restaurant" as an example, and the generated natural language text is "a certain commercial plaza has everything. If you are tired of shopping, you can go to a certain restaurant to sit down and eat something to replenish your energy. This place is quite lively and there are usually many people. You have to come early to park, otherwise you won’t be able to find a spot even after walking around for a long time."
[0039] In order to verify the effectiveness of the method, the embodiment of the present invention firstly generates 3000 geospatial information corpora as positive samples in two different ways; at the same time, collects the general information extraction corpus InstructIE dataset as negative samples, and constructs two datasets in a 1:1 ratio and processes them into Alpaca format, wherein 300 geospatial information corpora generated by the two methods and 600 general information extraction corpora are selected as test sets. One way is to generate by the proposed method, and the other is to generate by direct prompting. Next, the remaining parts of the two datasets except the test set are divided into two groups of few samples (1200 samples) and many samples (4800 samples) to ensure that the training process covers sample distributions of different scales. On this basis, ChatGLM3-6B is selected as the baseline model, and the instruction supervised fine-tuning method is used to train until the model converges to obtain different training models. Finally, the accuracy index of different models is verified on the test set to prove the effectiveness of the proposed method. The verification results are shown in Table 1 below.
[0040] Table 1 Verification results
[0041] The embodiment of the present invention determines the geographic elements in the study area based on the map service data of the study area, calculates the spatial relationship between the geographic elements, and obtains the structured triples of geographic spatial information; analyzes the acquired geographic spatial information text set corresponding to the study area to form a text description structure and clusters it to obtain multiple clustering results; identifies the common features of the text description structure in the multiple clustering results, analyzes each clustering result, and obtains the spatial relationship text description form; encodes the spatial relationship text description form, and inputs the structured triples of geographic spatial information and the encoded spatial relationship text description into a large language model for expansion based on a prompt template to generate natural language text; The language text is merged with the structured triples of geospatial information to obtain geospatial information samples; compared with the prior art, the embodiment of the present invention utilizes spatial relationship calculation to extract structured geospatial triples from map service data, thereby avoiding the high cost and inefficiency of manual annotation; the acquired geospatial information text set corresponding to the study area is analyzed to obtain the spatial relationship text description form, and the structured triples of geospatial information and the encoded spatial relationship text description are input into the large language model for expansion based on the prompt template, thereby avoiding the semantic deviation that may be caused by data enhancement, overcoming the limitations of rule templates in processing diverse corpora, and improving the diversity of geospatial information data generation.
[0042] Corresponding to the method for generating geographic space information samples described in the above embodiment, as follows Figure 2As shown, the embodiment of the present invention further provides a device 100 for generating a geographic space information sample, and the device 100 for generating a geographic space information sample includes: The acquisition module 101 is used to acquire the map service data of the study area and determine the geographical elements in the study area based on the map service data; The calculation module 102 is used to calculate the spatial relationship between geographic elements based on the map service data to obtain a structured triple of geographic spatial information; The first analysis module 103 is used to analyze the acquired geographic spatial information text set corresponding to the study area to form a text description structure, and cluster the text description structure to obtain multiple clustering results; The second analysis module 104 is used to identify common features of the text description structure in multiple clustering results in combination with the structural features of the geographic spatial information text set, and analyze each clustering result based on the common features to obtain a spatial relationship text description form; The expansion module 105 is used to encode the spatial relationship text description form, and input the structured triple of geospatial information and the encoded spatial relationship text description into the large language model, and the large language model expands the structured triple of geospatial information based on the prompt template and the encoded spatial relationship text description to generate natural language text;
[0043] The merging module 106 is used to merge the natural language text with the structured triple of geospatial information to obtain a geospatial information sample.
[0044] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0045] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0046] The embodiment of the present invention also provides a terminal device, such as Figure 3 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the above-mentioned method for generating geospatial information samples when executing the computer program D102.
[0047] The terminal device D10 may be a computing device such as a desktop computer, a notebook, a PDA, a server, a server cluster, a cloud server, etc. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art will appreciate that Figure 3 This is only an example of the terminal device D10 and does not constitute a limitation on the terminal device D10. The terminal device D10 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0048] The processor D100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSp), application-specific integrated circuits (ASIC), field-programmable gate arrays (FpGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0049] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0050] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0051] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0052] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a geographic spatial information sample is implemented.
[0053] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0054] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for generating a geographic space information sample, characterized in that: include: Step 1, obtaining map service data of a study area, and determining geographic elements within the study area based on the map service data; Step 2, calculating the spatial relationship between the geographic elements based on the map service data to obtain a structured triple of geographic spatial information; Step 3, analyzing the acquired geographic spatial information text set corresponding to the study area to form a text description structure, and clustering the text description structure to obtain multiple clustering results; Step 4, identifying common features of the text description structure in multiple clustering results in combination with the structural features of the geospatial information text set, and analyzing each clustering result based on the common features to obtain a spatial relationship text description form; Step 5, encoding the spatial relationship text description form, and inputting the geospatial information structured triples and the encoded spatial relationship text description into a large language model, wherein the large language model expands the geospatial information structured triples based on a prompt template and the encoded spatial relationship text description to generate a natural language text; Step 6: Merge the natural language text with the structured triple of geospatial information to obtain a geospatial information sample.
2. The method for generating geographic space information samples according to claim 1, characterized in that: Before determining the geographic elements in the study area based on the map service data, the method further includes: The map service data is format-converted and parsed, and abnormal data is removed to obtain pre-processed map service data, wherein the abnormal data includes human error standard data and missing data.
3. The method for generating geographic space information samples according to claim 2, characterized in that: The spatial relationship between the geographic elements is calculated based on the map service data to obtain a structured triple of geographic spatial information, including: Construct a spatial index for all geographic elements in the study area through R-tree; Randomly select a geographical element from all the geographical elements in the study area as a target geographical element, and obtain neighbor elements whose distance to the target geographical element is less than a preset threshold through the spatial index to form a nearest neighbor candidate set; Randomly select a number of neighbor elements from the nearest neighbor candidate set, and calculate the spatial relationship between each neighbor element and the target geographic element, to obtain a calculation result in the form of a triple; Repeatedly select the target geographic element, and calculate the spatial relationship between the target geographic element selected each time and each neighbor element, and obtain the calculation result in the form of a triple; All calculation results in triple form that meet the saving conditions are saved to obtain structured triples of geospatial information.
4. The method for generating geographic space information samples according to claim 3, characterized in that: Before saving all calculation results in triple form that meet the saving conditions, including: Count the spatial relationship categories in the calculation results of all triples; The saving conditions are set according to the spatial relationship category: ; in, Indicates The spatial relationship of the categories, Indicates The target weight ratio, Indicates the number of categories.
5. The method for generating geographic space information samples according to claim 4, characterized in that: The obtained geospatial information text set corresponding to the study area is analyzed to form a text description structure, including: According to the data characteristics of the geospatial information field, a set of geospatial information texts corresponding to the research area is obtained on public information websites; Performing part-of-speech tagging on the geospatial information text set by using a Qwen-max model to identify geographic entities, action descriptions, and modification information in the geospatial information text set; The grammatical structure of the sentences in the geospatial information text set after part-of-speech tagging is parsed through the dependency tree to form a text description structure.
6. The method for generating geographic space information samples according to claim 5, characterized in that: Clustering the text description structure to obtain multiple clustering results, including: Vectorize the text description structure using the Sentence-BERT model to obtain a high-dimensional text description structure; The high-dimensional text description structure is density clustered using a DBSCAN model to obtain multiple clustering results, each of which contains a group of similar high-dimensional text description structures.
7. The method for generating geographic space information samples according to claim 6, characterized in that: The large language model expands the structured triples of the geospatial information based on the prompt template and the encoded spatial relationship text description to generate a natural language text expression: ; in, represents natural language text, represents a large language model, Represents a structured triple of geospatial information, They all represent geographical elements. Representing geographic features With geographical elements The spatial relationship between Represents the encoded text description of the spatial relationship.
8. A device for generating geographic space information samples, characterized in that: include: An acquisition module, used for acquiring map service data of a study area, and determining geographic elements within the study area based on the map service data; A calculation module, used for calculating the spatial relationship between the geographic elements based on the map service data to obtain a structured triple of geographic spatial information; A first analysis module is used to analyze the acquired geographic spatial information text set corresponding to the study area to form a text description structure, and cluster the text description structure to obtain multiple clustering results; A second analysis module is used to identify common features of the text description structure in multiple clustering results in combination with the structural features of the geographic spatial information text set, and analyze each clustering result based on the common features to obtain a spatial relationship text description form; An expansion module is used to encode the spatial relationship text description form, and input the geospatial information structured triples and the encoded spatial relationship text description into a large language model, and the large language model expands the geospatial information structured triples based on a prompt template and the encoded spatial relationship text description to generate a natural language text; The merging module is used to merge the natural language text with the structured triple of geospatial information to obtain a geospatial information sample.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for generating a geospatial information sample according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a geospatial information sample according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Public opinion verification method based on knowledge reasoning technology
CN113220973A
Geographic entity space-time knowledge graph ontology library construction method
CN115269751A
Spatial scene spatial relationship natural language description generation method
CN116049501A
Construction method and device of description text and space scene sample set and storage medium
CN116821692A
Method and device for acquiring key information of scientific and technical literature in earth environment field
CN118194995A