Geographic knowledge scene-based retrieval enhancement generation method and retrieval device thereof

By constructing a multi-dimensional classifier for geographical problems and corresponding search evaluators, combined with the GeoPrompt prompt word template, the semantic understanding and similarity calculation accuracy of geographic information retrieval in the existing technology is solved, the search quality and accuracy are improved, and the answering ability of large language models in the field of geography is enhanced.

CN120011498APending Publication Date: 2025-05-16NANJING NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411965891.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing search enhancement generation methods process geographic information, it is difficult to accurately understand professional terms and specific geographical nouns in geographic texts, resulting in the impact of semantic understanding and similarity calculations, affecting the search quality and LLM's ability to answer geographical questions.

Method used

By constructing a multi-dimensional classifier for geographical problems, geographical problems are divided into seven types: geographical semantics, spatial location, geometric morphology, attribute characteristics, feature relationships, evolutionary processes and mechanisms of action, and a corresponding search evaluator is built to obtain search documents for different types of problems, combine geographical problems and search documents, and build a GeoPrompt prompt word template to guide model reasoning.

Benefits of technology

It realizes the correlation evaluation and sorting of geographical problems and search documents, improves the search quality and accuracy, provides users with better question-and-answer results, and enhances the answering ability of large language models in the field of geography.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011498A_ABST
    Figure CN120011498A_ABST
Patent Text Reader

Abstract

The invention discloses a geographic knowledge scenarized retrieval enhancement generation method and a retrieval device thereof, and the method comprises the following steps: 1, constructing a geographic problem multi-dimensional classifier, and dividing geographic problems into seven types: geographic semantics, spatial positions, geometric morphology, attributive characters, element relationships, evolution processes and action mechanisms; 2, constructing a retrieval evaluator for geographic semantics, spatial positions, geometrical shapes, attribute features, element relationships, evolution processes and action mechanisms; 3, a simple retrieval strategy is adopted for problems of geographic semantics, spatial positions, geometrical shapes, attribute features and element relation types; 4, adopting a composite retrieval strategy for the problems of the evolution process and the action mechanism type; and 5, combining the geographic question and the retrieval document, constructing a GeoPrompt cue word template to guide model reasoning, and obtaining a question answer. Unification of efficiency and precision is achieved, retrieval quality is improved, and better question and answer results can be provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geographic information processing technology, and in particular to a scenario-based retrieval enhancement generation method for geographic knowledge and a retrieval device thereof. Background Art

[0002] The current retrieval enhancement generation method mainly searches through cosine similarity, and the semantic understanding of geographic entities, the interaction mechanism between geographic entities, and the information support of the evolution process of geographic entities are weak, and a lot of human intervention is required. The main reason is that on the one hand, a large number of professional terms and specific geographic nouns contained in geographic texts are not fully considered. Since these terms and nouns rarely appear in ordinary corpora, the vector model used to vectorize the text may not accurately capture the semantics of these terms, thereby affecting the accuracy of similarity calculation. On the other hand, a large number of unstructured geographic texts contain multi-dimensional knowledge (time, space, geometric morphology, semantic description, etc.), and it is difficult for existing methods to accurately identify these different aspects of knowledge. The above reasons make the existing retrieval methods have low retrieval quality for geographic knowledge, affecting the ability of LLM to answer geographic questions. Summary of the invention

[0003] In order to solve the problems existing in the prior art, the present invention provides a scenario-based retrieval enhancement generation method for geographic knowledge and a retrieval device thereof. By constructing a multidimensional classifier for geographic questions, geographic questions are divided into seven types: geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process and mechanism of action, and a corresponding retrieval evaluator is constructed to obtain retrieval documents of different types of questions. The geographic questions and retrieval documents are then combined, and a GeoPrompt prompt word template is constructed to guide model reasoning, thereby achieving the unity of efficiency and accuracy, and improving the retrieval quality and accuracy, and being able to provide users with better question and answer results.

[0004] In order to solve the above technical problems, the present invention provides a scenario-based retrieval enhancement generation method for geographic knowledge, comprising the following steps:

[0005] Step 1: Construct a multi-dimensional classifier for geographic issues and classify geographic issues into seven types: geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process, and action mechanism;

[0006] Step 2: Construct a retrieval evaluator based on geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process and mechanism of action;

[0007] Step 3: Use simple search strategies for questions about geographic semantics, spatial location, geometric morphology, attribute characteristics, and feature relationship types;

[0008] Step 4: Use a composite search strategy for questions about evolutionary processes and action mechanisms;

[0009] Step 5: Combine geographic questions and search documents, build a GeoPrompt prompt word template to guide model reasoning, and obtain the answer to the question.

[0010] Furthermore, the specific steps of constructing a multi-dimensional classifier for geographic issues in step 1 are as follows:

[0011] Step 1.1: Preprocess the geographic text corpus;

[0012] As a preferred embodiment of the present invention, the step 1.1 performs data preprocessing on the geographic text corpus, including:

[0013] Selectively retain text corpus: only retain lines ending with a period, exclamation mark, question mark, and closing quotation mark, which helps ensure the integrity of the text;

[0014] Sentence length and number of sentences: Only materials containing more than 5 sentences are retained. For each material, only lines with a sentence length greater than 3 are retained. This helps filter out short and meaningless lines.

[0015] Remove specific content: remove JavaScript warnings and placeholder text "loremipsum" included in many web pages, and remove paragraphs of text containing curly braces, which are common in many programming languages ​​but not in natural language text;

[0016] Deduplication: remove text and paragraphs that appear multiple times in the data set;

[0017] Format standardization: For non-editable formats such as PDF and images, use OCR technology to identify the image content and save it in txt format;

[0018] Filter non-Chinese materials: Use langdetect to filter text materials that are not recognized as Chinese.

[0019] Step 1.2: Divide geographic problems into seven types: geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process and mechanism of action; use a multi-agent method to extract geographic knowledge triples from text corpus, determine the geographic dimension characteristics of triples, and construct a multi-dimensional labeled geographic problem dataset;

[0020] Step 1.3: Based on the pre-trained model, the multi-label geographic question dataset is used to fine-tune and construct a multi-label classifier.

[0021] As a preferred embodiment of the present invention, the step 1.3 constructs a multi-label classifier, including:

[0022] Set the learning rate lr to 2e-5, which determines the pace at which the optimizer adjusts the model parameters at each update;

[0023] Set the batch size batchSize to 128 to control the number of samples used in each gradient update, which helps the geographic "seven-dimensional" classifier improve training efficiency and better capture language features of different dimensions for multi-label classification;

[0024] Set the maximum sequence length maxLen to 128. Most of the questions asked by users in geographic knowledge quizzes are within 128 characters. This controls the maximum processing length of the input text and ensures the consistency of the input.

[0025] The hidden layer dimension (hiddenSize=768) is the representation dimension of each token in the model. The default value of BERT is 768, which also determines the capacity and representation ability of the model.

[0026] Set the number of training rounds epochs to 10 to control the number of iterations of the entire data set;

[0027] The dropout rate is set to 0.1 to prevent overfitting and to enhance the generalization ability of the model by randomly dropping some neurons.

[0028] Furthermore, the specific steps of constructing a multi-dimensional labeled geographic question dataset in step 1.2 are as follows:

[0029] Step 1.2.1, extract triples: construct a fact extraction agent through a multi-agent method to extract basic facts about geographic entities from geographic texts, construct a geographic element extraction agent to extract the involved geographic elements from basic fact texts, and construct an element relationship extraction agent to extract the relationship between elements from basic fact texts;

[0030] Step 1.2.2, determine the type: construct a type judgment agent to determine the expression dimension of the relationship in the triple, and use the seven attributes of geographic semantics, spatial location, geometric form, attribute characteristics, element relationship, evolution process and mechanism of action to represent the seven dimensions of geography, and return a Boolean value result to determine whether the question includes the description of this dimension. Due to the inevitable illusion problem of large models, the results generated by the model sometimes do not meet the requirements. Therefore, it is necessary to filter and eliminate data entries that do not meet the format requirements according to the settings of the seven-dimensional class, and at the same time eliminate entries whose return results of all dimensions are false, which usually means that the question raised does not belong to the scope of geography;

[0031] Step 1.2.3, Generate geographic question dataset: Construct a question generation agent to use a multi-step reasoning method to generate a geographic question dataset based on the four dimensional standards of clarity, pertinence, relevance and expertise;

[0032] In terms of the system settings of the multi-agent system, Qwen1.5-110B-Chat is used as the evaluator of the generated results, the embedding vector model is text2vec-base-chinese, and the generated results are saved in JSONL format files.

[0033] The multi-label geographic question dataset D category It is expressed as:

[0034] D category ={(T (1) ,L (1) ),(T (2) ,L (2) ),...,(T (N) ,L (N) )}

[0035] Where, T (n) Represents the problem description of the nth data point; Represents the 7-dimensional labeling of the nth data point, where L1 represents the geographic semantic dimension, L2 represents the spatial position dimension, L3 represents the geometric morphology dimension, L4 represents the attribute feature dimension, L5 represents the element relationship dimension, L6 represents the evolution process dimension, and L7 represents the action mechanism dimension.

[0036] Furthermore, the specific meanings of the seven geographic dimensions in step 1.2.2 are as follows:

[0037] Geographic semantics: the term explanation and classification system involved in representing geographic entities;

[0038] Spatial location: represents the topological relationship, direction relationship and distance relationship involved in geographic entities;

[0039] Geometric form: represents the morphological and structural features involved in representing geographic entities;

[0040] Attribute characteristics: represent the natural attributes and social attributes involved in geographical entities;

[0041] Element relationships: represent the causal and interactive relationships involved in geographic entities;

[0042] Evolution process: represents the time points, time segments and life cycles of changes involved in geographical entities;

[0043] Mechanism of action: represents the physical mechanism, chemical mechanism, biological mechanism, humanistic, social and economic mechanism involved in geographical entities and the transmission and transformation mechanism of geographical information.

[0044] Furthermore, the specific steps of constructing the retrieval evaluator in step 2 are as follows:

[0045] Step 2.1: Use a multi-agent approach to evaluate the semantic relevance between geographic questions and geographic facts, and construct a seven-dimensional geographic question-answer dataset;

[0046] Step 2.2: Based on the pre-trained model, the geographic seven-dimensional question-answering dataset is fine-tuned to construct a retrieval evaluator.

[0047] As a preferred embodiment of the present invention, the step 2.2 constructs a retrieval evaluator, including:

[0048] Set the learning rate to 2×10 -5 , helping the model gradually learn and capture the complex semantic relationships in geographic texts about geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationships, evolutionary processes, and mechanisms of action;

[0049] Set the batch size to 128 to balance memory usage and training speed, and effectively learn the semantic information of the seven dimensions of geography;

[0050] The maximum sequence length is set to 512 to ensure that the key information after the user's question and related knowledge fragments are covered, which helps to understand the connection between geographic texts;

[0051] The hidden layer size is set to 768 to provide sufficient expressive power to capture complex semantics.

[0052] Furthermore, the specific steps of constructing the geographic seven-dimensional question-answer pair dataset in step 2.1 are as follows:

[0053] Geographic seven-dimensional question-answering dataset D evaluator for:

[0054] D evaluator ={(T (1) ,A (1) ),(T (2) ,A (2) ),...,(T (N) ,A (N) )}

[0055] Where, T (n) A represents the problem description of the nth data point; (n) Represents the answer description of the nth data point.

[0056] Furthermore, the specific steps of adopting a simple search strategy for the questions of geographic semantics, spatial location, geometric form, attribute characteristics and element relationship type in step 3 are as follows:

[0057] Step 3.1: Formal expression of simple geographical questions: The question-answering process of simple questions can be expressed as:

[0058] Qsimple (q, DB) = {Y top-k (q,D)|D∈DB}

[0059] result = g llm (q,Q simple (q,DB))

[0060] In the formula, q represents the geographic question raised by the user; DB represents the knowledge base including all geographic knowledge; D represents geographic knowledge, including the five dimensions of geographic semantics, spatial location, geometric morphology, attribute characteristics and element relationship, expressed as a five-tuple<S,P,F,A,R> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents a simple question retrieval method; formula g llm Indicates the use of a large language model LLM for reasoning and generating answers;

[0061] Step 3.2, simple question retrieval strategy: The retrieval method for simple questions combines the cosine similarity score between the geographic question and the retrieval document in the vector space and the relevance score of the retrieval evaluator, and re-ranks the top-k data as the retrieval results after weighted average. The retrieval method for simple questions is:

[0062] q,D=W emb (q i ,D i )+P i +S i

[0063]

[0064] y(q,D)=w X1 x1+w X2 x2

[0065] Y top-k (q,D)={y i '|i=1,2,...,k}

[0066] Where q represents the geographic question vector raised by the user; D represents the geographic knowledge result vector obtained from the vector library; W emb represents the word embedding matrix; q i ,D i Represents word embedding, the input sequence is X = [x1, x2, ..., x n ]; position embedding P = [p1,p2,...,p n], indicating the position of each word in the sequence; paragraph embedding S is used to distinguish different sentences; X1 represents the cosine similarity between the question vector and the retrieved document vector; X2 represents the relevance score between the question vector and the retrieved document vector predicted by the retrieval evaluator; y(q,D) represents the comprehensive semantic similarity between the question vector and the retrieved document vector; Represent the weights of cosine similarity and retrieval evaluator relevance score respectively; Y top-k Indicates the top-k search results.

[0067] Furthermore, the specific steps of using a composite search strategy for questions about the evolution process and mechanism of action type in step 4 are as follows:

[0068] Step 4.1: Formal expression of complex geographical questions: The question-answering process of complex questions is expressed as follows:

[0069] Q composite (q, DB) = {Y top-k (q,D)|D∈DB}

[0070] result = g llm (q,Q composite (q,DB))

[0071] In the formula, q represents the geographical question raised by the user; DB represents the knowledge base including all geographical knowledge; D represents geographical knowledge, including the content of two dimensions: mechanism and evolution process, expressed as a binary<E,M> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents the compound question retrieval method; formula g llm Indicates the use of a large language model LLM for reasoning and generating answers;

[0072] Step 4.2, complex question retrieval strategy: The retrieval method for complex questions combines adaptive query expansion, arc consistency query and local optimization, and obtains the top-k data as the retrieval results after multiple rounds of retrieval. The retrieval method for complex questions is:

[0073] For a given user question q, the multi-agent system extracts element entities E = {e1, e2, ..., e n} and relation entity R = {r1,r2,...,r n}; Based on the entity and relationship extraction results, combined with synonym expansion, a keyword matching set X = {X1, X2, ..., X n}, where the keyword X i Corresponds to a set of geographic entities or feature relationships Let C = {c(X i ,X j)} is a set of binary constraints, each constraint c(X i ,X j ) describes a reasonable relationship between two keywords; think of it as a i ×D j The mapping from 0 to {0,1} is:

[0074] c(X i ,X j ):D i ×D j →{0,1}

[0075] Where c(X i ,X j )(d i ,d j )=1 means d i and d j Satisfy the constraints; c(X i ,X j )(d i ,d j )=0 means not satisfied;

[0076] Create a queue Q containing all arcs to be checked, that is, Q = { <X i ,X j >|c(X i ,X j )∈C}; for each arc taken from the queue <X i ,X j > By updating the function revise(X i ,X j )Get the matching results:

[0077]

[0078] Where D i ′ represents the updated X i New domains;

[0079] All variables and their valid domains are combined into S output, expressed as S = {(X1, D1′), (X2, D2′), ..., (X n ,D n ′)}; construct the question-answer pair of the question and the retrieved document, and score the relevance through the retrieval evaluator, and select the top-k results as the compound question retrieval output, that is: Y top-k (q,D)={S i '|i=1,2,...,k}.

[0080] Furthermore, the specific steps of constructing the GeoPrompt prompt word template to guide model reasoning in step 5 are as follows:

[0081] GeoPrompt general template definition. A typical GeoPrompt definition is as follows:

[0082] GeoPrompt = <question type, user question, search document>

[0083] GeoPrompt provides reference information and question prompts for model reasoning, designs a multi-step thinking process to prompt the model to reason and obtain answers to questions, and controls the form of text generated by the model to improve text quality.

[0084] The GeoPrompt has the form:

[0085] You are an expert in the {discipline classification} discipline research {research field} field. Please judge the correct answer to the following {question type} question based on the given reference knowledge text. Please do not answer when you are unsure. The following is the reference knowledge text: {knowledge text}. The following is my question: {user question}.

[0086] A retrieval device for the scenario-based retrieval enhancement generation method of geographic knowledge as described above comprises an analysis module: used to automatically generate a multi-label geographic question data set based on a seven-dimensional geographic classification method using a multi-agent system, and to construct a geographic question multi-label classifier based on the data set;

[0087] Retrieval module: used to generate a geographic seven-dimensional question-answer pair dataset based on multi-label geographic questions, and build a retrieval evaluator based on the dataset;

[0088] The first retrieval module: used to establish a formal expression of simple geographical questions and build a simple question retrieval strategy;

[0089] The second search module is used to establish a formal expression of complex geographical questions and construct a search strategy for complex questions;

[0090] Reasoning module: used to combine geographic questions and retrieve documents, and build GeoPrompt prompt word templates to guide model reasoning.

[0091] The beneficial effects of the present invention are:

[0092] The present invention constructs a knowledge retrieval model for different types of geographical questions, optimizes the retrieval process of the retrieval enhancement generation method in the field of geography, realizes the relevance evaluation and sorting of geographical questions and retrieval documents, and optimizes the answering ability of the large language model in the field of geography.

[0093] The present invention realizes the automatic classification of geographical issues from seven dimensions: geographical semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process and mechanism of action by setting a geographical seven-dimensional multi-label classifier. The geographical document retrieval evaluator of the present invention scores the relevance of geographical issues and retrieval documents based on the semantic similarity of the text. The two types of retrieval modes of the present invention improve the quality of retrieval documents for different types of questions. The present invention constructs a prompt word template GeoPrompt to integrate questions and retrieval documents, guides the model to understand and generate answers to questions based on reference material reasoning, improves the accuracy of the model in answering geographical questions, and enables a small parameter model to obtain performance that is similar to that of a large parameter model. In terms of technical adaptation, the present invention provides a plug-and-play method that can be applied to different types of retrieval enhancement generation technologies and improves their performance in the field of geography. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 is a flow chart of the present invention;

[0095] Figure 2 A flowchart for extracting geographic knowledge triples and generating question-answer pairs in a specific embodiment;

[0096] Figure 3 A result graph for geographic knowledge triple extraction and question generation in a specific embodiment;

[0097] Figure 4 It is a classification result diagram of a multi-label classifier in a specific embodiment;

[0098] Figure 5 It is a simple question retrieval strategy flow chart in a specific embodiment;

[0099] Figure 6 It is a flow chart of the complex question retrieval strategy in a specific embodiment;

[0100] Figure 7 This is a result graph of GeoPrompt construction in a specific embodiment. DETAILED DESCRIPTION

[0101] Embodiment 1:

[0102] like Figure 1 As shown, this embodiment discloses a scenario-based retrieval enhancement generation method for geographic knowledge, comprising the following steps:

[0103] Step 1: Based on the seven-dimensional geographic classification method, a multi-agent system is used to automatically generate a multi-label geographic question dataset, and a geographic question multi-label classifier is constructed based on the dataset;

[0104] Step 1.1: Preprocess the geographic text corpus, including:

[0105] Selectively retain text corpus: only retain lines ending with a period, exclamation mark, question mark, and closing quotation mark, which helps ensure the integrity of the text;

[0106] Sentence length and number of sentences: Only materials containing more than 5 sentences are retained. For each material, only lines with a sentence length greater than 3 are retained. This helps filter out short and meaningless lines.

[0107] Remove specific content: remove JavaScript warnings and placeholder text "loremipsum" included in many web pages, and remove paragraphs of text containing curly braces, which are common in many programming languages ​​but not in natural language text;

[0108] Deduplication: remove text and paragraphs that appear multiple times in the data set;

[0109] Format standardization: For non-editable formats such as PDF, pictures, etc., use OCR technology to identify the image content and save it in txt format; filter non-Chinese materials: use langdetect to filter text materials that are not recognized as Chinese.

[0110] Step 1.2: Figure 2 As shown in the figure, a multi-agent method is used to extract geographical knowledge triples from text corpus, determine the geographical dimension characteristics of triples, and construct a multi-label geographical question dataset, including:

[0111] Triple extraction: A fact extraction agent is constructed through a multi-agent approach to extract basic facts about geographic entities from geographic texts. The text length is between 100 and 150 characters, including six aspects: geographic location, geographic features, climate features, natural resources, human activities, and environmental issues. A geographic element extraction agent is constructed to extract the involved geographic elements from the basic fact text, and an element relationship extraction agent is constructed to extract the relationship between elements from the basic fact text;

[0112] Type judgment: Construct a type judgment agent to judge the expression dimension of the relationship in the triple, and express the seven dimensions of geography through seven attributes: geographic semantics, spatial location, geometric form, attribute characteristics, element relationship, evolution process and mechanism of action. By returning a Boolean value result, it is judged whether the question includes the description of this dimension. Due to the inevitable illusion problem of large models, the results generated by the model sometimes do not meet the requirements. Therefore, it is necessary to filter and eliminate data entries that do not meet the format requirements according to the settings of the seven-dimensional class, and at the same time eliminate entries with false return results for all dimensions, which usually means that the question raised does not belong to the scope of geography;

[0113] Question Generation: Construct a question generation agent to generate geographic questions using a multi-step reasoning approach based on four criteria: clarity, relevance, expertise, and specificity.

[0114] In terms of multi-agent system settings, Qwen1.5-110B-Chat is used as the evaluator of the generated results, the embedding vector model is text2vec-base-chinese, and the generated results are saved in JSONL format files.

[0115] The constructed multi-label geographic question dataset D category It is expressed as:

[0116] D category ={(T (1) ,L (1) ),(T (2) ,L (2) ),...,(T (N) ,L (N) )}

[0117] Where, T (n) Represents the problem description of the nth data point; Represents the 7-dimensional labeling of the nth data point, where L1 represents the geographic semantic dimension, L2 represents the spatial position dimension, L3 represents the geometric morphology dimension, L4 represents the attribute feature dimension, L5 represents the element relationship dimension, L6 represents the evolution process dimension, and L7 represents the mechanism dimension.

[0118] The results of geographic knowledge triple extraction and question generation are as follows: Figure 3 shown.

[0119] Step 1.3: Based on the pre-trained model, fine-tune and construct a multi-label classifier using the multi-label geographic question dataset, specifically including:

[0120] Set the learning rate lr to 2e-5, which determines the pace at which the optimizer adjusts the model parameters at each update;

[0121] Set the batch size batchSize to 128 to control the number of samples used in each gradient update, which helps the geographic "seven-dimensional" classifier improve training efficiency and better capture language features of different dimensions for multi-label classification;

[0122] Set the maximum sequence length maxLen to 128. Most of the questions asked by users in geographic knowledge quizzes are within 128 characters. This controls the maximum processing length of the input text and ensures the consistency of the input.

[0123] The hidden layer dimension (hiddenSize=768) is the representation dimension of each token in the model. The default value of BERT is 768, which also determines the capacity and representation ability of the model.

[0124] Set the number of training rounds epochs to 10 to control the number of iterations of the entire data set;

[0125] The dropout rate is set to 0.1 to prevent overfitting and to enhance the generalization ability of the model by randomly dropping some neurons.

[0126] The classification results of the multi-label classifier are as follows Figure 4 shown.

[0127] Step 2: Generate a geographic seven-dimensional question-answer pair dataset based on multi-label geographic questions, and build a retrieval evaluator based on the dataset;

[0128] Step 2.1: Use a multi-agent approach to evaluate the semantic relevance between geographic questions and geographic facts, and construct a seven-dimensional geographic question-answer dataset, including:

[0129] Use step 1.2 to generate a question dataset. The answer evaluation is based on prompt engineering. Qwen1.5-110B-Chat is used as the answer evaluator. The evaluation criteria are:

[0130] Clarity: Questions should be clear and specific, avoiding vagueness or generalities;

[0131] Targeted: The question should be asked about an entity in the text entity table;

[0132] Relevance: The question should be closely related to the text content, ensuring that the answer can be found in the text;

[0133] Professionalism: Questions should be professional in the field of geomorphology and avoid questions that are too simple or common sense.

[0134] Geographic seven-dimensional question-answering dataset D evaluator for:

[0135] D evaluator ={(T (1) ,A (1) ),(T (2) ,A (2) ),...,(T (N) ,A (N) )}

[0136] Where, T (n) A represents the problem description of the nth data point; (n) Represents the answer description of the nth data point.

[0137] Step 2.2: Based on the pre-trained model, the geographic seven-dimensional question-answering dataset is fine-tuned to construct a retrieval evaluator, specifically including:

[0138] Set the learning rate to 2×10 -5 , helping the model gradually learn and capture the complex semantic relationships in geographic texts about geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationships, evolutionary processes, and mechanisms of action;

[0139] The batch size is set to 128 to balance memory usage and training speed, effectively learning the semantic information of the seven dimensions of geography;

[0140] The maximum sequence length is set to 512 to ensure that the key information after the user's question and the splicing of related knowledge fragments are covered, which helps to understand the connection between geographic texts;

[0141] The hidden layer size is 768 to provide enough expressive power to capture complex semantics.

[0142] Step 3: Establish a formal expression of simple geographic questions and construct a simple question retrieval strategy, including:

[0143] Step 3.1: Formal expression of simple geographic questions: A scenario-based retrieval and enhanced generation method for geographic knowledge is used to express the question-answering process of simple questions as follows:

[0144] Q simple (q, DB) = {Y top-k (q,D)|D∈DB}

[0145] result = g llm (q,Q simple (q,DB))

[0146] In the formula, q represents the geographic question raised by the user; DB represents the knowledge base including all geographic knowledge; D represents geographic knowledge, including the five dimensions of geographic semantics, spatial location, geometric morphology, attribute characteristics and element relationship, expressed as a five-tuple<S,P,F,A,R> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents a simple question retrieval method; formula g llm Indicates the use of a large language model (LLM) for reasoning and generating answers.

[0147] Step 3.2: Simple question search strategy: Figure 5 As shown in the figure, the retrieval method for simple questions combines the cosine similarity score of the geographic question and the retrieval document in the vector space and the relevance score of the retrieval evaluator, and re-ranks and selects the top-k data as the retrieval results after weighted average. The retrieval method is:

[0148] q,D=W emb (q i ,D i )+P i +S i

[0149]

[0150] Y top-k (q,D)={y i '|i=1,2,...,k}

[0151] Where q represents the geographic question vector raised by the user; D represents the geographic knowledge result vector obtained from the vector library; W emb represents the word embedding matrix; q i ,D i Represents word embedding, the input sequence is X = [x1, x2, ..., x n ]; position embedding P = [p1,p2,...,p n ], indicating the position of each word in the sequence; paragraph embedding S is used to distinguish different sentences; X1 represents the cosine similarity between the question vector and the retrieved document vector; X2 represents the relevance score between the question vector and the retrieved document vector predicted by the retrieval evaluator; y(q,D) represents the comprehensive semantic similarity between the question vector and the retrieved document vector; Represent the weights of cosine similarity and retrieval evaluator relevance score respectively; Y top-k Indicates the top-k search results.

[0152] Step 4: Establish a formal expression of complex geographic questions and construct a complex question retrieval strategy, including:

[0153] Step 4.1: Formal expression of complex geographic questions: A scenario-based retrieval and enhanced generation method for geographic knowledge is used to express the question-answering process of complex questions as follows:

[0154] Q composite (q, DB) = {Y top-k (q,D)|D∈DB}

[0155] result = g llm (q,Q composite (q,DB))

[0156] In the formula, q represents the geographical question raised by the user; DB represents the knowledge base including all geographical knowledge; D represents geographical knowledge, including the content of two dimensions: mechanism and evolution process, expressed as a binary<E,M> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents the compound question retrieval method; formula g llmIndicates the use of a large language model (LLM) for reasoning and generating answers.

[0157] Step 4.2: Complex question search strategy: Figure 6 As shown in the figure, the retrieval method for complex problems combines adaptive query expansion, arc consistency query and local optimization. After multiple rounds of retrieval, the top-k data are obtained as retrieval results. The retrieval method is:

[0158] For a given user question q, the multi-agent system extracts element entities E = {e1, e2, ..., e n} and relation entity R = {r1,r2,...,r n According to the entity and relationship extraction results, combined with synonym expansion, a keyword matching set X = {X1, X2, ..., X n}, where the keyword X i Corresponding to a set of geographic entities or feature relationships D i ={d i1 ,d i2 ,...,d imn}. Let C = {c(X i ,X j )} is a set of binary constraints, each constraint c(X i ,X j ) describes a reasonable relationship between two keywords. It can be regarded as i ×D j The mapping from {0,1} is:

[0159] c(X i ,X j ):D i ×D j →{0,1}

[0160] Where c(X i ,X j )(d i ,d j )=1 means d i and d j Satisfy the constraints; c(X i ,X j )(d i ,d j )=0 means not satisfied.

[0161] Create a queue Q containing all arcs to be checked, that is, Q = { <X i ,X j >|c(X i ,X j )∈C}. For each arc taken from the queue <X i ,Xj > By updating the function revise(X i ,X j )Get matching results

[0162]

[0163] Where D i ′ represents the updated X i New domain.

[0164] All variables and their valid domains are combined into S output, expressed as S = {(X1, D1′), (X2, D2′), ..., (X n ,D n ′)}. The question-answer pair of the question and the retrieved document is scored for relevance by the retrieval evaluator, and the top-k results are selected as the compound question retrieval output, that is: Y top-k (q,D)={S i '|i=1,2,...,k}.

[0165] Step 5: Combine geographic questions and search documents to build a GeoPrompt prompt word template to guide model reasoning, including:

[0166] GeoPrompt general template definition. A typical GeoPrompt definition is as follows:

[0167] GeoPrompt = <question type, user question, search document>

[0168] GeoPrompt provides reference information and question prompts for model reasoning, designs a multi-step thinking process to prompt the model to reason and obtain answers to questions, and controls the form of text generated by the model to improve text quality.

[0169] like Figure 7 As shown, the GeoPrompt is in the form of:

[0170] You are an expert in {discipline classification} subject research {research field}. Please judge the correct answer to the following {question type} question based on the given reference knowledge text. Please do not answer when you are unsure. The following is the reference knowledge text: {knowledge text}. The following is my question: {user question}

[0171] Embodiment 2:

[0172] The following combination Figure 2 , Figure 3 and Figure 4 This paper gives an example of building a multi-label geographic question dataset and a geographic seven-dimensional question-answer dataset. The specific implementation process is as follows:

[0173] In order to clearly describe the implementation process of the technical solution, the description paragraph of Shanxi Province in the "Encyclopedia of Chinese Geography" is used to illustrate the results of triple extraction and geographic question dataset generation, providing a general example for the implementation process. The processed text is as follows:

[0174] Shanxi Province is called "Jin" for short. It is located in the north of China, west of the North China Plain, and east of the middle reaches of the Yellow River. It covers an area of ​​about 160,000 square kilometers. The population is 34.74 million (2010). It is named after being west of Taihang Mountain. The provincial people's government is located in the ancient Bingzhou area of ​​Taiyuan City. It was the Jin State in the Spring and Autumn Period. It was the Zhao territory in the Warring States Period, and also belonged to Han and Wei. The Qin Dynasty established Taiyuan, Hedong, Shangdang, Yanmen, Yunzhong, Dai and other counties. It was Bingzhou in the Han Dynasty. It was Hedong Road in the Tang Dynasty. It was Hedong Road and Xijing Road (belonging to Liao) in the Song Dynasty. It belonged to the Shanxi Xingzhongshu Province in the Yuan Dynasty, and the Hedong Shanxi Road Suzheng Lianfangsi and Xuanweishisi were established. The Ming Dynasty established the Shanxi Provincial Administration. It was Shanxi Province in the Qing Dynasty. The northern part of the province belonged to Chahar Province from 1949 to 1952. It is 800 to 1,500 meters above sea level, called the Shanxi Plateau, which is part of the Loess Plateau. There are Luliang, Taihang, Wutai, Hengshan, Taiyue, Zhongtiao and other mountains. There are basins such as Datong, Zhezhou, Taiyuan, Linfen and Yuncheng in the territory, which are all major agricultural areas. The Yellow River circulates in the west and south, with rapids such as Hukou, Baobu and Longmen, rich in hydropower resources, and tributaries such as Fen, Qin and other rivers. Sanggan, Putuo and Zhang rivers are the upper sources of each river. It belongs to the warm temperate semi-humid semi-arid continental monsoon climate, with an average annual temperature of 4-14℃ and an average annual precipitation of 400-600 mm, decreasing from southeast to northwest. Agriculture mainly produces wheat, sorghum, corn, cotton, etc. in the south of Yanmen Pass, and millet, oats and sesame in the north. Animal husbandry is well developed. It is rich in mineral resources, and its coal reserves rank first in the country.

[0175] First, use the basic geographical facts contained in the text Figure 2 The fact extraction agent is used to extract, and the geographic feature extraction agent and feature relationship extraction agent are used to obtain the triples of geographic facts. The geographic entity and triple extraction results of the above text are shown in Table 1:

[0176] Table 1 Geographic entity and triple extraction results of text

[0177]

[0178]

[0179] According to the triple extraction results, the type judgment agent is used to judge the geographical seven-dimensional classification of the triple, and the question generation agent is used to generate geographical questions. Some question generation results and dimension judgment results are shown in Table 2:

[0180] Table 2: Some question generation results and dimension judgment results

[0181]

[0182]

[0183] According to the geographic questions generated in step 1.2 as the questions of the question-answer pair, the basic facts extracted from the paragraph are used as the answers to the question-answer pair to construct the corresponding dimension question-answer pairs. The generation results of some geographic question-answer pairs are shown in Table 3:

[0184] Table 3. Results of generating some geographic question-answer pairs

[0185]

[0186] Figure 4 Results of a partial geographic question-answer pair dataset used to train the geographic semantic retrieval evaluator are given.

[0187] Embodiment 3:

[0188] The following combination Figure 5 A specific retrieval process for a geomorphology judgment question is given. The retrieval method for simple questions combines the cosine similarity score of the geographical question and the retrieval document in the vector space and the relevance score of the retrieval evaluator. After weighted averaging, the top-k data are re-ranked and selected as the retrieval results.

[0189] In order to clearly describe the implementation process of the technical solution, a simple question about attribute characteristics, "Is the longitudinal profile of the periglacial valley very steep?" is taken as an example to provide a general example for the implementation process. The cosine similarity results of the retrieval documents and the retrieval evaluator weight evaluation results for this question are shown in Table 4:

[0190] Table 4. Cosine similarity results of retrieved documents and retrieval evaluator re-evaluation results

[0191]

[0192]

[0193] Embodiment 4:

[0194] The following combination Figure 6 A specific retrieval process for a geomorphology judgment question is given. The retrieval method for complex questions combines adaptive query expansion, arc consistency query and local optimization. After multiple rounds of retrieval, the top-k data are obtained as retrieval results.

[0195] In order to clearly describe the implementation process of the technical solution, the compound question "Is the formation of tectonic lake basins independent of external forces such as weathering and erosion of the crust, and mainly caused by internal forces?" is taken as an example to provide a general example for the implementation process. The binary constraints for the retrieval documents for this question are as follows:

[0196] {'query':{'bool':{'should':[{'match':{'text':'Tectonic lake basin'}},{'match':{'text':'Crust'}},

[0197] {'match':{'text':'Weathering'}},{'match':{'text':'Erosion'}},{'match':{'text':'Internal Force'}}],

[0198] 'minimum_should_match':1,'filter':{'bool':{'should':[{'bool':{'must':[{'term':

[0199] {'entities.entity':'Tectonic lake basin'}},{'term':{'entities.relation':'Formation'}},{'term':{'entities.entity':'Internal forces'}}]}}],'minimum_should_match':1}}}}}

[0200] The updated matching scores of the retrieved documents and the retrieval evaluator weight evaluation results are shown in Table 5:

[0201] Table 5. Retrieval document update matching scores and retrieval evaluator re-evaluation results

[0202]

[0203]

[0204]

[0205] Embodiment 5:

[0206] This embodiment also provides a retrieval device for a scenario-based retrieval enhancement generation method of geographic knowledge, which is characterized by comprising an analysis module: used to automatically generate a multi-label geographic question data set based on a seven-dimensional geographic classification method using a multi-agent system, and construct a geographic question multi-label classifier based on the data set;

[0207] Retrieval module: used to generate a geographic seven-dimensional question-answer pair dataset based on multi-label geographic questions, and build a retrieval evaluator based on the dataset;

[0208] The first retrieval module: used to establish a formal expression of simple geographical questions and build a simple question retrieval strategy;

[0209] The second search module is used to establish a formal expression of complex geographical questions and construct a search strategy for complex questions;

[0210] Reasoning module: used to combine geographic questions and retrieve documents, and build GeoPrompt prompt word templates to guide model reasoning.

Claims

1. A scenario-based retrieval enhancement generation method for geographic knowledge, characterized in that: The following steps are involved: Step 1: Construct a multi-dimensional classifier for geographic issues and classify geographic issues into seven types: geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process, and action mechanism; Step 2: Construct a retrieval evaluator based on geographic semantics, spatial location, geometric morphology, attribute characteristics, element relationship, evolution process and mechanism of action; Step 3: Use simple search strategies for questions about geographic semantics, spatial location, geometric morphology, attribute characteristics, and feature relationship types; Step 4: Use a composite search strategy for questions about evolutionary processes and action mechanisms; Step 5: Combine geographic questions and search documents, build a GeoPrompt prompt word template to guide model reasoning, and obtain the answer to the question.

2. A scenario-based retrieval enhancement generation method for geographic knowledge according to claim 1, characterized in that: The specific steps of constructing a multi-dimensional classifier for geographic issues in step 1 are as follows: Step 1.1: Preprocess the geographic text corpus; Step 1.2: Divide geographic problems into seven types: geographic semantics, spatial location, geometric form, attribute characteristics, element relationship, evolution process and mechanism; A multi-agent approach is used to extract geographic knowledge triples from text corpora, determine the geographic dimension features of triples, and construct a multi-dimensional labeled geographic question dataset. Step 1.3: Based on the pre-trained model, the multi-label geographic question dataset is used to fine-tune and construct a multi-label classifier.

3. A scenario-based retrieval enhancement generation method for geographic knowledge according to claim 2, characterized in that: The specific steps for constructing a multi-dimensional labeled geographic question dataset in step 1.2 are as follows: Step 1.2.1, extract triples: construct a fact extraction agent through a multi-agent method to extract basic facts about geographic entities from geographic texts, construct a geographic element extraction agent to extract the involved geographic elements from basic fact texts, and construct an element relationship extraction agent to extract the relationship between elements from basic fact texts; Step 1.2.2, determine the type: construct a type determination agent to determine the expression dimension of the relationship in the triple, and use the seven attributes of geographic semantics, spatial location, geometric form, attribute characteristics, element relationship, evolution process and mechanism to represent the seven dimensions of geography. By returning the Boolean value result, determine whether the problem includes the description of the dimension, eliminate the data entries that do not meet the format requirements, and eliminate the entries whose return results of all dimensions are false; Step 1.2.3, Generate geographic question dataset: Construct a question generation agent to use a multi-step reasoning method to generate a geographic question dataset based on the four dimensional standards of clarity, pertinence, relevance and expertise; The multi-label geographic question dataset D category It is expressed as: D category ={(T (1) ,L (1) ),(T (2) ,L (2) ),...,(T (N) ,L (N) )} Where, T (n) Represents the problem description of the nth data point; Represents the 7-dimensional labeling of the nth data point, where L1 represents the geographic semantic dimension, L2 represents the spatial position dimension, L3 represents the geometric morphology dimension, L4 represents the attribute feature dimension, L5 represents the element relationship dimension, L6 represents the evolution process dimension, and L7 represents the action mechanism dimension.

4. A method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 3, characterized in that: The specific meanings of the seven geographic dimensions in step 1.2.2 are as follows: Geographic semantics: the term explanation and classification system involved in representing geographic entities; Spatial location: represents the topological relationship, direction relationship and distance relationship involved in geographic entities; Geometric form: represents the morphological and structural features involved in representing geographic entities; Attribute characteristics: represent the natural attributes and social attributes involved in geographical entities; Element relationships: represent the causal and interactive relationships involved in geographic entities; Evolution process: represents the time points, time segments and life cycles of changes involved in geographical entities; Mechanism of action: represents the physical mechanism, chemical mechanism, biological mechanism, humanistic, social and economic mechanism involved in geographical entities and the transmission and transformation mechanism of geographical information.

5. The method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 1, characterized in that: The specific steps of constructing the retrieval evaluator in step 2 are as follows: Step 2.1: Use a multi-agent approach to evaluate the semantic relevance between geographic questions and geographic facts, and construct a seven-dimensional geographic question-answer dataset; Step 2.2: Based on the pre-trained model, the geographic seven-dimensional question-answering dataset is fine-tuned to construct a retrieval evaluator.

6. A method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 5, characterized in that: The specific steps of constructing the geographic seven-dimensional question-answer pair dataset in step 2.1 are as follows: Geographic seven-dimensional question-answering dataset D evaluator for: D evaluator ={(T (1) ,A (1) ),(T (2) ,A (2) ),...,(T (N) ,A (N) )} Where, T (n) A represents the problem description of the nth data point; (n) Represents the answer description of the nth data point.

7. The method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 1, characterized in that: The specific steps of using a simple search strategy for questions about geographic semantics, spatial location, geometric form, attribute characteristics, and element relationship types in step 3 are as follows: Step 3.1: Formal expression of simple geographical questions: The question-answering process of simple questions can be expressed as: Q simple (q,DB)={Y top-k (q,D)∣D∈DB} result=g llm (q,Q simple (q,DB)) In the formula, q represents the geographic question raised by the user; DB represents the knowledge base including all geographic knowledge; D represents geographic knowledge, including the five dimensions of geographic semantics, spatial location, geometric morphology, attribute characteristics and element relationship, expressed as a five-tuple<S,P,F,A,R> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents a simple question retrieval method; formula g llm Indicates the use of a large language model LLM for reasoning and generating answers; Step 3.2, simple question retrieval strategy: The retrieval method for simple questions combines the cosine similarity score between the geographic question and the retrieval document in the vector space and the relevance score of the retrieval evaluator, and re-ranks the weighted average to select the top-k data as the retrieval result. The retrieval method for simple questions is: q,D=W emb (q i ,D i )+P i +S i Y top-k (q,D)={y i '∣i=1,2,...,k} Where q represents the geographic question vector raised by the user; D represents the geographic knowledge result vector obtained from the vector library; W emb represents the word embedding matrix; q i ,D i Represents word embedding, the input sequence is X = [x1, x2, ..., x n ]; position embedding P = [p1,p2,...,p n ], indicating the position of each word in the sequence; paragraph embedding S is used to distinguish different sentences; X1 represents the cosine similarity between the question vector and the retrieved document vector; X2 represents the relevance score between the question vector and the retrieved document vector predicted by the retrieval evaluator; y(q,D) represents the comprehensive semantic similarity between the question vector and the retrieved document vector; Represent the weights of cosine similarity and retrieval evaluator relevance score respectively; Y top-k Indicates the top-k search results.

8. The method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 1, characterized in that: The specific steps of using the composite search strategy for the evolution process and mechanism type questions in step 4 are as follows: Step 4.1: Formal expression of complex geographical questions: The question-answering process of complex questions is expressed as follows: Q composite (q,DB)={Y top-k (q,D)∣D∈DB} result=g llm (q,Q composite (q,DB)) In the formula, q represents the geographical question raised by the user; DB represents the knowledge base including all geographical knowledge; D represents geographical knowledge, including the content of two dimensions: mechanism and evolution process, expressed as a binary<E,M> ; top-k represents the number of retrieved documents; Formula Y top-k (q,D) represents the compound question retrieval method; formula g llm Indicates the use of a large language model LLM for reasoning and generating answers; Step 4.2, complex question retrieval strategy: The retrieval method for complex questions combines adaptive query expansion, arc consistency query and local optimization, and obtains the top-k data as the retrieval results after multiple rounds of retrieval. The retrieval method for complex questions is: For a given user question q, the multi-agent system extracts element entities E = {e1, e2, ..., e n } and relation entity R = {r1,r2,...,r n }; Based on the entity and relationship extraction results, combined with synonym expansion, a keyword matching set X = {X1, X2, ..., X n }, where the keyword X i Corresponds to a set of geographic entities or feature relationships Let C = {c(X i ,X j )} is a set of binary constraints, each constraint c(X i ,X j ) describes a reasonable relationship between two keywords; think of it as a i ×D j The mapping from {0,1} is: c(X i ,X j ):D i ×D j →{0,1} Where c(X i ,X j )(d i ,d j )=1 means d i and d j Satisfy the constraints; c(X i ,X j )(d i ,d j )=0 means not satisfied; Create a queue Q containing all arcs to be checked, that is, Q = { <X i ,X j >|c(X i ,X j )∈C}; for each arc taken from the queue <X i ,X j > By updating the function revise(X i ,X j )Get the matching results: Where D i ′ represents the updated X i New domains; All variables and their valid domains are combined into S output, expressed as S = {(X1, D1′), (X2, D2′), ..., (X n ,D n ′)}; construct the question-answer pair of the question and the retrieved document, score the relevance through the retrieval evaluator, and select the top-k results as the compound question retrieval output, that is: Y top-k (q,D)={S i '|i=1,2,...,k}.

9. The method for enhancing the generation of scenario-based retrieval of geographic knowledge according to claim 1, characterized in that: The specific steps of constructing the GeoPrompt prompt word template to guide model reasoning in step 5 are as follows: GeoPrompt generic template definition: GeoPrompt = <question type, user question, search document> GeoPrompt provides reference information and question prompts for model reasoning, designs a multi-step thinking process to prompt the model to reason and obtain answers to questions, and controls the form of text generated by the model to improve text quality.

10. A retrieval device according to any one of claims 1 to 9 for the enhanced generation method of scenario-based retrieval of geographic knowledge, characterized in that: include: Analysis module: used to automatically generate a multi-label geographic problem dataset based on the seven-dimensional geographic classification method using a multi-agent system, and build a multi-label classifier for geographic problems based on the dataset; Retrieval module: used to generate a geographic seven-dimensional question-answer pair dataset based on multi-label geographic questions, and build a retrieval evaluator based on the dataset; The first retrieval module: used to establish a formal expression of simple geographical questions and build a simple question retrieval strategy; The second search module is used to establish a formal expression of complex geographical questions and construct a search strategy for complex questions; Reasoning module: used to combine geographic questions and retrieve documents, and build GeoPrompt prompt word templates to guide model reasoning.

Citation Information

Cited By

  • Output result generation method and system, electronic equipment and storage medium

    CN120277131A