Geographic scene triple extraction method for open text
By optimizing the multi-debate mechanism of sentence classification and multi-agent collaboration of open texts in geography, the problems of poor generalization of geographic information extraction and limited understanding of contextual relationships in the existing technology are solved, and a high-accurate geographic triple extraction is achieved.
Patent Information
- Application Number
- CN202411948674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
When extracting key information from the open field of geography, the prior art has poor generalization and limited understanding of contextual relationships. It lacks the complex semantic understanding of geographical knowledge such as causal relationships and hierarchical relationships, resulting in one-sided or incomplete extraction relationships.
By classifying the open text of geography Chinese into parallel sentences, hierarchical progressive sentences and multi-semantic element aggregation sentences, and processing them for different sentence patterns, a multi-debate mechanism is introduced to optimize the processing process, and multi-agent collaboration is used to complete the extraction of complex semantic triplets under the condition of few sample prompts.
It significantly improves the accuracy of extracting strong semantic knowledge triplets, reduces time, manpower and material resources costs, achieves the purpose of automated extraction of triplets, and performs excellently in evaluation indicators such as accuracy, recall, F1 value and BLEU.
Smart Images

Figure CN119990269A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of open triple extraction, in particular to a method for extracting geographic scene triples of open text. Background Art
[0002] Triples can improve information integration, semantic understanding and intelligent reasoning capabilities by constructing a structured entity relationship network, and effectively support knowledge application and decision-making in a big data environment. Open information extraction (OIE) does not need to predefine relationships and entity types in triple extraction, and can extract diverse relationships and entities, broadening the depth and breadth of triples. However, when facing geographical text knowledge, due to the characteristics of geographical knowledge such as spatial semantic complexity, temporal dimension dynamics, multi-scale geographical hierarchical structure, multiple geographical relationship intersections and complex geographical semantic causal relationships, the existing OIE extraction based on rules, statistics and neural networks depends on the coverage of the knowledge base; the existing methods based on large language models are mostly oriented to general domain texts. When extracting general information, they have the advantages of high versatility and accuracy and low labor cost, but they have poor generalization in geographical OIE, and limited understanding of contextual relationships. They lack understanding of complex semantics of geographical knowledge such as causal relationships and hierarchical relationships, resulting in problems such as one-sided or incomplete extracted relationships. It can be seen that when extracting effective key information in the open field of geography, it is impossible to use information extraction methods based on general large language models. Instead, experts from multiple fields are required to manually supplement, adjust, screen and build models. This not only increases time and economic costs, but also the volume and category of information obtained may change due to changes in the selection of training samples. Therefore, how to improve the accuracy of extracting strong semantic knowledge triples has become a problem that needs to be solved urgently. Summary of the invention
[0003] In order to solve the problems existing in the prior art, the present invention provides a method for extracting geographical scene triples from open texts. The method classifies geographical Chinese open texts into parallel sentences, hierarchical progressive sentences and multi-semantic element aggregation sentences, processes different sentence patterns, and introduces a multi-debate mechanism to optimize the processing process. The method can complete the extraction of complex semantic triples under the condition of few sample prompts by utilizing multi-agent collaboration without the need for a large amount of labeled data training or fine-tuning, and improve the accuracy of extracting strong semantic knowledge triples.
[0004] In order to solve the above technical problems, the present invention provides a method for extracting geographic scene triples from open text, comprising the following steps:
[0005] Step 1: Using geography-related data, construct a geography sentence classifier for parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences: According to the geography knowledge text input by the user, the geography sentence classifier classifies the geography knowledge text input by the user into three types: parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences;
[0006] Step 2: Construct a triple extraction module for different sentence characteristics to extract geography-related entities from geographical knowledge texts, and construct sentence characteristic prompt words for relationship extraction agents according to different sentence characteristics, and then call relationship extraction agents for different sentence patterns to form extraction relations to form triples;
[0007] Step 3: Build an optimization solution generation module for different sentence characteristics, build sentence characteristic prompt words for the inspection agent according to the different sentence characteristics, and then call the inspection agent for different sentences to provide inspection opinions for each triple inspection agent;
[0008] Step 4, construct a multi-agent debate optimization module for different sentence characteristics, construct sentence characteristic prompts for debate agents according to different sentence characteristics, and then call debate agents for different sentences to conduct debates to check opinions, and construct sentence characteristic prompts for decision agents according to different sentence characteristics, and then call decision agents for different sentences to give final modification opinions; the modification agent completes the modification of the triples according to the final modification opinions and the summary agent removes duplicate triples and outputs all triples.
[0009] Furthermore, in step 1, the parallel sentence of the geographical knowledge text is defined as: comprising two or more independent clauses, connected by parallel conjunctions, expressing parallel information; formally expressed as:
[0010] T parallcl →{T 1 (E1,R1,E2),T 2 (E1,R1,E3),...}→{(E1,R1,E2),(E1,R1,E3),...}
[0011] In the above formula, T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T parallcl Represents the parallel sentence structure in geography sentences.
[0012] Furthermore, in step 1, the hierarchical progressive sentence of the geographical knowledge text is defined as: a compound sentence with a logical progressive relationship is formed by gradually unfolding multiple levels of information or processes; this sentence pattern emphasizes the hierarchy and causal relationship between various elements, and aims to gradually explore the formation mechanism or evolution process of a certain geographical phenomenon; the formal expression is:
[0013] T progressive →{T 1 (E1,R1,E2),T 2 (E2,R1,E3)......}→{(E1,R1,E2),((E1,R1,E2),R2,E3),......}
[0014] In the above formula, {T 1 ,T 2 ,...} represents the set of all triples extracted from the hierarchical progressive sentence, T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T progressive Represents the hierarchical progressive sentence structure in geography sentences; in the subsequent triple T (n+1) E (n+1) It is composed of the previous triple T n The results obtained are affected.
[0015] Furthermore, in step 1, the multi-semantic element aggregation sentence of the geographical knowledge text is defined as: not only describing the relationship between different geographical phenomena, but also showing their relevance in terms of time, space, and causality, and the characteristic of the sentence is that it contains multiple levels of information, can simultaneously convey complex geographical concepts and processes, and is suitable for in-depth analysis and explanation of natural phenomena; the formal expression is:
[0016] T aggregated →{T parallel ,T progressive ,...}→{(E1,R1,E2),(E1,R1,E3),((E1,R1,E2),R2,E3),...}
[0017] The multi-semantic element aggregation sentence contains both parallel information in parallel sentences and causal logic in hierarchical progressive sentences. First, the parallel structure and progressive structure are separated from the sentence, and the initial triple set T under each structure is extracted. parallel and T progressiveThen, the triplets under these two structures are aggregated to form a complex association description, including direct parallel relationships and progressive relationships. In the parallel part, the head entity E1 of each triple is parallelly associated with multiple tail entities E2, E3,...; in the progressive part, the head entity or tail entity of the subsequent triple may depend on the result of the previous triple, resulting in a nested relationship, such as ((E1, R1, E2), R2, E3).
[0018] Furthermore, in step 2, a triple extraction module for different sentence characteristics is constructed to extract geography-related entities from geographical knowledge texts, and sentence characteristic prompt words for relationship extraction agents are constructed according to different sentence characteristics. The specific steps of calling relationship extraction agents for different sentence patterns to form extraction relations and triples are as follows:
[0019] Step 2.1: For the input geography text, build a triple extraction module targeting different sentence characteristics to extract all geography professional entities;
[0020] Step 2.2: Construct a relationship based on different sentence features to extract the sentence feature prompt words of the intelligent agent;
[0021] Step 2.3: Call the relation extraction agent for different sentence patterns to form extraction relations;
[0022] Step 2.4: The output results of the entity extraction agent will all be provided to the relationship extraction agents for different sentence patterns. The relationship extraction agents for different sentence patterns will sequentially perform the overall semantic understanding, triple form constraints, qualifier condition supplementation and repetitive correction task steps, and output the generated triples.
[0023] Furthermore, the specific steps of constructing a relationship according to different sentence characteristics to extract the sentence characteristic prompt words of the intelligent agent in step 2.2 are as follows:
[0024] For parallel sentences, parallel sentences contain multiple independent clauses, which are connected by coordinating conjunctions to express parallel information, so special attention should be paid to the relationship between the clauses in parallel sentences;
[0025] For hierarchical progressive sentences, the relational names are not limited to simple conjunctions, but emphasize specific causal relationships, action mechanisms, and evolutionary processes to ensure that entities can be connected to form hierarchical progressive relationships;
[0026] For sentences with multiple semantic elements, when extracting relationships, we should start from the overall semantics of the sentence, pay attention to the multi-level information within the sentence, respect the description of the original sentence, and not over-simplify it.
[0027] Furthermore, the specific steps of calling the relation extraction agent for different sentence patterns in step 2.3 are as follows:
[0028] Step 2.3.1, overall semantic understanding: The triples formed are not fine-grained triples, but macro-semantic triples. It is not necessary to form triples for each word. Therefore, entities that cannot form triples with correct meanings are discarded.
[0029] Step 2.3.2, triple form constraint: pay attention to the order of the head entity and the tail entity, if the order is inconsistent, correct it;
[0030] Step 2.3.3: Supplement the restrictive conditions: Pay attention to the modifiers and restrictive conditions in the original text to ensure that the semantics of the triples are correct;
[0031] Step 2.3.4, repeatability correction: remove duplicate triplets.
[0032] Furthermore, the specific steps of constructing the sentence characteristic prompt words of the inspection agent according to different sentence characteristics in step 3 are as follows:
[0033] For parallel sentences, the generated triples are provided one by one to the checking agents for different sentence patterns in the optimization solution generation module. When checking entities, attention should be paid to the possibility of multiple different subject switches in the sentence to avoid semantic confusion and ensure that the triples can accurately restore the meaning of the original text. At the same time, a coreference resolution step is introduced to handle different expressions of the same entity, reduce redundant information, and improve the refinement of the triples. When checking relationship names, attention should be paid to the selection of relationship names to avoid confusion.
[0034] For hierarchical progressive sentences, in view of the causal chain triples in the hierarchical progressive sentences, adjacent triple information is provided in each inspection link to avoid misjudgment of inspection and debate due to information isolation.
[0035] For sentences aggregated with multiple semantic elements, after the relation extraction agent for different sentence patterns generates triples, the triple classification agent for sentences aggregated with multiple semantic elements will divide the triples into parallel sentence triples and hierarchical progressive sentence triples, and then execute the corresponding step 4 process.
[0036] Furthermore, the specific steps of calling the checking agent for different sentence patterns in step 3 are as follows:
[0037] Overall inspection: Check the semantic order and extraction granularity of triples and original text;
[0038] Entity check: Check whether the subject of the triple is correct. If qualifiers and modifiers are missing, add them.
[0039] Relationship name check: pay attention to the subject of the sentence. Some unclear relationships can be reasonably supplemented, and obvious relationships must all come from the original text;
[0040] Semantic check: Pay attention to the correlation and progressive relationship between the previous and following semantics to ensure that the triples express the original meaning of the complete sentence.
[0041] Furthermore, in step 4, the debate agent debates, and the decision agent gives the final modification opinion, and then the triples are modified and all triples are output as follows:
[0042] The debate agents for different sentence structures include the affirmative and negative sides. They debate on the four checks provided by the check agents for different sentence structures. The debate theme is: whether the conclusion is correct. There are five rounds of debate in total. After the debate, the decision agent observes the debate history and gives the final check conclusion. Finally, the modification agent modifies the triples according to the final check conclusion and outputs all triples after removing duplicate triples through the summary agent.
[0043] The detailed steps are as follows:
[0044] Step 4.1: Construct debate agents for different sentence structures. The debate agents include the affirmative side and the negative side. They debate on the four inspection conclusions provided by the inspection agents for different sentence structures, namely entity inspection, relation name inspection, semantic inspection and overall inspection. The debate topic is: whether the inspection conclusion is correct;
[0045] To build debate agents for different sentence patterns, we need to first build debate agents for different sentence patterns. Sentence pattern prompts: For parallel sentences, sometimes some relationships are more obscure, such as the parallel relationship or inclusion relationship described in parallel sentences. In these cases, we need to supplement or infer the relationship name on the basis of ensuring the original meaning, so as to better reflect the structure and semantics of the parallel sentence; for hierarchical progressive sentences, we need to ensure the semantic progressive relationship in the original sentence, and pay attention to the causal relationship between adjacent triples.
[0046] Step 4.2: To ensure that the two agents reach a consensus as much as possible, debate five rounds in total. During the debate, remind the agents to pay attention to the fact that the two debaters should give their own opinions based on the inspection opinions, respond closely to the opponent's opinions, and do not have to fully agree with the opponent's opinions. The answers should be brief and fully state the opinions.
[0047] Step 4.3: After the debate, build a decision-making agent for different sentence patterns, and let the decision-making agent observe the debate history and give the final inspection conclusion;
[0048] Before building decision-making agents for different sentence patterns, it is necessary to build sentence feature prompts for decision-making agents for different sentence patterns: for parallel sentences, when making decisions, choose the side with the correct opinions on parallel relationship expression, parallel semantics, and modifier or qualifier addition; for hierarchical progressive sentences, choose the side with the correct opinions on hierarchical progressive relationship and causal chain entity extraction;
[0049] Step 4.4: Finally, the modification agent modifies the triples according to the final inspection conclusion and outputs all triples after eliminating duplicate triples through the summary agent.
[0050] The beneficial effects of the present invention are:
[0051] Compared with the existing knowledge mining methods that require tens of thousands of labeled data for training or fine-tuning, the present invention can complete the extraction of complex semantic triples under the conditions of a small number of sample prompts through multi-agent collaboration, which significantly reduces the time, manpower and material costs required for triple extraction of related complex semantic texts, and ultimately achieves the purpose of automatic extraction of triples.
[0052] At the same time, compared with other information extraction methods of zero-sample or small-sample large language models, the framework of the present invention does not limit entities and relational words, does not need to enumerate all possible relationships and preliminary named entities, and has the ability to flexibly handle complex geographical Chinese semantic texts. Taking the Qwen2-7B-Instruct model as an example, it can improve the performance of small-parameter large language models, reaching or even exceeding the capabilities of large language models with 72B and 11B parameters such as Qwen2-72B-Instruct and Qwen1.5-110B-Chat. Ultimately, it performs well in terms of accuracy, recall, F1 value, and BLEU and other evaluation indicators in triple extraction. Compared with the 30%-50% accuracy of using a large language model to ask prompt words, the results extracted by this framework can reach an accuracy of more than 80%, and the recall rate is 25% higher than that of previous methods. The framework can recognize triples in sentences that previous methods have not been able to recognize, and can adapt to the current rapid iteration and update of knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flow chart of the present invention;
[0054] Figure 2 It is a flow chart of a module for extracting triples of multi-semantic elements aggregated sentences in a specific embodiment;
[0055] Figure 3 A flowchart of a module for generating a multi-semantic element aggregation sentence optimization solution in a specific embodiment;
[0056] Figure 4It is a flowchart of the debate process of multi-semantic element aggregation sentence in a specific embodiment;
[0057] Figure 5 This is an example diagram comparing the results of the present invention and the prompt question extraction results in a specific embodiment. DETAILED DESCRIPTION
[0058] Embodiment 1:
[0059] like Figure 1 As shown, the present invention provides a method for extracting geographic scene triples from open text, comprising the following steps:
[0060] Step 1: Using geography-related data, construct a geographical sentence classifier for parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences.
[0061] Step 1.1: According to the geographical knowledge text input by the user, the geographical sentence classifier classifies the geographical knowledge text input by the user into three types: parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences.
[0062] Step 1.2: In the present invention, a parallel sentence in a geographical knowledge text is defined as: containing two or more independent clauses connected by parallel conjunctions to express parallel information. The formal expression is:
[0063] T parallcl →{T 1 (E1,R1,E2),T 2 (E1,R1,E3),...}→{(E1,R1,E2),(E1,R1,E3),...} (1)
[0064] In formula (1), T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T parallcl Represents the parallel sentence structure in geography sentences.
[0065] Step 1.3: In the present invention, the hierarchical progressive sentence of the geographical knowledge text is defined as: a compound sentence with a logical progressive relationship is formed by gradually unfolding multiple levels of information or processes. This sentence pattern emphasizes the hierarchy and causal relationship between various elements, and aims to gradually explore the formation mechanism or evolution process of a certain geographical phenomenon. The formal expression is:
[0066] T progressive →{T 1 (E1,R1,E2),T2 (E2,R1,E3)......}→{(E1,R1,E2),((E1,R1,E2),R2,E3),......}(2)
[0067] In formula (2), {T 1 ,T 2 ,...} represents the set of all triples extracted from the hierarchical progressive sentence, T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T progressive Represents the hierarchical sentence structure in geography. In the subsequent triple T (n+1) E (n+1) It is composed of the previous triple T n The results obtained are affected.
[0068] Step 1.4: In the present invention, the multi-semantic element aggregation sentence of geographical knowledge text is defined as: not only describing the relationship between different geographical phenomena, but also showing their relevance in time, space, and causality, and the characteristic of the sentence is that it contains multiple levels of information, can simultaneously convey complex geographical concepts and processes, and is suitable for in-depth analysis and explanation of natural phenomena. The formal expression is:
[0069] T aggregated →{T parallel ,T progressive ,...}→{(E1,R1,E2),(E1,R1,E3),((E1,R1,E2),R2,E3),...}(3)
[0070] The multi-semantic element aggregation sentence contains both parallel information in parallel sentences and causal logic in hierarchical progressive sentences. First, the parallel structure and progressive structure are separated from the sentence, and the initial triple set T under each structure is extracted. parallel and T progressive Then, the triples under these two structures are aggregated to form a complex association description, including direct parallel relationships and progressive relationships. In the parallel part, the head entity E1 of each triple can be parallelly associated with multiple tail entities E2, E3, ... In the progressive part, the head entity or tail entity of the subsequent triple may depend on the result of the previous triple, resulting in a nested relationship, such as ((E1, R1, E2), R2, E3).
[0071] Step 2: Construct a triple extraction module targeting different sentence patterns to extract geography-related entities from geographic knowledge texts, and call relation extraction agents for different sentence patterns to form extraction relations and triples according to different sentence patterns.
[0072] Step 2.1. Construct the basic task requirements of the entity extraction agent in the triple extraction module targeting different sentence characteristics: Extract all geographical professional entities from the input geographical text.
[0073] Step 2.2: Construct a triple extraction module for different sentence patterns. The basic task requirements for the relationship extraction agent are: Based on the previously extracted geographical entities, extract the relationships between them from the text to form triples, and give your reasons for extracting the relationships.
[0074] Step 2.3: Construct a relation extraction agent for different sentence structures. Steps for relation extraction task:
[0075] Step 2.3.1, overall understanding, the triples formed are not fine-grained triples, but macro-semantic triples, and it is not necessary to form triples for each word. Therefore, please discard entities that cannot form correct meaning triples, and do not force them to form triples.
[0076] Step 2.3.2: Constrain the triple form. Pay attention to the order of the head entity and the tail entity. If the order is inconsistent, correct it.
[0077] Step 2.3.3: Supplement the restrictive conditions. Pay attention to the modifiers and restrictive conditions in the original text to ensure that the semantics of the triples are correct.
[0078] Step 2.3.4: Repeatability correction, remove duplicate triplets.
[0079] Step 2.4: The basic task requirements for constructing a triple classification agent for multi-semantic element aggregation sentences are: divide the input triples into parallel sentence triples and hierarchical progressive sentence triples.
[0080] Step 2.5: Construct sentence feature prompt words for the relationship extraction agent according to different sentence features. For parallel sentences, parallel sentences contain multiple independent clauses, which are connected by parallel conjunctions to express parallel information. Therefore, special attention should be paid to the relationship between the clauses in the parallel sentences. For hierarchical progressive sentences, the relationship name is not limited to simple conjunctions, but emphasizes the specific causal relationship, mechanism of action and evolution process to ensure that the entities can be connected to form a hierarchical progressive relationship. For multi-semantic element aggregation sentences, when extracting relationships, we should start from the overall semantics of the sentence, pay attention to the multi-level information within the sentence, respect the description of the original sentence, and not over-simplify it.
[0081] Step 2.6: Input the constructed basic task requirements, sentence feature prompt words, and task steps into the entity extraction agent, the relationship extraction agent, and the triple classification agent for multi-semantic element aggregation sentences.
[0082] Step 2.7: The output results of the entity extraction agent will all be provided to the relationship extraction agents for different sentence patterns. The relationship extraction agents for different sentence patterns will execute the task steps in sequence and output the generated triples.
[0083] Step 2.8, for parallel sentences, the generated triple results are provided one by one to the inspection agents for different sentence patterns in the optimization scheme generation module; for hierarchical progressive sentences, considering the logical relationship between triples, the triples to be checked and the remaining triples are provided to the triple classification agent for multi-semantic element aggregate sentences; for multi-semantic element aggregate sentences, after the relation extraction agents for different sentence patterns output triples, the triple classification agent for multi-semantic element aggregate sentences divides the triples into parallel sentence triples and hierarchical progressive sentence triples, and then executes their respective corresponding processes.
[0084] Step 3: Construct an optimization solution generation module targeting different sentence patterns, call inspection agents for different sentence patterns according to different sentence patterns, and provide inspection opinions for each triple inspection agent.
[0085] Step 3.1: Construct optimization solutions for different sentence patterns. The basic task requirements of the generation module for different sentence patterns are: Check the problems in each triple.
[0086] Step 3.2: Construct inspection agents for different sentence structures. Steps for inspection tasks:
[0087] Step 3.2.1, overall check: check the semantic order and extraction granularity of the triples and the original text;
[0088] Step 3.2.2, Entity Check: Check whether the subject of the triple is correct. If qualifiers and modifiers are missing, add them.
[0089] Step 3.2.3, check the relationship name: pay attention to the subject of the sentence. Some unclear relationships can be reasonably supplemented, and obvious relationships must all come from the original text;
[0090] Step 3.2.4, semantic check: pay attention to the correlation and progressive relationship between the previous and following semantics, and ensure that the triples express the original meaning of the complete sentence.
[0091] Step 3.3: Construct sentence feature prompt words for the inspection agent according to different sentence features. For parallel sentences, when checking entities, pay attention to the possibility of multiple different subject switches in the sentence to avoid semantic confusion and ensure that the triples can accurately restore the meaning of the original text. At the same time, introduce the coreference resolution step to handle different expressions of the same entity, reduce redundant information, and improve the refinement of triples. When checking relationship names, pay attention to the choice of relationship names to avoid confusion. For hierarchical progressive sentences, for the causal chain triples in hierarchical progressive sentences, adjacent triple information is provided in each inspection link to avoid misjudgment of inspection and debate due to information isolation.
[0092] Step 3.4: Input the constructed basic task requirements, sentence feature prompts, and task steps into the inspection agent.
[0093] Step 3.5: After the inspection agents for different sentence patterns have completed the inspection process in sequence, for parallel sentences, the inspection opinions and triples to be modified are provided to the multi-agent debate optimization module for different sentence characteristics; for hierarchical progressive sentences, the inspection opinions, triples to be modified and the remaining triples are provided to the multi-agent debate optimization module for different sentence characteristics.
[0094] Step 4: Construct a multi-agent debate optimization module for different sentence characteristics, call debate agents for different sentence patterns to debate on inspection opinions according to different sentence characteristics, and the decision agent gives the final modification opinion. The modification agent completes the modification of the triples according to the final modification opinion, and the summary agent removes duplicate triples and outputs all triples.
[0095] Step 4.1: Construct a multi-agent debate optimization module for different sentence characteristics. Basic task requirements for debate agents for different sentence characteristics: Debate on the inspection opinions. Closely respond to the opponent's latest arguments and state your position.
[0096] Step 4.2: Construct the basic task requirements of the decision-making agent for different sentence characteristics in the multi-agent debate optimization module for different sentence characteristics: Make the correct judgment after the debate between the previous two agents and choose the one you think is correct.
[0097] Step 4.3: Construct the basic task requirements of modifying the agent in the multi-agent debate optimization module for different sentence characteristics: Modify the triples according to the provided modification suggestions on the triples.
[0098] Step 4.4: Construct a multi-agent debate optimization module targeting different sentence patterns and summarize the basic task requirements of the agents: remove semantically repeated triples from the provided triples.
[0099] Step 4.5: Construct sentence feature prompts for debate agents based on different sentence features. For parallel sentences, sometimes some relationships are more obscure, such as the parallel relationship or inclusion relationship described in parallel sentences. In these cases, it is required to supplement or infer the relationship name on the basis of ensuring the original meaning, so as to better reflect the structure and semantics of the parallel sentence; for hierarchical progressive sentences, it is necessary to ensure the semantic progressive relationship in the original sentence, and pay attention to the causal relationship between adjacent triples.
[0100] Step 4.6: Construct sentence feature prompts for decision agents based on different sentence features. For parallel sentences, choose the side with correct opinions on parallel relationship expression, correct opinions on parallel semantics, and correct opinions on adding modifiers or qualifiers when making decisions; for hierarchical progressive sentences, choose the side with correct opinions on hierarchical progressive relationship and correct opinions on causal chain entity extraction.
[0101] Step 4.7: Input the constructed basic task requirements, sentence characteristic prompts, and task steps into the debate agent and decision agent for different sentence patterns.
[0102] Step 4.8: After generating the four conclusions of the inspection, the inspection agent for different sentence structures will provide them to the debate agents for different sentence structures for debate. During the debate, the agent is reminded to pay attention to the fact that the two debaters should give their own opinions based on the inspection opinions, respond closely to the opponent's opinions, and do not have to fully agree with the opponent's opinions. The answers should be brief and fully express the opinions.
[0103] Step 4.9: After observing the debate process of the two debaters, the decision-making agent for different sentence structures makes its own judgment and chooses the correct side, while ensuring that the content that appears in the debate process and is not in the original sentence is not added to the review opinion. The decision-making agent for different sentence structures outputs the final four review opinions.
[0104] Step 4.10: Provide the final modification opinions generated by the decision-making agent for different sentence patterns as input to the modification agent. The modification agent completes the modification of the triples according to the final modification opinions, and the summary agent removes duplicate triples and outputs all triples.
[0105] In a specific example of geomorphology knowledge text experiment information, the prompt word template of the constructed agent is as follows:
[0106] You are {character} and your mission is {mission name}.
[0107] ***{Sentence definition}***
[0108] ***{Task Notes}***
[0109] ***{Key points for handling this sentence}***
[0110] ***{Task Result Example}***
[0111] Finally, it returns according to the set format. This is the input text: ```{text}```. The answer process is all output in Chinese. Output format: {output format}, do not output other irrelevant content.
[0112] Output: The output of the agent.
[0113] In a specific example of geomorphology knowledge text experimental information, since the execution process of multi-semantic element aggregation sentence involves the main processes of parallel sentences and hierarchical progressive sentences, the execution process of multi-semantic element aggregation sentence is provided, and the specific steps include:
[0114] Step 1: According to the geographical knowledge text input by the user, the geographical sentence classifier classifies the geographical knowledge text input by the user into three types: parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences; Figure 2 As shown, a sentence with multiple semantic elements is taken as an example.
[0115] Step 2: Construct a triple extraction module for different sentence characteristics to extract geography-related entities from geographical knowledge texts, and construct sentence characteristic prompt words for relationship extraction agents based on different sentence characteristics, and then call relationship extraction agents for different sentence patterns to form extraction relations and triples. Figure 2 As shown, the entity extraction agent extracted the entities: "karst area, rainwater dissolution, stone buds, funnels, sinkholes, surface water system destruction, underground rivers, surface rivers, river valleys" and provided them to the relationship extraction agent, which generated preliminary triples.
[0116] Step 3: Construct an optimization solution generation module for different sentence characteristics, construct sentence characteristic prompt words for the inspection agent according to the different sentence characteristics, and then call the inspection agent for different sentences to provide inspection opinions for each triple inspection agent; Figure 3 As shown, the relationship extraction agent provides the extracted preliminary triples to the classification agent, and the classification agent classifies the preliminary triples into "parallel sentence triples" and "hierarchical progressive sentence triples", and provides them to the inspection agents for parallel sentences and hierarchical progressive sentences for inspection respectively to obtain inspection conclusions.
[0117] Step 4: Construct a multi-agent debate optimization module for different sentence characteristics, construct sentence characteristic prompts for debate agents according to different sentence characteristics, and then call debate agents for different sentences to check the opinions, and construct sentence characteristic prompts for decision agents according to different sentence characteristics, and then call decision agents for different sentences to give final modification opinions; the modification agent completes the modification of the triples according to the final modification opinions, and the summary agent removes duplicate triples and outputs all triples. Figure 4 As shown in the figure, a debate process for the relationship name check conclusion is given. After five rounds of debate, the decision-making agent gives the final check conclusion and provides it to the modification agent to modify the preliminary triples. Finally, the summary agent removes duplicates and outputs all triples.
[0118] like Figure 5 As shown, a comparison example of the results of the present invention and the results of asking questions through prompt is provided. The present invention has the following advantages:
[0119] 1. Extracting triples only through prompts cannot adapt to geographical texts with multiple geographical relationships and complex geographical semantic causal relationships, such as Figure 3 In the embodiment of the present invention, colloquial or implicit relationship names such as "also known as" and "surrounded by" are extracted.
[0120] 2. The present invention can avoid missing triples in the original sentence and greatly improve the richness of the extracted triples, such as Figure 3 As shown, the number of triples extracted by the present invention far exceeds the number of questions asked by the prompt.
[0121] 3. The present invention can supplement the limiting conditions missing in entities in different sentence patterns. Limiting conditions are indispensable factors for correctly describing geographical concepts and geographical principles in geography.
[0122] 4. The present invention can further improve the accuracy of triple extraction of large language models on geographical Chinese open texts, and does not require labeled data sets for training, only a few sample prompts are needed. It is a few-sample geographical Chinese open text triple extraction method.
[0123] Embodiment 2:
[0124] This embodiment provides a multi-agent geographical Chinese open text triple extraction system, including:
[0125] The geographical sentence classification module constructs a geographical sentence classifier for parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences through geography-related data.
[0126] The intelligent agent group prompt template module constructs an intelligent agent group prompt template corresponding to each sentence pattern. According to the geographical knowledge text input by the user, the geographical sentence classifier classifies the geographical knowledge text input by the user into sentence patterns; different intelligent agent group prompt templates are executed according to the sentence pattern classification results.
[0127] The triple extraction module targeting the characteristics of different sentence patterns can extract geography-related entities from geographical knowledge texts, and call the relationship extraction agents for different sentence patterns to form triples based on the extraction relations.
[0128] An optimization solution generation module is developed for different sentence patterns. The inspection agent for different sentence patterns is called according to different sentence patterns, and inspection opinions are provided for each triple inspection agent.
[0129] The multi-agent debate optimization module for different sentence characteristics calls debate agents for different sentence patterns to conduct debates based on the characteristics of different sentence patterns, and the decision agent gives the final modification opinions. The modification agent completes the modification of the triples based on the final modification opinions, and the summary agent removes duplicate triples and outputs all triples.
[0130] Embodiment three:
[0131] This embodiment provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a multi-agent geography Chinese open text triple extraction method as described in Example 1 of the present invention.
[0132] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0133] In the several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
Claims
1. A method for extracting geographic scene triples from open text, characterized in that: The following steps are involved: Step 1: Using geography-related data, construct a geography sentence classifier for parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences: According to the geography knowledge text input by the user, the geography sentence classifier classifies the geography knowledge text input by the user into three types: parallel sentences, multi-semantic element aggregation sentences, and hierarchical progressive sentences; Step 2: Construct a triple extraction module for different sentence characteristics to extract geography-related entities from geographical knowledge texts, and construct sentence characteristic prompt words for relationship extraction agents according to different sentence characteristics, and then call relationship extraction agents for different sentence patterns to form extraction relations to form triples; Step 3: Build an optimization solution generation module for different sentence characteristics, build sentence characteristic prompt words for the inspection agent according to the different sentence characteristics, and then call the inspection agent for different sentences to provide inspection opinions for each triple inspection agent; Step 4, construct a multi-agent debate optimization module for different sentence characteristics, construct sentence characteristic prompts for debate agents according to different sentence characteristics, and then call debate agents for different sentences to conduct debates to check opinions, and construct sentence characteristic prompts for decision agents according to different sentence characteristics, and then call decision agents for different sentences to give final modification opinions; the modification agent completes the modification of the triples according to the final modification opinions and the summary agent removes duplicate triples and outputs all triples.
2. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: In step 1, the parallel sentence of the geographical knowledge text is defined as: containing two or more independent clauses, connected by parallel conjunctions, expressing parallel information; formally expressed as: T parallcl →{T 1 (E1,R1,E2),T 2 (E1,R1,E3),...}→{(E1,R1,E2),(E1,R1,E3),...} In the above formula, T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T parallcl Represents the parallel sentence structure in geography sentences.
3. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: In step 1, the hierarchical progressive sentence of the geographical knowledge text is defined as: a complex sentence with a logical progressive relationship is formed by gradually unfolding multiple levels of information or processes; the formal expression is: T progressive →{T 1 (E1,R1,E2),T 2 (E2,R1,E3)......}→{(E1,R1,E2),((E1,R1,E2),R2,E3),......} In the above formula, {T 1 ,T 2 ,...} represents the set of all triples extracted from the hierarchical progressive sentence, T 1 ,T 2 Represents the triple formed by the semantic elements in the sentence, E n is the entity word in the triple, R n is the relationship word between the head entity and the tail entity, T progressive Represents the hierarchical progressive sentence structure in geography sentences; in the subsequent triple T (n+1) E (n+1) It is composed of the previous triple T n The results obtained are affected.
4. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: In step 1, the multi-semantic element aggregation sentence of the geographical knowledge text is defined as: describing the relationship between different geographical phenomena and showing their relevance in time, space and causality. The sentence contains multiple levels of information and can simultaneously convey complex geographical concepts and processes. The formal expression is: T aggregated →{T parallel ,T progressive ,...}→{(E1,R1,E2),(E1,R1,E3),((E1,R1,E2),R2,E3),...} The multi-semantic element aggregation sentence contains both parallel information in parallel sentences and causal logic in hierarchical progressive sentences. First, the parallel structure and progressive structure are separated from the sentence, and the initial triple set T under each structure is extracted. parallel and T progressive , and then aggregate the triples under these two structures to form a complex association description, including direct parallel relationships and progressive relationships; in the parallel part, the head entity E1 of each triple is parallelly associated with multiple tail entities E2, E3,...; in the progressive part, the head entity or tail entity of the subsequent triple may depend on the result of the previous triple, resulting in a nested relationship.
5. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: In step 2, a triple extraction module for different sentence characteristics is constructed to extract geography-related entities from geographical knowledge texts, and sentence characteristic prompt words for relationship extraction agents are constructed according to different sentence characteristics. The specific steps of calling relationship extraction agents for different sentence patterns to form extraction relations and triples are as follows: Step 2.1: For the input geography text, build a triple extraction module targeting different sentence characteristics to extract all geography professional entities; Step 2.2: Construct a relationship based on different sentence features to extract the sentence feature prompt words of the intelligent agent; Step 2.3: Call the relation extraction agent for different sentence patterns to form extraction relations; Step 2.4: The output results of the entity extraction agent will all be provided to the relationship extraction agents for different sentence patterns. The relationship extraction agents for different sentence patterns will sequentially perform the overall semantic understanding, triple form constraints, qualifier condition supplementation and repetitive correction steps, and output the generated triples.
6. The method for extracting geographic scene triples from open text according to claim 5, characterized in that: The specific steps of constructing a relationship according to different sentence characteristics in step 2.2 to extract the sentence characteristic prompt words of the intelligent agent are as follows: For parallel sentences, parallel sentences contain multiple independent clauses that are connected by coordinating conjunctions to express parallel information; For hierarchical progressive sentences, the relational names are not limited to simple conjunctions, but emphasize specific causal relationships, action mechanisms, and evolutionary processes to ensure that entities can be connected to form hierarchical progressive relationships; For sentences with multiple semantic elements, when extracting relations, we start from the overall semantics of the sentence and focus on the multi-level information within the sentence.
7. The method for extracting geographic scene triples from open text according to claim 5, characterized in that: The specific steps of calling the relation extraction agent for different sentence patterns in step 2.3 are as follows: Step 2.3.1, overall semantic understanding: the triples formed are macro-semantic triples, and entities that cannot form correct meaning triples are discarded; Step 2.3.2, triple form constraint: pay attention to the order of the head entity and the tail entity, if the order is inconsistent, correct it; Step 2.3.3: Supplement the restrictive conditions: Pay attention to the modifiers and restrictive conditions in the original text to ensure that the semantics of the triples are correct; Step 2.3.4, repeatability correction: remove duplicate triplets.
8. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: The specific steps of constructing the sentence characteristic prompt words of the inspection agent according to different sentence characteristics in step 3 are as follows: For parallel sentences, when checking entities, pay attention to the possibility of multiple different subject switches in the sentence, ensure that the triples can accurately restore the meaning of the original text, and handle different expressions of the same entity; when checking relationship names, pay attention to the choice of relationship names to avoid confusion; For hierarchical progressive sentences, for the causal chain triples in hierarchical progressive sentences, adjacent triple information is provided in each inspection link to avoid misjudgment of inspection and debate caused by isolated information; For sentences aggregated with multiple semantic elements, after the relation extraction agent for different sentence patterns generates triples, the triple classification agent for sentences aggregated with multiple semantic elements will divide the triples into parallel sentence triples and hierarchical progressive sentence triples, and then execute the corresponding step 4 process.
9. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: The specific steps of calling the inspection agent for different sentence patterns in step 3 are as follows: Overall inspection: Check the semantic order and extraction granularity of triples and original text; Entity check: Check whether the subject of the triple is correct. If qualifiers and modifiers are missing, add them. Relationship name check: pay attention to the subject of the sentence. Some unclear relationships need to be supplemented reasonably. Obvious relationships must all come from the original text. Semantic check: Pay attention to the correlation and progressive relationship between the previous and following semantics to ensure that the triples express the original meaning of the complete sentence.
10. The method for extracting geographic scene triples from open text according to claim 1, characterized in that: In step 4, the debate agent debates, and the decision agent gives the final modification opinion, and then modifies the triples and outputs all triples. The specific steps are as follows: Step 4.1: Construct debate agents for different sentence structures. The debate agents include the affirmative side and the negative side. The debate is conducted based on the four checks provided by the check agents for different sentence structures, namely, overall check, entity check, relation name check, and semantic check. The debate topic is: check whether the conclusion is correct; To build debate agents for different sentence patterns, we need to first build debate agents for different sentence patterns. Sentence pattern feature prompts: for parallel sentences, we need to supplement or infer the relationship names on the basis of ensuring the original meaning; for hierarchical progressive sentences, we need to ensure the semantic progressive relationship of the original sentence, and pay attention to the causal relationship between adjacent triples. Step 4.2, debate five rounds in total; during the debate, remind the agent to pay attention to the two debaters to give their own opinions based on the inspection opinions, closely respond to the opponent's opinions, and do not have to fully agree with the opponent's opinions; the answers should be brief and fully state the opinions; Step 4.3: After the debate, build a decision-making agent for different sentence patterns, and let the decision-making agent observe the debate history and give the final inspection conclusion; Before building decision-making agents for different sentence patterns, it is necessary to build sentence feature prompts for decision-making agents for different sentence patterns: for parallel sentences, when making decisions, choose the side with the correct opinions on parallel relationship expression, parallel semantics, and modifier or qualifier addition; for hierarchical progressive sentences, choose the side with the correct opinions on hierarchical progressive relationship and causal chain entity extraction; Step 4.4: Finally, the modification agent modifies the triples according to the final inspection conclusion and outputs all triples after eliminating duplicate triples through the summary agent.