A geological disaster risk double control method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610928742.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]然而,传统地质灾害风险管控方法在数据处理与应用中存在明显局限:一方面,对多类型原始数据的提取处理能力不足,难以构建统一、精准的知识载体,导致数据价值无法充分挖掘;另一方面,查询方式灵活性差,受限于固定流程无法满足复杂业务的开放式查询需求,且所用模型对地质灾害业务的数据分析准确性与效率较低,进一步导致预警响应不及时、风险识别不精细
根据所述图谱查询语句以及所述地质灾害知识图谱得到分析结果;
Smart Images

Figure CN122596674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of geological disaster risk prevention and control technology, and in particular to a method, device, equipment and storage medium for dual control of geological disaster risks. Background Technology
[0002] With the increasing demand for geological disaster prevention and control, geological disaster-related data are becoming more diversified, complex, and large-scale, encompassing both structured data tables of potential hazards / risk areas and unstructured disaster reports, among other raw data. Effective processing is urgently needed to achieve precise risk management.
[0003] However, traditional geological disaster risk management methods have obvious limitations in data processing and application: on the one hand, they lack the ability to extract and process various types of raw data, making it difficult to build a unified and accurate knowledge carrier, which results in the inability to fully explore the value of the data; on the other hand, the query methods are inflexible, and the fixed process cannot meet the open query needs of complex businesses. Moreover, the models used have low accuracy and efficiency in data analysis of geological disaster businesses, which further leads to untimely early warning response and imprecise risk identification.
[0004] Therefore, how to efficiently process raw data of multi-source geological disasters and improve the efficiency and accuracy of dual control of geological disaster risks is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the geological disaster risk dual control method, device, equipment, and storage medium provided in this application embodiment can integrate knowledge graph and model optimization technology to achieve accurate analysis and efficient management of geological disaster risks, providing reliable technical support for disaster early warning and hidden danger management, and is applicable to various geological disaster risk dual control scenarios. The geological disaster risk dual control method, device, equipment, and storage medium provided in this application embodiment are implemented as follows: This application provides a method for dual risk control of geological disasters, including: Obtain raw data on geological disasters; The raw data on geological hazards are extracted and processed to obtain a geological hazard knowledge graph; Obtain prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and the prompt words to obtain the optimized risk dual control model; Obtain the query statement and input the query statement into the optimized risk dual control model to obtain the graph query statement; The analysis results are obtained based on the map query statement and the geological disaster knowledge map. The analysis results are parsed and processed to obtain the risk dual control results.
[0006] In some embodiments, the raw geological disaster data includes structured data and unstructured data, and the extraction and processing of the raw geological disaster data to obtain a geological disaster knowledge graph includes: The unstructured data is divided into blocks according to a preset length to obtain processed unstructured data; The processed unstructured data is subjected to dual-strategy extraction to obtain a table file. The dual-strategy extraction includes semantic extraction and preset dictionary rule matching extraction. The structured data is processed by attribute extraction and merging to obtain a graph query statement; The geological disaster knowledge graph is obtained based on the table file and the graph query statement.
[0007] In some embodiments, obtaining the prompt words involves constructing a training dataset based on the geological disaster knowledge graph, and optimizing the risk dual control model according to the training dataset and the prompt words to obtain an optimized risk dual control model, including: Based on the dual control requirements for geological disasters and the entity and relation information of the geological disaster knowledge graph, the prompt words are obtained. The prompt words include geological disaster business annotations, graph structure descriptions, task definitions, and problem examples. The training dataset is obtained from the query statement dataset and the geological disaster question-answer pairs generated based on the geological disaster knowledge graph; The risk dual control model is optimized based on the training dataset and the prompt words to obtain the optimized risk dual control model.
[0008] In some embodiments, the dual-strategy extraction process on the processed unstructured data to obtain a table file includes: Semantic extraction is performed on the processed unstructured data to obtain semantic extraction results; The processed unstructured data is subjected to character matching and extraction to obtain the rule matching extraction results; A table file is obtained based on the semantic extraction results and the rule matching extraction results.
[0009] In some embodiments, the step of extracting and merging attributes from the structured data to obtain a graph query statement includes: The structured data is subjected to relation extraction processing to obtain relation extraction results; The structured data is subjected to attribute merging processing to obtain the attribute merging result; The graph query statement is obtained based on the relationship extraction results and attribute merging results.
[0010] In some embodiments, after parsing the analysis results to obtain the risk dual control results, the method further includes: Obtain the feedback processing result, and optimize the parameters of the geological disaster knowledge graph and the optimized risk dual control model based on the feedback processing result to obtain the optimized geological disaster knowledge graph and parameters. The feedback processing result is the user feedback operation in the analysis result. The raw geological disaster data is cleaned to obtain processed raw geological disaster data; A graph space is constructed based on the geological disaster knowledge graph to obtain a dedicated graph space; The processed original geological disaster data, the dedicated map space, and the risk dual control results are loaded into the map.
[0011] In some embodiments, the structured data includes at least one of a hazard point table, a risk zone table, and a monitoring point table; the unstructured data includes geological disaster reports, which include at least one of the disaster occurrence area, disaster type, prevention and early warning measures, disaster level, and risk level.
[0012] This application provides a dual-risk control device for geological disasters, comprising: Obtain raw data on geological disasters; The raw data on geological hazards are extracted and processed to obtain a geological hazard knowledge graph; Obtain prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and the prompt words to obtain the optimized risk dual control model; Obtain the query statement and input the query statement into the optimized risk dual control model to obtain the graph query statement; The analysis results are obtained based on the map query statement and the geological disaster knowledge map. The analysis results are parsed and processed to obtain the risk dual control results.
[0013] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.
[0014] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.
[0015] This application provides a method, apparatus, equipment, and storage medium for dual risk control of geological disasters. The method involves: first, acquiring raw geological disaster data; extracting and processing the raw data to construct a geological disaster knowledge graph; then, acquiring prompt words, constructing a training dataset based on the knowledge graph, and optimizing the dual risk control model by combining the training dataset and prompt words; next, acquiring a query statement, inputting it into the optimized model to generate a graph query statement; obtaining analysis results based on the graph query statement and the knowledge graph; and finally, parsing the analysis results to obtain the dual risk control results. This approach integrates knowledge graph and model optimization techniques to achieve accurate analysis and efficient management of geological disaster risks, providing reliable technical support for disaster early warning and hazard mitigation. It is applicable to various dual risk control scenarios for geological disasters and addresses the technical problems mentioned in the background section. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A schematic diagram illustrating the implementation process of a dual-risk control method for geological disasters provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation process of obtaining a geological disaster knowledge graph, provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a dual-control device for geological disaster risk provided in an embodiment of this application; Figure 4 The risk management personnel provided in this application embodiment are shown in the map loading and positioning map; Figure 5 The number of geological disaster business Q&A risk areas and map interaction diagram provided for the embodiments of this application; Figure 6 This is a question-and-answer and map interaction diagram for the early warning area provided in the embodiments of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0019] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.
[0020] Figure 1 This is a schematic diagram illustrating the implementation process of a dual-risk control method for geological disasters provided in this application embodiment, including steps 101 to 106. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order of a dual-control method for geological disaster risks. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.
[0021] Step 101: Obtain raw data on geological hazards.
[0022] In this embodiment, the raw geological disaster data is first obtained, which includes structured and unstructured data. The structured data is geological disaster business data in the form of multiple tables, such as tables of hidden danger points, risk areas, and monitoring points. Each table has a primary key (such as hidden danger point ID and risk area ID) and corresponding disaster-related attributes (such as hidden danger point name, risk level, number of people threatened, disaster coordinates, and field survey record number in the hidden danger point table, and risk area name, hidden danger point ID, and warning level in the risk area table). The unstructured data is mainly geological disaster reports, which record information such as the disaster location (such as the specific location of province, city, county, township, and village), disaster type (such as landslide and debris flow), prevention and early warning measures (such as monitoring frequency and emergency evacuation routes), disaster level (such as small, medium, and large), and hazard level.
[0023] Step 102: Extract and process the original geological disaster data to obtain a geological disaster knowledge graph.
[0024] In this embodiment, the unstructured geological disaster report is first divided into blocks according to a preset length, retaining the overlapping context of adjacent blocks during the block division. Then, a dual strategy of semantic extraction and preset dictionary rule matching extraction is used to extract entities and relationships from the block-divided unstructured data. Semantic extraction involves inputting the block-divided document into a risk control model, which is guided by pre-designed prompts to extract entities (e.g., region, disaster, hazard point), entity attributes (e.g., administrative level of the region, time of disaster occurrence), and relationships between entities (e.g., xx county - occurrence - debris flow). Preset dictionary rule matching extraction uses a pre-collected entity matching dictionary (containing place names from provincial to village levels and common geological disaster names) to perform character matching on the block-divided document, identifying successfully matched content as entities. The results of the two extraction strategies are then input into the risk control model, which merges duplicate entities and relationships (e.g., if the same xx county entity is extracted multiple times in different blocks, the risk control model integrates its attribute information to retain a unique entity). Finally, the merged triple data consisting of entities, relationships, and attributes is converted into a tabular file.
[0025] Relationship extraction is performed on structured data in multi-table format. Attribute column names from any two tables are input into the risk dual-control model, which determines whether semantically consistent attribute columns exist (e.g., the hidden danger point ID in the hidden danger point table and the contained hidden danger point ID in the risk area table have semantic consistency). If they exist, implicit relationship edges are added between the two tables (e.g., risk area - contained - hidden danger point). Next, attribute merging is performed. The table names and attribute column names of all tables are input into the risk dual-control model, which extracts common attributes from different tables (e.g., county / district attributes included in both the hidden danger point table and the risk area table). These common attributes are then categorized and merged, and given a unified name (e.g., classifying counties / districts as regional entities). Simultaneously, a relationship is established between the source entity nodes containing the original attributes (e.g., hidden danger point nodes, risk area nodes) and the merged entity nodes (e.g., regional nodes). Finally, the processed relational data and attribute-merged data are converted into graph query statements conforming to the graph database syntax.
[0026] The table files obtained from the above unstructured data processing and the graph query statements obtained from the structured data processing are imported into the graph database. The graph database stores entities, entity attributes, and relationships between entities, and finally constructs a complete geological disaster knowledge graph.
[0027] Step 103: Obtain prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and prompt words to obtain the optimized risk dual control model.
[0028] In this embodiment, the prompts are constructed by combining the requirements of dual control of geological disasters with the features of the geological disaster knowledge graph. They include geological disaster business annotations, a description of the geological disaster knowledge graph structure, task definitions, and problem examples. The geological disaster business annotations are used to adapt to specific geological disaster logic; for example, geographical location requires the province, city, county, township, village, and group. Specific location warning levels are divided into red, orange, yellow, and blue. Disaster scale needs to be labeled as small, medium, large, and extra-large. The geological disaster knowledge graph structure description uses natural language to explain the entity types (such as hazard points, risk areas, monitoring points, regions, and disasters) and the relationships between entities (such as risk area - contains - hazard point / monitoring point - associated - warning level / region - occurrence - disaster). The task definition explicitly requires the dual control risk model to convert the user's input natural language query into a graph query.
[0029] This application collects a dataset of general domain graph query statements, specifically the Text-to-CQL English dataset. The graph structure of the English dataset is then input into a risk dual-control model in a natural language format: entities: [entity1, entity2, ...]; relations: [relation1{entity1, entity2}, relation2{}, ...]. The risk dual-control model translates this into Chinese, simultaneously translating the question-answer pairs (i.e., combinations of natural language queries and corresponding graph query statements) in the dataset, while preserving the core syntax of the graph query statements to obtain a general domain Chinese dataset. Subsequently, using this general domain Chinese dataset as a few-sample example, and combining the entities, relations, and geological disaster business logic of the geological disaster knowledge graph (such as the inclusion relationship between hazard points and risk areas, and the rules for classifying warning levels), the risk dual-control model rewrites the entities and relations in the example to generate geological disaster-specific question-answer pairs. Finally, the general domain Chinese dataset and the geological disaster-specific question-answer pairs are merged to form a training dataset of approximately 15,000 data points.
[0030] Qwen1.5-14B was selected as the basic risk dual-control model. LoRa lightweight technology was used to update some parameters of the model. The training rounds were set to 20, and the sample batch size for each training was 32. AdamW was selected as the model optimization algorithm. The constructed training dataset and the above prompt words were input into the basic risk dual-control model. The model gradually adjusted its parameters by learning the question-answering logic in the training dataset and the business rules in the prompt words. After training, the optimized risk dual-control model was obtained. This model can accurately convert natural language query statements into query statements that conform to graph syntax and understand geological disaster business-specific information.
[0031] Step 104: Obtain the query statement and input it into the optimized risk dual control model to obtain the graph query statement.
[0032] In this embodiment, a geological disaster risk query statement input by the user is obtained. The query statement is in natural language form, such as querying how many risk zones of each level are in Laolong Town, the amount of threatened property at a certain hidden danger point (named "Huang Siguang's house back slope"), and the number of GNSS monitoring points with a red warning level in Longchuan County. The above query statement is input into the optimized risk dual control model. The model performs semantic understanding and format conversion on the natural language query statement based on the preset geological disaster business annotations, map structure descriptions, and task definitions in the prompt words. For example, the query of how many risk zones of each level are in Laolong Town is converted into a map query statement that conforms to the map database syntax, and finally the map query statement is output.
[0033] Step 105: Obtain the analysis results based on the map query statement and the geological disaster knowledge map.
[0034] In this embodiment, the obtained graph query statement is input into a database storing a geological hazard knowledge graph. The database retrieves the corresponding entities, attributes, and relationships from the knowledge graph based on the logic of the query statement. For example, for a graph query statement asking how many risk zones of each level exist in Laolong Town, the database retrieves the risk zone entities associated with the corresponding regional entity of Laolong Town, and classifies and counts the number of risk zone entities according to risk level attributes (e.g., 1 extremely high-risk zone, 33 high-risk zones, 149 medium-risk zones, and 558 low-risk zones). For a graph query statement asking about the amount of threatened property at a certain hazard point, the database retrieves the threatened property attribute value corresponding to the entity at that hazard point (e.g., 203,000 yuan). After the retrieval is completed, the database outputs analysis results containing the retrieved data, presented in a structured form (e.g., key-value pairs containing risk level and corresponding quantity, and key-value pairs containing entity attributes and attribute values).
[0035] Step 106: Analyze the analysis results to obtain the risk dual control results.
[0036] In this embodiment, the structured analysis results output from the database are parsed and converted into an intuitive result format that meets the operational needs of geological disaster risk management. For example, the risk area statistics of Laolong Town are analyzed as follows: 1 extremely high-risk area, 33 high-risk areas, 149 medium-risk areas, and 558 low-risk areas. The number of risk areas of each level in Laolong Town is as follows: 1 extremely high-risk area, 33 high-risk areas, 149 medium-risk areas, and 558 low-risk areas. The "threatened property" attribute value of 203,000 yuan for a certain hidden danger point is analyzed as the amount of threatened property of a certain hidden danger point (named "slope behind Huang Siguang's house") being 203,000 yuan. The final risk control results can be directly used for geological disaster risk classification and management (such as formulating differentiated monitoring plans based on the number of risk areas) and hidden danger investigation and treatment (such as determining treatment priorities based on the amount of threatened property of hidden danger points).
[0037] This application's embodiments construct a comprehensive, end-to-end systematic solution for dual-control of geological disaster risks, covering the entire chain from data acquisition, knowledge graph construction, model optimization, query transformation, result analysis, to dual-control output. It solves the problems of fragmented data processing, fixed query dependencies, and disjointed analytical logic found in traditional methods, achieving closed-loop management from raw data to risk control results. The knowledge graph provides a structured storage medium for data, addressing the pain point of unifying the management of multi-source geological disaster data; the optimized model can accurately understand natural language queries, solving the problems of insufficient flexibility and inability to adapt to complex geological disaster business scenarios in traditional queries, ultimately improving the precision of geological disaster risk identification and the efficiency of control decisions.
[0038] In the above Figure 1 Based on the above, this application also provides a schematic diagram of the implementation process for obtaining a geological disaster knowledge graph, as shown below. Figure 2 As shown, steps 201 to 204 are included: Step 201: Divide the unstructured data into blocks according to a preset length to obtain the processed unstructured data.
[0039] In this embodiment, the unstructured data is a geological disaster report, which includes information such as the disaster location (e.g., Shekeng Village, Hezhen Town, Longchuan County, Heyuan City, Guangdong Province), disaster type (e.g., landslide and debris flow), prevention and early warning measures (e.g., on-site monitoring three times a month), and disaster level (e.g., small-scale disaster). Because the risk dual-control model has a context window limitation and cannot directly process complete long documents, the geological disaster report is divided into blocks according to a preset length (e.g., each block contains 300-500 characters). During block division, overlapping contextual content between adjacent blocks is preserved (e.g., the previous block ends with "A debris flow occurred in Longchuan County in June, threatening 5 people," and the next block begins with the same core information) to avoid incomplete entity information due to block fragmentation, ensuring consistency when extracting entities (e.g., the Longchuan County debris flow), ultimately resulting in multi-segment processed unstructured data.
[0040] Step 202: Perform dual-strategy extraction on the processed unstructured data to obtain a table file.
[0041] In this embodiment, the segmented unstructured data is input into the risk dual-control model block by block, and the model is guided by pre-designed prompts to extract information in a fixed format. The prompts explicitly require the model to identify entities such as regional disaster hazard points from the text, as well as the attributes of the entities (such as the time of disaster occurrence and the number of people threatened) and the relationships between entities (such as region-occurrence-disaster). For example, for the text "A mudslide disaster occurred in Longchuan County in June, threatening 5 people, and the field investigation record number is HS3-YZ040", the risk dual-control model will extract the entities Longchuan County (attribute: region) and mudslide (attribute: disaster type), and the relationship Longchuan County-occurrence-mudslide (attributes: occurrence time June, number of people threatened 5, investigation number HS3-YZ040), forming a semantic extraction result.
[0042] An entity matching dictionary is pre-constructed, consisting of two parts: First, complete place names from the provincial, municipal, county, and village levels collected from the national administrative place name database (e.g., Heshi Town, Longchuan County, Heyuan City, Guangdong Province); second, common disaster names compiled from geological disaster classification standards (e.g., landslides, mudslides, and ground subsidence). The segmented unstructured data is then matched character-by-character with this dictionary. If a word in the text matches a word in the dictionary (e.g., mudslide in Longchuan County), it is identified as an entity and recorded, forming the rule-matching extraction result.
[0043] The semantic extraction results are merged with the rule matching extraction results and input into the risk dual control model for deduplication. The model will identify and merge duplicate entities and relationships (for example, if the same Longchuan County entity is extracted twice in different blocks, the model will integrate all its attribute information to retain a unique entity). Finally, the deduplicated entity-relationship-attribute triple data (such as Longchuan County, occurrence, mudslide, June, 5 people, HS3-YZ040) is converted into a tabular file (such as CSV format). The columns of the table correspond to Entity 1, Relationship Entity 2, Attribute 1 (occurrence time), Attribute 2 (number of people threatened), Attribute 3 (investigation number), etc.
[0044] Step 203: Extract and merge attributes from the structured data to obtain the graph query statement.
[0045] In this embodiment of the application, the structured data is geological disaster business data in the form of multiple tables, including a hidden danger point table, a risk zone table, a monitoring point table, etc. Each table contains a primary key (such as the hidden danger point ID in the hidden danger point table and the risk zone ID in the risk zone table) and attribute columns (such as the name of the hidden danger point in the hidden danger point table, the county where the hidden danger point is located, and the risk zone name in the risk zone table, which includes the hidden danger point ID and the warning level).
[0046] Input the attribute column names of any two tables into the risk dual control model, and the model will determine whether there are semantically consistent attribute columns. For example, after inputting the hazard point ID from the hazard point table and the contained hazard point ID from the risk area table into the model, the model will identify that both columns are unique identifiers pointing to hazard points, and are semantically consistent. Then, it will add an implicit relationship edge (i.e., risk area-contained-hazard point) to the two tables, complete the attribute extraction to clarify the relationship between the tables, and form the relationship extraction result.
[0047] Input the table names and attribute column names of all tables (such as the county where the hazard point table is located, the city to which the risk area table belongs, and the village where the monitoring point table is located) into the risk dual control model. The model will extract the same type of attributes from different tables (such as the county, city, and village being geographical location-related attributes), and classify and merge these attributes and assign them a unified name (such as merging the above attributes into a regional entity). At the same time, the model will establish a relationship between the source entity nodes containing the original attributes (such as hazard point nodes and risk area nodes) and the merged entity nodes (such as regional nodes) (such as hazard point - located in - regional risk area - belonging to - region), forming the attribute merging result.
[0048] The extracted relationship results (inter-table relationship edges) and attribute merging results (merged entities and associations) are converted into graph query statements that conform to the graph database syntax. For example, for the relationship of risk area-containing-potential point, a query statement is generated to establish this relationship in the graph; for the association between a region entity and a source entity, a query statement is generated to define the entity association.
[0049] Step 204: Obtain the geological disaster knowledge graph based on the table file and the graph query statement.
[0050] In this embodiment, the obtained table file is imported into the graph database using a third-party open-source tool. The tool automatically parses the entity, relationship, and attribute information in the table and creates corresponding nodes and edges in the database. At the same time, the graph query statement obtained from the structured data processing is input into the graph database sentence by sentence. The database executes the statement to establish implicit relationships between tables, create merged entities (such as regions), and associates. Finally, the processing results of unstructured data and structured data are integrated in the graph database to form entities including regional disaster hazard points, risk area monitoring points, etc.
[0051] This application's embodiments adapt to the characteristics of unstructured data through block processing, avoiding entity fragmentation caused by the context window limitation of the dual-risk control model; the dual-strategy extraction balances attribute / relationship acquisition and entity accuracy in unstructured data, and attribute extraction and merging solve the problems of no explicit association and attribute redundancy in structured data. The table file obtained from unstructured data processing provides complete entity-relationship-attribute triples, and the graph query statement obtained from structured data processing ensures that the logic between tables can be implemented. The knowledge graph data constructed by combining the two is complete and the relationships are clear.
[0052] In some embodiments, prompt words are obtained, a training dataset is constructed based on the geological disaster knowledge graph, and the risk dual control model is optimized based on the training dataset and prompt words to obtain an optimized risk dual control model. This includes obtaining prompt words based on the geological disaster dual control requirements and the entity information and relationship information of the geological disaster knowledge graph. The prompt words include geological disaster business annotations, graph structure descriptions, task definitions, and problem examples.
[0053] Specifically, the construction of prompts aims to adapt to the business scenario of dual control of geological disaster risks and assist the dual control risk model in understanding business logic. It includes four parts: geological disaster business annotations, map structure descriptions, task definitions, and problem examples.
[0054] The geological hazard operational annotations address the proprietary logic and terminology of the geological hazard field, clearly defining annotation rules to avoid model misunderstandings. For example, in accordance with risk classification and control requirements, the annotations categorize early warning levels into red, orange, yellow, and blue, with red indicating the highest risk level. In accordance with geographic location requirements, the annotations must specify the location at six levels: province, city, county, township, village, and group, such as "Shekeng Village, Heshi Town, Longchuan County, Heyuan City, Guangdong Province." In accordance with disaster scale classification requirements, the annotations must indicate the disaster scale as small, medium, large, or extra-large, based on the number of people threatened and the amount of property damaged (small: number of people threatened < 5 and property damage < 100,000 yuan; medium: 5 ≤ number of people threatened < 20 and 100,000 yuan ≤ property damage < 500,000 yuan, etc.).
[0055] The core entity types and relationships between entities in the geological hazard knowledge graph are clearly described in natural language to ensure that the model understands the logic of the graph data. Entity types include regions (including province, city, county, etc.), hazard points (including hazard point name, ID, etc.), risk zones (including risk level, warning level, etc.), disasters (including disaster type, occurrence time, etc.), and monitoring points (including monitoring frequency, warning status, etc.). Relationships between entities include region-occurrence-disaster risk zone-containment-hazard point monitoring point-monitoring-hazard point region-jurisdiction-risk zone, etc. For example, the entity "risk zone" and the entity "hazard point" are associated through a "containment" relationship; one risk zone can contain multiple hazard points, and each hazard point belongs to only one risk zone.
[0056] Define the core task of the risk dual-control model, that is, receive the natural language query statement input by the user, convert it into a graph query statement that conforms to the syntax of the graph database, and the query statement needs to accurately associate the entities and relationships in the knowledge graph to ensure that the corresponding data can be retrieved.
[0057] Provide sample natural language query-graph query statement pairs for typical query scenarios to assist the model in learning the conversion logic. For example, Sample 1: Natural language query "Query what are the high-risk areas in Dengyun Town", the corresponding graph query statement needs to associate the area (Dengyun Town), the risk area (high-risk level), and the area-jurisdiction-risk area relationship; Sample 2: Natural language query "Query the number of people threatened by the hidden danger point on the slope behind Huang Siguang's house", the corresponding graph query statement needs to associate the hidden danger point (name: the slope behind Huang Siguang's house) entity and the number of threatened people attribute. Integrate the above four parts to obtain the prompt words adapted to this embodiment.
[0058] Furthermore, obtain the training data set according to the query statement data set and the geological disaster question-answer pairs generated based on the geological disaster knowledge graph.
[0059] Specifically, select the Text-to-CQL English data set. This data set contains about 10,000 pieces of data, covering various question types such as entity query, relationship query, and attribute statistics, and the syntax of the graph query statement is comprehensive. Since this data set is in English, it needs to be converted into Chinese first: input the graph structure in the data set in the natural language format of entity: [entity1, entity2,...]; relationship: [relationship1{entity1, entity2}, relationship2{},...] into the risk dual-control model, and the model translates it into Chinese (such as translating Entity:County into entity: county); at the same time, input the natural language query-graph query statement question-answer pairs in the data set into the model, and require the model to retain the core syntax of the graph query statement and only convert the natural language query and entity / relationship names into Chinese, and finally obtain the general domain Chinese query statement data set.
[0060] Using a general-domain Chinese query dataset as a sample example, and combining the entities, relationships, and business logic of the geological disaster knowledge graph, a risk dual-control model generates geological disaster-specific question-and-answer pairs in batches. During generation, the model needs to rewrite the example based on the actual data logic in the knowledge graph. For example, referring to the question-and-answer logic for querying the number of entities in a certain region in the general dataset, and combining the region-jurisdiction-risk area relationship in the geological disaster knowledge graph, a natural language query is generated: query the number of risk areas of each level in Longchuan County; a graph query statement is generated: a question-and-answer pair that associates "region (Longchuan County)", "risk area", and "region-jurisdiction-risk area" relationships, and groups the number of entities by risk level. Referring to the question-and-answer logic for querying the attribute of a certain entity in the general dataset, and combining the hazard point-threatened number attribute in the geological disaster knowledge graph, a natural language query is generated: query the number of people threatened by a certain hazard point (ID: HS3-YZ040); a graph query statement is generated: a question-and-answer pair that retrieves the "number of people threatened" attribute of the entity "hazard point (ID: HS3-YZ040)". Approximately 5000 geological disaster-specific question-and-answer pairs are generated in this way.
[0061] By merging a general-domain Chinese query dataset (approximately 10,000 entries) with a geological disaster-specific question-and-answer pair dataset (approximately 5,000 entries) and removing duplicate data (such as question-and-answer pairs with the same scenario), a training dataset of approximately 15,000 entries was obtained. This dataset contains both the logical patterns of general geographic graph queries and covers the specific scenario of dual control of geological disaster risks.
[0062] Furthermore, the risk dual control model is optimized based on the training dataset and prompt words to obtain the optimized risk dual control model.
[0063] Specifically, comparing the performance of different risk dual-control models in the graph query statement conversion task, Qwen1.5-14B was selected as the base model. This model has strong semantic understanding and format conversion capabilities, and good adaptability to medium-sized training data, which can meet the query conversion needs in geological disaster scenarios.
[0064] To reduce training resource consumption and shorten the training cycle, the LoRa lightweight technique is used to optimize the model. This technique only updates some low-rank matrix parameters of the model without adjusting all model parameters, thus significantly reducing the computational and storage resources required for training while ensuring optimization results.
[0065] The training parameters are configured as follows: the number of training rounds is set to 20 (to ensure that the model fully learns the data patterns while avoiding overfitting), the sample batch size is set to 32 (to balance training efficiency and the amount of data information in a single training session), and the optimization algorithm is AdamW (suitable for training risk dual-control models, which can effectively control the parameter update magnitude).
[0066] The constructed training dataset and prompt words are input into the basic risk dual-control model to start the training process. During training, the model uses the prompt words as business guidance to learn the corresponding logic of natural language queries and graph queries in the training dataset. For example, it understands the meaning of specialized terms such as warning level risk areas through prompt words, and learns the sentence conversion rules for querying specific warning level risk areas in a certain region by combining training data. After training, the model's effectiveness is verified (e.g., by inputting an unfamiliar geological disaster query, verifying whether the converted graph query is accurate). Once the verification is successful, the optimized risk dual-control model is obtained. This model can accurately understand the business logic of geological disasters and convert users' natural language queries into query statements that conform to graph syntax.
[0067] This application's embodiments clarify the model's understanding boundaries by including geological disaster business annotations and map structure descriptions in the prompts, avoiding misunderstandings of geological disaster-specific terms such as deep displacement and tilting in the dual-control system. Task definitions and problem examples provide the model with clear query transformation logic, reducing the risk of format deviation. Combining general domain datasets with geological disaster-specific question-and-answer pairs covers both the basic syntax logic of map queries and adapts to the specific scenarios of geological disaster risk management, avoiding the problems of poor generalization ability or insufficient scenario coverage caused by a single dataset.
[0068] In some embodiments, a dual-strategy extraction process is performed on the processed unstructured data to obtain a tabular file, including: performing semantic extraction processing on the processed unstructured data to obtain a semantic extraction result.
[0069] Specifically, the processed unstructured data consists of geological disaster report fragments divided into blocks of preset length (300-500 characters per block, retaining overlapping contextual content between adjacent blocks). For example, a block might contain the information about a landslide disaster that occurred in Shekeng Village, Hezhen Town, Longchuan County, Heyuan City, Guangdong Province in June 2024. On-site investigation indicated that the hazard threatened 5 people and 203,000 yuan worth of property. The field investigation record number was HS3-YZ040, and monitoring was required twice monthly. Semantic extraction processing is achieved through a risk dual-control model. The specific operation is as follows: First, pre-designed prompts are input into the risk dual-control model. These prompts explicitly require the extraction of three core entities—region, disaster, and hazard point—from the input geological disaster report fragment, along with the attributes of each entity (such as the administrative level of the region, the time of the disaster, and the number of people threatened by the hazard point) and the relationships between entities (such as the relationship between the region and the occurrence of the disaster, and the correlation between the hazard point and the monitoring frequency). The results are then output in a fixed format: Entity 1-Relationship-Entity 2-Attribute List. After inputting the segmented report fragments into the risk control model, the model performs semantic parsing based on prompts, extracting: Entity 1, Shekeng Village, Heshi Town, Longchuan County (attribute: region, administrative level: village), Entity 2, landslide (attribute: disaster type, occurrence time: June 2024), relationship occurrence, as well as Entity 1, hazard point (number HS3-YZ040) (attribute: 5 people threatened, 203,000 yuan of property threatened, investigation number HS3-YZ040), Entity 2, monitoring (attribute: monitoring frequency: twice a month), and relationship association. After processing all segmented data in this way, the semantic extraction results containing multiple sets of entity-relationship-attribute information are obtained.
[0070] Furthermore, character matching and extraction are performed on the processed unstructured data to obtain the rule matching extraction results.
[0071] Specifically, the dictionary contains two core entity data categories. The first is place names of areas prone to geological disasters, selected from the national administrative place name database, covering five levels of administrative units: provincial (e.g., Guangdong Province), municipal (e.g., Heyuan City), county (e.g., Longchuan County), township (e.g., Heshi Town), and village (e.g., Shekeng Village). The second is a compilation of common geological disaster names, such as landslides, mudslides, ground subsidence, and ground fissures. These place names and disaster names are organized into a dictionary according to the format of entity type-entity name, for example, region-Guangdong Province disaster-landslides.
[0072] Each piece of processed unstructured data is compared character by character with an entity matching dictionary. If a word in the report fragment contains a word that exactly matches an entity name in the dictionary (such as "landslide in Longchuan County, Guangdong Province" in the report), the word and its corresponding entity type are recorded. For example, for the aforementioned report fragment of a landslide disaster in Shekeng Village, Hezhen Town, Longchuan County, Heyuan City, Guangdong Province, the matching can extract the following: region - Guangdong Province - Heyuan City - Longchuan County - Hezhen Town - Shekeng Village disaster - landslide.
[0073] The matching results of all data blocks are summarized, and duplicate records are removed to obtain the rule matching extraction results.
[0074] Furthermore, a table file is obtained based on the semantic extraction results and the rule matching extraction results.
[0075] Specifically, the semantic extraction results and rule matching extraction results are input into the risk dual-control model, which performs two processing steps. First, it completes entity attributes: Since the rule matching extraction results only contain entities and types (e.g., region - Longchuan County), the model finds the corresponding entity's (e.g., Longchuan County's) attributes (e.g., administrative level, affiliated city-level unit Heyuan City) from the semantic extraction results and completes them in the rule matching results. Second, it merges duplicate entities: If the same entity (e.g., Longchuan County) is recorded in both types of results, the model integrates all its attributes (e.g., administrative level, associated disaster type), removes duplicate information, and retains a unique and complete entity record. For example, the integrated record for Longchuan County would be: Entity type: Region; Entity name: Longchuan County; Attributes: Administrative level: County; Affiliated city-level: Heyuan City; Associated disaster type: Landslide; Disaster occurrence time: June 2024.
[0076] The integrated entity-relationship-attribute information is organized into a table (such as CSV format) according to preset column names. The table column names include: Entity 1 type, Entity 1 name, Relationship type, Entity 2 type, Entity 2 name, Attribute 1 name, Attribute 1 value, Attribute 2 name, Attribute 2 value, etc. For example, for information about a landslide disaster in Longchuan County, the corresponding row records in the table are: Entity 1 type region, Entity 1 name Longchuan County, Relationship type occurrence, Entity 2 type disaster, Entity 2 name landslide, Attribute 1 name occurrence time, Attribute 1 value June 2024, Attribute 2 name number of people threatened, Attribute 2 value 5 people.
[0077] Fill all the integrated information into the table in this format to obtain a table file containing complete entity-relationship-attribute information.
[0078] This application's embodiments proactively uncover relationships and attributes between entities through semantic extraction, overcoming the limitation of traditional dictionary matching which can only extract entities but not related information. Character matching extraction accurately identifies core entities such as place names and disaster types, avoiding entity recognition biases caused by semantic ambiguity in the risk dual-control model. The combination of these two strategies significantly improves the reliability of unstructured data extraction. By integrating and deduplicating results, duplicate data caused by block processing or dual-strategy extraction is eliminated, generating standardized table files. This ensures that the entity-relationship-attribute triples input to the knowledge graph are non-redundant and formatted uniformly, reducing the error correction cost of subsequent graph construction and improving data utilization efficiency.
[0079] In some embodiments, performing attribute extraction and merging processing on structured data to obtain a graph query statement includes: performing relation extraction processing on structured data to obtain relation extraction results.
[0080] Specifically, the structured data consists of multi-table data in the geological disaster business scenario, including a hazard point table, a risk zone table, and a monitoring point table. The structure and core attributes of each table are as follows: The hazard point table contains attributes such as hazard point ID (primary key), the number of people threatened in the county where the hazard point is located, etc.; the risk zone table contains attributes such as risk zone ID (primary key), the risk zone name including hazard point ID, and the warning level, etc.; the monitoring point table contains attributes such as monitoring point ID (primary key), the monitoring point name associated with the hazard point ID, and the monitoring frequency, etc.
[0081] The risk control model inputs a combination of attribute column names from any two tables. Pre-designed prompts guide the model to determine if semantically consistent attribute columns exist that can be used to establish inter-table relationships. For example, inputting the hazard point ID from the hazard point table and the contained hazard point ID from the risk zone table into the model prompts it to analyze whether the two columns point to the same unique identifier. After semantic judgment, the model determines that both columns correspond to the unique identifiers of hazard points, and are semantically identical. This establishes a risk zone-containing-hazard point relationship edge between the two tables, clarifying that a risk zone can be associated with multiple hazard points, and each hazard point belongs to a risk zone through its contained hazard point ID. Similarly, inputting the hazard point ID from the hazard point table and the associated hazard point ID from the monitoring point table into the model determines that the two columns are semantically consistent, establishing a monitoring point-monitoring-hazard point relationship edge, clarifying the correspondence between monitoring points and hazard points.
[0082] After iterating through all the attribute column name combinations in the above manner and extracting the relationships between all tables, the relationship extraction results are summarized. The results are recorded in the format of Table 1 name-relationship type-Table 2 name-related attribute column, for example, Risk Area Table-Contains-Hidden Danger Point Table-Contains Hidden Danger Point ID & Hidden Danger Point ID Monitoring Point Table-Monitoring-Hidden Danger Point Table-Related Hidden Danger Point ID & Hidden Danger Point ID.
[0083] Furthermore, the structured data is processed by attribute merging to obtain the attribute merging result.
[0084] Specifically, the table name-attribute column name list of all tables is input into the risk dual-control model. The prompts require the model to extract attribute columns describing the same type of information from different tables, categorize them, and assign them a unified entity type name. Simultaneously, the model establishes the association between the original table entities and the merged entities. For example, from the county in the hazard point table, the city in the risk area table, and the village in the monitoring point table, the model identifies three columns that describe geographical location information. These attributes are then categorized and merged, assigning a unified entity type name: "Region." At the same time, the model establishes the association between the original table entities and the region entities: for the hazard point table, a hazard point-location-region association is established, clarifying that each hazard point is associated with the corresponding county-level region entity through its county attribute; for the risk area table, a risk area-belonging-region association is established, clarifying that each risk area is associated with the corresponding city-level region entity through its city attribute; for the monitoring point table, a monitoring point-located-region association is established, clarifying that each monitoring point is associated with the corresponding village-level region entity through its village attribute.
[0085] In addition, the model will merge other similar attributes. For example, it will merge the number of people threatened by the hazard point table and the estimated number of people threatened by the risk area table into a single "number of people threatened" attribute, uniformly categorizing them under the risk information description category. After all attributes are classified, named, and their relationships are established, the attribute merging results are summarized.
[0086] Furthermore, the graph query statement is obtained based on the relationship extraction results and attribute merging results.
[0087] Specifically, the processed data is converted into graph query statements. Combining this with the syntax rules of the graph database, the relationship extraction results are transformed into query statements for creating entity relationships between tables. For example, for the risk area-containing-hazard point relationship extraction results, a statement is generated to create this relationship in the graph. The statement clearly defines the type identifiers of the risk area entity and the hazard point entity, as well as the association logic for attribute matching using the hazard point ID and the hazard point ID. A similar relationship creation statement is generated for the monitoring point-monitoring-hazard point relationship.
[0088] Simultaneously, the attribute merging results are converted into query statements for creating merged entities and their relationships. For example, for the merged results of regional entities, a statement is generated to create regional entity nodes, which includes core attributes such as regional name and administrative level (e.g., province, city, county, village). For the relationship between hazard point - location - region, a statement is generated to establish the association between the hazard point entity and the regional entity, clarifying the association logic and attribute matching rules between the two.
[0089] After all the transformations are completed, a series of graph query statements that conform to the graph database syntax are obtained. These statements can be executed directly in the graph database, realizing the implementation of inter-table relationships and merging attributes in the structured data in the graph.
[0090] This application's embodiments can clearly identify implicit relationships between tables (such as risk area-containment-potential point) through relationship extraction, eliminating the need for manual annotation of foreign keys and resolving the pain point of fragmented data logic across multiple tables. Attribute merging reduces data redundancy, constructs unified entity types, and improves the readability and query efficiency of the knowledge graph. The generated graph query statements conform to the graph database syntax, ensuring that the results of structured data processing can be directly imported into the knowledge graph without additional format conversion, lowering the technical implementation threshold. At the same time, the clear relationships between tables and the unified entity types also provide clear logic for subsequent graph queries, reducing syntax errors and logical deviations in query statements.
[0091] In some embodiments, after parsing the analysis results to obtain the risk dual control results, the method further includes: obtaining feedback processing results, optimizing the parameters of the geological hazard knowledge graph and the optimized risk dual control model based on the feedback processing results, and obtaining the optimized geological hazard knowledge graph and parameters. The feedback processing results are the user feedback operations in the analysis results.
[0092] Specifically, the interactive interface of the geological disaster risk dual-control platform includes three feedback buttons: like, dislike, and refresh. After viewing the analysis results, users can choose an action based on the accuracy and readability of the results. A like indicates that the analysis results are accurate and clearly expressed; a dislike indicates that the data is correct but there are grammatical or logical issues; and a refresh indicates that the analysis results are incorrect or do not meet the query requirements. The system records the user's action type and the corresponding analysis results, generating feedback processing results.
[0093] If the feedback is a thumbs up: there is no need to adjust the parameters of the geological disaster knowledge graph and the risk dual control model. The query scenario and processing logic corresponding to the analysis result will be marked as a high-quality case for reference in subsequent model iterations.
[0094] If the feedback is a negative comment: Address the language logic issue by optimizing the parameters of the risk dual-control model. For example, adjust the statement organization logic of the model's output layer, or supplement the rules for concise prompts, and retrain the model to correct the output expression. If the feedback mentions vague attribute descriptions in the geological disaster knowledge graph, supplement or correct the attribute information of the corresponding entities in the geological disaster knowledge graph.
[0095] If the feedback is a refresh: First, check and analyze the reasons for the error in the results. If the error is due to a logical deviation in the query statement of the risk dual-control model, then use CoT (Chain-of-Thought) and step-by-step prompting techniques to adjust the model's statement conversion parameters. For example, let the model parse the query requirements step by step (first determine that the query area is Longchuan County, and then filter the risk areas with a red warning level). If the error is due to data errors in the geological disaster knowledge graph (such as incorrect entry of the risk area warning level), then correct the attribute values of the corresponding entities in the graph, and finally obtain the optimized geological disaster knowledge graph and risk dual-control model parameters.
[0096] Furthermore, the original geological disaster data is cleaned to obtain processed original geological disaster data.
[0097] Specifically, for structured data such as hazard point tables and risk zone tables, three main aspects are addressed. First, duplicate data is removed, such as deleting multiple duplicate records for the same hazard point (filtered by the unique identifier of the hazard point ID). Second, erroneous data is corrected, such as uniformly correcting fifty people entered in the threat number field to 50 people, and correcting incorrectly written red in the warning level field to red. Third, missing data is supplemented, such as supplementing the corresponding latitude and longitude coordinates for records with missing hazard point coordinates, based on the description in the unstructured geological disaster report that the hazard point is located 500 meters east of XX village.
[0098] For unstructured data such as geological disaster reports, two main aspects are addressed. First, irrelevant information is removed, such as deleting administrative notices from the report that are unrelated to geological disaster risks. Second, the format of expression is standardized, such as correcting different time statements in the report, such as stating "debris flow disaster occurred in June 2024" as "debris flow disaster occurred in June 2024," to ensure consistency in subsequent data extraction.
[0099] After the above cleaning is completed, the raw geological disaster data is obtained with uniform format, accurate data, and no redundancy.
[0100] Furthermore, a graph space is constructed based on the geological disaster knowledge graph to obtain a dedicated graph space.
[0101] Specifically, a map space creation request is initiated in the map database. Based on the region or business scenario of geological disaster risk management, the map space is named (e.g., the dedicated map space for dual control of geological disaster risk in Longchuan County), and the storage parameters of the map space (e.g., the number of data shards and the number of replicas) are configured to ensure the security of data storage and access efficiency.
[0102] The entities (such as regional hazard points and risk areas), relationships (such as occurrence and inclusion), and attributes (such as warning level and number of people threatened) in the optimized geological disaster knowledge graph are fully mapped to a dedicated graph space, enabling the graph space to possess the core data logic of the knowledge graph. At the same time, access permissions for the graph space are set, allowing only authorized users of the geological disaster risk dual control platform to operate the data, thus ensuring data security.
[0103] Through the above operations, a dedicated graph space is obtained that is logically consistent with the geological disaster knowledge graph and has independent data. This graph space can be used specifically to store the processed original geological disaster data and the results of dual risk control.
[0104] Furthermore, the processed original geological disaster data, dedicated map space, and risk control results are loaded onto the map.
[0105] Specifically, two types of standards are established. First, data format standards: The processed raw geological disaster data (such as coordinates of hazard points and the extent of risk zones), entity attributes in the dedicated map space (such as risk zone warning levels and the number of people threatened by hazard points), and risk control results (such as the number of risk zones at each level) are uniformly converted into a format recognizable by the map front-end (such as JSON format). Coordinate data must conform to the map's latitude and longitude coordinate system. Second, interface interaction standards: A data transmission interface is designed using logic similar to an ask-and-answer interface to ensure that when the back-end transmits data to the map front-end, it can accurately return the data type, data content, and data identifier (such as risk zone data + 32 red warning zones + Longchuan County area).
[0106] The backend transmits the processed raw geological disaster data, related data in the dedicated map space (such as the inclusion relationship between hazard points and risk areas), and risk dual control results to the map frontend via an interface, according to a unified standard. After receiving the data, the frontend parses it. For example, when parsing risk area data, different colors are assigned to different risk areas according to the warning level (red corresponds to red warning, orange corresponds to orange warning), and the geographical location of the risk area is marked on the map; when parsing hazard point data, hazard point icons are marked at the corresponding coordinates on the map, and a floating prompt message with the hazard point name and the number of people threatened is displayed.
[0107] Ultimately, the processed raw geological hazard data, dedicated geographic spatial data, and risk control results are displayed in a visualized, interconnected manner on a map. Users can intuitively view the risk distribution, hazard location, and risk level in different areas, such as... Figure 4-6 As shown, Figure 4 The risk management personnel provided in this application embodiment are shown in the map loading and positioning map. Figure 5 This application provides an example of a Q&A platform for geological disaster risk zones and an interactive map. Figure 6 This is a question-and-answer and map interaction diagram for the early warning area provided in the embodiments of this application.
[0108] This application's embodiments optimize knowledge graph data and model parameters based on user feedback, overcoming the limitations of traditional methods that lack feedback mechanisms and cannot adapt models and graphs to business iterations, thus improving the system's long-term accuracy and adaptability. Data cleaning ensures that subsequent geological disaster data is formatted uniformly and accurately; a dedicated graph space enables independent storage of geological disaster dual-control data, avoiding confusion with other business data, while also supporting access control to ensure data security. By loading data and results onto a map, users can intuitively view information such as risk zone distribution, hazard point locations, and warning levels, solving the problems of scattered information and difficulty in quickly locating risks in traditional text results. This provides intuitive decision support for geological disaster emergency response and tiered management, improving management efficiency.
[0109] While this application provides method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the method can be executed sequentially according to this embodiment or the accompanying drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0110] like Figure 3 As shown in the illustration, this application also provides a dual-control device 300 for geological disaster risks. The device includes: Module 301 is used to acquire raw data on geological hazards; Processing module 302 is used to extract and process raw geological disaster data to obtain a geological disaster knowledge graph; The acquisition module 301 is also used to acquire prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and prompt words to obtain the optimized risk dual control model; The acquisition module 301 is also used to acquire query statements, input the query statements into the optimized risk dual control model, and obtain the graph query statement; Processing module 302 is also used to obtain analysis results based on the map query statement and the geological disaster knowledge map; The parsing module 303 is used to parse and process the analysis results to obtain the risk dual control results.
[0111] In some embodiments, the processing module 302 is further configured to divide the unstructured data into blocks according to a preset length to obtain processed unstructured data; The processing module 302 is also used to perform dual-strategy extraction processing on the processed unstructured data to obtain a table file. The dual-strategy extraction includes semantic extraction and preset dictionary rule matching extraction. Processing module 302 is also used to extract and merge attributes from structured data to obtain a graph query statement; The processing module 302 is also used to obtain a geological disaster knowledge graph based on the table file and the graph query statement.
[0112] In some embodiments, the acquisition module 301 is further configured to obtain prompt words based on the dual control requirements of geological disasters and the entity information and relationship information of the geological disaster knowledge graph. The prompt words include geological disaster business annotations, graph structure descriptions, task definitions, and problem examples. The acquisition module 301 is also used to obtain the training dataset based on the query statement dataset and the geological disaster question-and-answer pairs generated based on the geological disaster knowledge graph; The processing module 302 is also used to optimize the risk dual control model based on the training dataset and prompt words to obtain the optimized risk dual control model.
[0113] In some embodiments, the processing module 302 is further configured to perform semantic extraction processing on the processed unstructured data to obtain semantic extraction results; The processing module 302 is also used to perform character matching and extraction processing on the processed unstructured data to obtain the rule matching extraction result; The acquisition module 301 is also used to obtain a table file based on the semantic extraction results and the rule matching extraction results.
[0114] In some embodiments, the processing module 302 is further configured to perform relation extraction processing on the structured data to obtain relation extraction results; Processing module 302 is also used to perform attribute merging processing on structured data to obtain attribute merging results; The processing module 302 is also used to obtain a graph query statement based on the relationship extraction results and attribute merging results.
[0115] In some embodiments, the acquisition module 301 is further configured to acquire feedback processing results, optimize the parameters of the geological disaster knowledge graph and the optimized risk dual control model based on the feedback processing results, and obtain the optimized geological disaster knowledge graph and parameters. The feedback processing results are the user feedback operations in the analysis results. The processing module 302 is also used to clean the original geological disaster data to obtain the processed original geological disaster data; Processing module 302 is also used to construct a graph space based on the geological disaster knowledge graph to obtain a dedicated graph space; The processing module 302 is also used to load the processed geological disaster raw data, dedicated map space, and risk dual control results onto the map.
[0116] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0117] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0118] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module implementing the method or a structure within a hardware component.
[0119] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.
[0120] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.
[0121] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.
[0122] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.
[0123] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0124] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0125] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.
Claims
1. A method for dual control of geological disaster risks, characterized in that, include: Obtain raw data on geological disasters; The raw data on geological hazards are extracted and processed to obtain a geological hazard knowledge graph; Obtain prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and the prompt words to obtain the optimized risk dual control model; Obtain the query statement and input the query statement into the optimized risk dual control model to obtain the graph query statement; The analysis results are obtained based on the map query statement and the geological disaster knowledge map. The analysis results are parsed and processed to obtain the risk dual control results.
2. The method according to claim 1, characterized in that, The raw geological disaster data includes structured data and unstructured data. The extraction and processing of the raw geological disaster data to obtain a geological disaster knowledge graph includes: The unstructured data is divided into blocks according to a preset length to obtain processed unstructured data; The processed unstructured data is subjected to dual-strategy extraction to obtain a table file. The dual-strategy extraction includes semantic extraction and preset dictionary rule matching extraction. The structured data is processed by attribute extraction and merging to obtain a graph query statement; The geological disaster knowledge graph is obtained based on the table file and the graph query statement.
3. The method according to claim 1, characterized in that, The process of obtaining prompt words involves constructing a training dataset based on the geological disaster knowledge graph, and optimizing the risk dual-control model based on the training dataset and the prompt words to obtain an optimized risk dual-control model, including: Based on the dual control requirements for geological disasters and the entity and relation information of the geological disaster knowledge graph, the prompt words are obtained. The prompt words include geological disaster business annotations, graph structure descriptions, task definitions, and problem examples. The training dataset is obtained from the query statement dataset and the geological disaster question-answer pairs generated based on the geological disaster knowledge graph; The risk dual control model is optimized based on the training dataset and the prompt words to obtain the optimized risk dual control model.
4. The method according to claim 2, characterized in that, The process of extracting unstructured data using a dual-strategy method to obtain a table file includes: Semantic extraction is performed on the processed unstructured data to obtain semantic extraction results; The processed unstructured data is subjected to character matching and extraction to obtain the rule matching extraction results; A table file is obtained based on the semantic extraction results and the rule matching extraction results.
5. The method according to claim 1, characterized in that, The process of extracting and merging attributes from the structured data to obtain a graph query statement includes: The structured data is subjected to relation extraction processing to obtain relation extraction results; The structured data is subjected to attribute merging processing to obtain the attribute merging result; The graph query statement is obtained based on the relationship extraction results and attribute merging results.
6. The method according to claim 1, characterized in that, After parsing the analysis results to obtain the risk dual control results, the process further includes: Obtain the feedback processing result, and optimize the parameters of the geological disaster knowledge graph and the optimized risk dual control model based on the feedback processing result to obtain the optimized geological disaster knowledge graph and parameters. The feedback processing result is the user feedback operation in the analysis result. The raw geological disaster data is cleaned to obtain processed raw geological disaster data; A graph space is constructed based on the geological disaster knowledge graph to obtain a dedicated graph space; The processed original geological disaster data, the dedicated map space, and the risk dual control results are loaded into the map.
7. The method according to claim 1, characterized in that, The structured data includes at least one of a hazard point table, a risk zone table, and a monitoring point table; the unstructured data includes geological disaster reports, which include at least one of the following: disaster location, disaster type, prevention and early warning measures, disaster level, and hazard level.
8. A dual-risk control device for geological disasters, characterized in that, include: The acquisition module is used to acquire raw data on geological disasters. The processing module is used to extract and process the raw geological disaster data to obtain a geological disaster knowledge graph; The acquisition module is also used to acquire prompt words, construct a training dataset based on the geological disaster knowledge graph, and optimize the risk dual control model according to the training dataset and the prompt words to obtain the optimized risk dual control model. The acquisition module is also used to acquire a query statement, input the query statement into the optimized risk dual control model, and obtain a graph query statement; The processing module is also used to obtain analysis results based on the map query statement and the geological disaster knowledge map; The parsing module is used to parse and process the analysis results to obtain the risk dual control results.
9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.