Water conservancy scene entity relation extraction method based on large model fine tuning

CN122817482APending Publication Date: 2026-09-25AEROSPACE CLOUD SPACE SPACE INFORMATION TECHNOLOGY (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010445.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]在水利业务场景中,常采用通过自然语言提问方式查询水利设施(如河流、水库)的属性及关系,例如“新开河-金钟河流向哪里?”;水利行业知识图谱构建需要大量结构化数据支撑,而高质量的命名实体识别(NER)语料是构建知识图谱的基础;水利行业具有专业性强、实体关系复杂、专业术语多的特点;而通用语料无法满足专业需求传统实体关系抽取方法主要依赖预定义规则或统计模型,但存在泛化能力差、依赖大量标注数据等问题

Benefits of technology

[0012]与现有技术相比,本发明的有益效果是:针对水利业务场景进行了定制化的数据标注与模型微调,降低了数据量与计算成本,且直接基于用户问题文本生成实体关系,通过低资源依赖的LoRA技术,实现了轻量化的模型和推理引擎的高性能低成本部署;具有更良好的信息适应性,以及交底的推广部署成本占用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817482A_ABST
    Figure CN122817482A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of natural language processing and knowledge graph construction, and discloses a water conservancy scene entity relationship extraction method based on large model fine tuning, which comprises the following steps: data collection and labeling, data preprocessing, model quantization fine tuning, model merging and deployment training, and terminal service and feedback; the water conservancy business scene is customized for data labeling and model fine tuning, the data quantity and the calculation cost are reduced, entity relationship is directly generated based on user question text, the LoRA technology with low resource dependence is used, the light-weight model and the high-performance low-cost deployment of the reasoning engine are realized, the information adaptability is better, and the cost occupancy is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of natural language processing and knowledge graph construction technology, specifically a method for extracting entity relationships in water conservancy scenarios based on large-scale model fine-tuning. Background Technology

[0002] In water conservancy business scenarios, it is common to use natural language queries to query the attributes and relationships of water conservancy facilities (such as rivers and reservoirs), for example, "Where does the Xinkai River-Jinzhong River flow?". The construction of knowledge graphs in the water conservancy industry requires a large amount of structured data support, and high-quality named entity recognition (NER) corpora are the foundation for building knowledge graphs. The water conservancy industry is characterized by strong professionalism, complex entity relationships, and a large number of professional terms. However, general corpora cannot meet the professional needs. Traditional entity relationship extraction methods mainly rely on predefined rules or statistical models, but they have problems such as poor generalization ability and dependence on a large amount of labeled data.

[0003] In recent years, deep learning technology has become mainstream and widely used in various fields. In water conservancy business scenarios, the commonly used technical solutions include: LSTM+CRF SPO triple extraction: entities are identified through sequence labeling models, and relations are extracted through relation classification models, but the text must contain complete SPO (subject-predicate-object) triple information; CASREL layered SPO extraction: entities and relations are extracted simultaneously using a hierarchical structure, but the input text still requires complete triples and high data labeling quality. These methods have certain effects in general domains, but when applied to water conservancy scenarios, they lack semantic understanding of water conservancy terminology (such as "basin" and "water system"), which easily leads to mis-extraction. Furthermore, because user questions are usually short and incomplete (such as only mentioning the subject without specifying the object), the extraction effect of traditional methods is not good. Summary of the Invention

[0004] The purpose of this invention is to provide a method for extracting entity relationships in water conservancy scenarios based on large model fine-tuning, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for extracting entity relationships from water conservancy scenarios based on large model fine-tuning includes the following steps: Data collection and annotation: Collect user question texts and water conservancy business data in the water conservancy field, and annotate the entities, types and relationships of user question texts and water conservancy business data to obtain annotated data and build a training dataset; Data preprocessing converts the labeled dataset into structured information required for instruction fine-tuning, the structured information including instruction fields, input fields, and output fields; Model quantization fine-tuning involves loading a pre-trained large language base model and injecting trainable low-rank matrix parameters into the base model based on low-rank adaptation LoRA technology to obtain the fine-tuned LoRA adapter parameters. Model merging and deployment training: The LoRA adapter parameters are merged with the base model to generate the final model, and the model is deployed based on the vLLM inference engine; Terminal services and feedback involve obtaining users' natural language questions online, loading the final model to perform information matching and reasoning on the database, and extracting and outputting structured entity and relation extraction results.

[0006] As a further aspect of the present invention: the water conservancy business data includes water conservancy project planning documents, hydrological monitoring reports, and flood control and drought relief work records; The user question text includes user question-and-answer dialogue records, and the user question text is used to obtain the demand feature points of structured information.

[0007] As a further aspect of the present invention: in the step of annotating entities, types, and relationships: The annotation format is a JSON object, which contains an array of entities and an array of relationships. In the entity array, each element includes an entity name and an entity type, represented as: {entities: [entity: "entity name", type: "entity type"], [entity: "entity name", type: "entity type"]}; In the relation array, each element includes a relation type and a corresponding head entity and tail entity, represented as: {relation: "relation type", head: "head entity", tail: "tail entity"}.

[0008] As a further aspect of the present invention: the data preprocessing step includes: The labeled data is converted into the Alpaca instruction format supported by the framework, and each sample is converted into a three-field JSON object containing extracted entities and relations, user question text, and labeled data, specifically represented as: {instruction: "Extract entities and relations", input: "User question text", output: "Label data"}.

[0009] As a further embodiment of the present invention: in the step of model quantization fine-tuning: The original parameters of the base model are frozen, and trainable low-rank matrix parameters are injected into the query matrix and value matrix of the self-attention module based on the low-rank adaptive LoRA technique. The low-rank matrix constrains the weight update amount to a low-dimensional subspace through matrix decomposition, and sets the rank to 8, the scaling factor to 16, and the dropout rate to 0.05. The injected low-rank parameters are optimized using an autoregressive cross-entropy loss function to map the water conservancy natural language problem into a structured entity relation output, thus obtaining the LoRA adapter parameters.

[0010] As a further aspect of the present invention: the training dataset includes a training set and a validation set, wherein the training set and the validation set are mutually exclusive, and the method further includes the following steps: The final model is input using the validation set, and the fine-tuned final model is evaluated to obtain validation evaluation results. Based on the verification and evaluation results, the accuracy of entity recognition is calculated by matching and comparing with the labeled true values, and it is determined whether the recognition accuracy of the final model meets the preset standard.

[0011] As a further aspect of the present invention: the base model is the Qwen2.5-1.5B-Instruct large language model, which has a small parameter size and is compatible with the vLLM inference engine.

[0012] Compared with existing technologies, the beneficial effects of this invention are: customized data annotation and model fine-tuning for water conservancy business scenarios, reducing data volume and computing costs; and direct generation of entity relationships based on user question text, achieving high-performance and low-cost deployment of lightweight models and inference engines through low-resource-dependent LoRA technology; it has better information adaptability and lower promotion and deployment costs. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the basic steps of a method for extracting entity relationships in a water conservancy scenario based on fine-tuning of a large model.

[0014] Figure 2 This is a flowchart illustrating the logic of a method for extracting entity relationships in a water conservancy scenario based on fine-tuning of a large model. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0016] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0017] like Figure 1 , Figure 2 The figure shows a method for extracting entity relationships in a water conservancy scene based on large model fine-tuning, provided by an embodiment of the present invention, including the following steps: S10, Data Collection and Labeling: Collect user question texts and water conservancy business data in the water conservancy field, and label the entities, types and relationships of user question texts and water conservancy business data to obtain labeled data and establish a training dataset. S20, Data preprocessing, converting the labeled dataset into structured information required for instruction fine-tuning, the structured information including instruction fields, input fields and output fields; S30, Model Quantization Fine-tuning: Load the pre-trained large language base model, and inject trainable low-rank matrix parameters into the base model based on low-rank adaptation LoRA technology to obtain the fine-tuned LoRA adapter parameters. S40, Model merging and deployment training: Merge the LoRA adapter parameters with the base model to generate the final model, and deploy the model based on the vLLM inference engine; S50, Terminal Services and Feedback, obtains users' natural language questions online, loads the final model to perform information matching and reasoning on the database, and extracts and outputs structured entity and relation extraction results.

[0018] This embodiment presents a method for extracting entity relationships in water conservancy scenarios based on large-scale model fine-tuning. It includes two parts: offline training and online service. Through a systematic generation strategy, and based on a large amount of real-world water conservancy business data and user question text, the model is professionally guided to achieve the construction of a knowledge graph for the water conservancy industry and the establishment of a relationship extraction model. Furthermore, the design of the corpus fully incorporates multiple real-world scenarios in the water conservancy industry, achieving a professional response to industry needs.

[0019] Compared to existing model technologies, this embodiment features customized data annotation and model fine-tuning for water conservancy business scenarios, reducing data volume and computational costs. It also directly generates entity relationships based on user question text and achieves high-performance, low-cost deployment of lightweight models and inference engines through the low-resource-dependent LoRA technology. It has better information adaptability and lower promotion and deployment costs.

[0020] In another preferred embodiment of the present invention, the water conservancy business data includes water conservancy project planning documents, hydrological monitoring reports, and flood control and drought relief work records; The user question text includes user question-and-answer dialogue records, and the user question text is used to obtain the demand feature points of structured information.

[0021] Furthermore, in the step of annotating entities, types, and relationships: The annotation format is a JSON object, which contains an array of entities and an array of relationships. In the entity array, each element includes an entity name and an entity type, represented as: {entities: [entity: "entity name", type: "entity type"], [entity: "entity name", type: "entity type"]}; In the relation array, each element includes a relation type and a corresponding head entity and tail entity, represented as: {relation: "relation type", head: "head entity", tail: "tail entity"}.

[0022] Furthermore, the data preprocessing step includes: The labeled data is converted into the Alpaca instruction format supported by the framework, and each sample is converted into a three-field JSON object containing extracted entities and relations, user question text, and labeled data, specifically represented as: {instruction: "Extract entities and relations", input: "User question text", output: "Label data"}.

[0023] In this embodiment, user question texts in the water conservancy field are collected, covering entity types such as river systems, rivers, reservoirs, monitoring stations, locations, and water conservancy facilities, as well as various water conservancy professional relationships such as inclusion relationships, location relationships, watershed relationships, hydrological relationships, and management relationships. Specifically, inclusion relationships are: river system → containing rivers; location relationships are: river → administrative division; watershed relationships are: entity → watershed; hydrological relationships are: upstream river, downstream river, inflow river; and management relationships are: personnel → affiliated unit. Each text is annotated by professional annotators, and the annotation format is a JSON object containing entities and relations fields. For example, the annotation result for the question "Where does the Xinkai River-Jinzhong River flow?" is: {entities: [entity: "Xinkaihe", type: "river"], [entity: "Jinzhonghe", type: "river"]}; {relation: "flowing into the river", head: "Xinkai River", tail: "Jinzhong River"}; Construct a training dataset of the required size (e.g., 1000 records) in the above manner, and then divide it into a training set (e.g., 800 records) and a validation set (e.g., 200 records, i.e., an 8:2 split with the training set).

[0024] The LoRA (Low-Rank Adaptation) technique is used for quantization fine-tuning, updating only some parameters, which significantly reduces the computational resource requirements. The fine-tuning goal is to minimize the loss function of entity relation extraction. The model performance is evaluated through a validation set. Then, the LoRA adapter parameters are merged with the base model to generate the final deployable model. As for the entity relation extraction mechanism, its principle is that the model learns based on instructions, maps user questions into structured outputs, and captures the semantic association between entities and relations through an attention mechanism.

[0025] In another preferred embodiment of the present invention, the model quantization fine-tuning step includes: The original parameters of the base model are frozen, and trainable low-rank matrix parameters are injected into the query matrix and value matrix of the self-attention module based on the low-rank adaptive LoRA technique. The low-rank matrix constrains the weight update amount to a low-dimensional subspace through matrix decomposition, and sets the rank to 8, the scaling factor to 16, and the dropout rate to 0.05. The injected low-rank parameters are optimized using an autoregressive cross-entropy loss function to map the water conservancy natural language problem into a structured entity relation output, thus obtaining the LoRA adapter parameters.

[0026] Furthermore, the training dataset includes a training set and a validation set, wherein the training set and the validation set are disjoint, and the method further includes the following steps: The final model is input using the validation set, and the fine-tuned final model is evaluated to obtain validation evaluation results. Based on the verification and evaluation results, the accuracy of entity recognition is calculated by matching and comparing with the labeled true values, and it is determined whether the recognition accuracy of the final model meets the preset standard.

[0027] Furthermore, the base model uses the Qwen2.5-1.5B-Instruct large language model, which has a small parameter size and is compatible with the vLLM inference engine.

[0028] In this embodiment, the pre-trained Qwen2.5-1.5B-Instruct large language model is loaded as the base model. The low-rank matrix constrains the weight update amount within a low-dimensional subspace through matrix factorization, and the rank is set to 8, so that the number of trainable parameters is only about 1.1 million, accounting for 0.07% of the total parameters of the base model. The converted Alpaca format training samples are used, and the injected low-rank parameters are optimized using the autoregressive cross-entropy loss function, so that the model learns to map the water conservancy natural language problem into a structured entity relation output. After training, the LoRA adapter parameters are obtained. This adapter carries the extraction capability of the water conservancy domain with minimal storage overhead. Among them, the Qwen2.5-1.5B-Instruct model has a small number of parameters and high inference efficiency, which is suitable for lightweight deployment in water conservancy scenarios.

[0029] Based on the above embodiments, the specific process of entity relationship extraction in water conservancy scenarios can be summarized as follows: Data collection phase: Systematically collect various water conservancy business-related data, including but not limited to professional documents such as water conservancy project planning documents, hydrological monitoring reports, flood control and drought relief work records, as well as raw text data such as user Q&A dialogue records accumulated by the water conservancy consulting service platform.

[0030] Data annotation work: Based on the predefined entity type system (such as "river", "reservoir", "hydropower station" etc.) and relation type system (such as "belonging basin", "inflowing river", "dam site location" etc.), the collected text data is annotated with fine granularity to form a training dataset with a high degree of structure and reliable quality.

[0031] Data preprocessing stage: The labeled structured data is converted into the Alpaca instruction format supported by the LLamaFactory framework. The specific operations include: constructing each labeled sample into a triple data structure containing "instruction-input-output" to ensure that the data format meets the requirements of model training.

[0032] Base model loading: Initialize the pre-trained large language model Qwen2.5-1.5B-Instruct. This model has 1.5 billion parameters and performs well in general domains, making it suitable as a base model for knowledge extraction tasks in the water conservancy field.

[0033] LoRA parameter configuration: Carefully set the key hyperparameters for LoRA fine-tuning, including but not limited to: rank (8), scaling factor (alpha) (16), dropout rate (0.05), etc. These parameters can be used to precisely control the number of trainable parameters and the fine-tuning effect.

[0034] Fine-tuning training execution: Start the training process on the prepared training dataset, and use LoRA technology to update only a small number of new adapter parameters in the model, while freezing most of the original parameters of the base model, so as to achieve efficient parameter fine-tuning.

[0035] Model Validation and Evaluation: The fine-tuned model is comprehensively evaluated using an independently partitioned validation dataset (which was not used in the training process), with a focus on key metrics such as the F1 score for entity extraction and the accuracy for relation extraction.

[0036] Performance benchmark assessment: Set strict performance threshold standards. If the model's performance on the validation set reaches or exceeds the predetermined threshold (e.g., entity F1 > 0.85), proceed to the next stage; otherwise, return to adjust the LoRA hyperparameters or supplement training data and retrain.

[0037] Model merging operation: The parameters of the trained LoRA adapter are merged with those of the original Qwen2 pedestal model to generate a complete and independent final model file. This file contains all model parameters, which is convenient for subsequent deployment and application.

[0038] Model saving steps: Save the merged complete model to the disk storage system in a standard format, and record complete training metadata information, including hyperparameter configuration, training logs, performance metrics, etc.

[0039] Training process complete / deployment ready: The model training phase is now complete, and the generated final model is ready for deployment and can be integrated into a practical water conservancy knowledge graph construction system for use.

[0040] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0041] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the disclosure in the specification and embodiments. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0042] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for extracting entity relationships in water conservancy scenarios based on large model fine-tuning, characterized in that, Includes the following steps: Data collection and annotation: Collect user question texts and water conservancy business data in the water conservancy field, and annotate the entities, types and relationships of user question texts and water conservancy business data to obtain annotated data and build a training dataset; Data preprocessing converts the labeled dataset into structured information required for instruction fine-tuning, the structured information including instruction fields, input fields, and output fields; Model quantization fine-tuning involves loading a pre-trained large language base model and injecting trainable low-rank matrix parameters into the base model based on low-rank adaptation LoRA technology to obtain the fine-tuned LoRA adapter parameters. Model merging and deployment training: The LoRA adapter parameters are merged with the base model to generate the final model, and the model is deployed based on the vLLM inference engine; Terminal services and feedback involve obtaining users' natural language questions online, loading the final model to perform information matching and reasoning on the database, and extracting and outputting structured entity and relation extraction results.

2. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning as described in claim 1, characterized in that, The water conservancy business data includes water conservancy project planning documents, hydrological monitoring reports, and flood control and drought relief work records; The user question text includes user question-and-answer dialogue records, and the user question text is used to obtain the demand feature points of structured information.

3. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning according to claim 2, characterized in that, In the steps of annotating entities, types, and relationships: The annotation format is a JSON object, which contains an array of entities and an array of relationships. In the entity array, each element includes an entity name and an entity type, represented as: {entities: [entity: "entity name", type: "entity type"], [entity: "entity name", type: "entity type"]}; In the relation array, each element includes the relation type and the corresponding head entity and tail entity, represented as: {relation: "relation type", head: "head entity", tail: "tail entity"}.

4. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning according to claim 3, characterized in that, The data preprocessing step includes: The labeled data is converted into the Alpaca instruction format supported by the framework, and each sample is converted into a three-field JSON object containing extracted entities and relations, user question text, and labeled data, specifically represented as: {instruction: "Extract entities and relations", input: "User question text", output: "Label data"}.

5. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning according to claim 1, characterized in that, In the model quantization fine-tuning steps: The original parameters of the base model are frozen, and trainable low-rank matrix parameters are injected into the query matrix and value matrix of the self-attention module based on the low-rank adaptive LoRA technique. The low-rank matrix constrains the weight update amount to a low-dimensional subspace through matrix decomposition, and sets the rank to 8, the scaling factor to 16, and the dropout rate to 0.

05. The injected low-rank parameters are optimized using an autoregressive cross-entropy loss function to map the water conservancy natural language problem into a structured entity relation output, thus obtaining the LoRA adapter parameters.

6. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning according to claim 5, characterized in that, The training dataset includes a training set and a validation set, wherein the training set and the validation set are mutually exclusive, and the method further includes the following steps: The final model is input using the validation set, and the fine-tuned final model is evaluated to obtain validation evaluation results. Based on the verification and evaluation results, the accuracy of entity recognition is calculated by matching and comparing with the labeled true values, and it is determined whether the recognition accuracy of the final model meets the preset standard.

7. The method for extracting entity relationships in a water conservancy scene based on large model fine-tuning according to claim 6, characterized in that, The base model uses the Qwen2.5-1.5B-Instruct large language model, which has a small parameter size and is compatible with the vLLM inference engine.