Water and soil conservation large language model data set construction method

The construction of professional data sets of soil and water conservation through large language models solves the problems of data format diversity, low sorting efficiency and difficulty in updating, and realizes efficient and intelligent data management and application, covers a variety of professional scenarios of soil and water conservation, and provides high-quality Q&A support.

CN120338061APending Publication Date: 2025-07-18NORTHWEST A & F UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179720.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems with data format diversity in the field of soil and water conservation, low manual sorting efficiency, insufficient coverage of professional scenarios and difficulty in updating and maintaining, and the existing data sets cannot meet the needs of rapid development.

Method used

A large language model is used to build a professional data set for soil and water conservation. Through automated text analysis, question-and-answer generation and intelligent update and maintenance, the data set is efficiently constructed and continuously optimized, including data collection, analysis, question-and-answer generation and update steps. ByteDance's text analysis API and doubao-pro-128k large language model are used, and data processing and update are combined with vector matching technology.

Benefits of technology

It realizes automatic analysis and unified management of professional soil and water conservation documents, improves data processing efficiency, generates high-quality Q&A pairs, ensures the consistency and timeliness of data quality, supports multi-format document processing, covers nine application scenarios, and improves usage efficiency and data set update capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338061A_ABST
    Figure CN120338061A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water and soil conservation, in particular to a method for constructing a water and soil conservation professional data set by using a large language model, which comprises the following steps of: S1, collecting literature literature such as water and soil conservation schemes, planning design, laws and regulations, scientific achievements and the like in recent ten years; s2, converting various documents into a unified format text database by using a text analysis interface; and S3, calling a large language model based on the prompt word cluster to generate professional question and answer pairs for nine application scenes. And S4, detecting the data overlap ratio through a vector matching technology, and realizing automatic intelligent updating of the data set. Through the method, a specialized water and soil conservation text database can be constructed, and a training data set for water and soil loss prediction, measure arrangement, benefit evaluation and other scenes is automatically generated, so that high-quality data support is provided for optimization of a water and soil conservation large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil and water conservation, and particularly relates to a method for constructing a professional dataset of soil and water conservation by using large language model technology. The present invention also relates to the application of artificial intelligence technology in the field of soil and water conservation, especially in the construction of professional knowledge bases and the development of intelligent question-answering systems. Background Art

[0002] With the development of artificial intelligence technology, large language models are increasingly widely used in various professional fields. In the field of soil and water conservation, there are a large number of professional literature, technical specifications and practical experiences. The effective integration and utilization of this knowledge is of great significance for improving the efficiency of soil and water conservation work. However, there are the following technical problems at present: First, the problem of data format diversity: The professional materials in the field of soil and water conservation have diverse formats, including pictures, PDFs, Word, etc., making it difficult to manage and utilize them uniformly. Especially for scanned copies and handwritten documents in historical materials, there are great difficulties in digitalization and structuring; Second, the problem of manual collation efficiency: The existing dataset construction methods mainly rely on manual collation, which is not only inefficient but also of uneven quality. Professionals need to invest a lot of time in data screening, collation and annotation, and it is difficult to ensure the consistency of annotation quality; Third, insufficient coverage of professional scenarios: There is a lack of professional question-and-answer datasets for specific application scenarios in the field of soil and water conservation, and existing general datasets cannot meet the needs of professional work in soil and water conservation, especially in highly professional fields such as technical solution formulation and benefit assessment; Fourth, the lack of an update and maintenance mechanism: It is difficult to update and maintain the dataset, and it cannot reflect the latest technological progress and practical experiences in a timely manner. Existing datasets are often static and lack a dynamic update mechanism, making it difficult to adapt to the rapid development of soil and water conservation technology. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for constructing a large language model dataset for soil and water conservation, aiming to solve the above technical problems. This method realizes the efficient construction and continuous optimization of the professional dataset of soil and water conservation through automated text parsing, large language model question-and-answer generation, and intelligent update and maintenance. The present invention also provides an extensible framework to support the unified processing and management of different types of professional materials.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A method for constructing a large language model dataset for soil and water conservation, comprising the following steps:

[0006] S1. Data collection step: Collect the text data related to soil and water conservation in the recent 10 years, specifically including: industry-related cases such as soil and water conservation plans, soil and water conservation programs, and small watershed management designs (including national key soil and water conservation projects, provincial soil and water conservation demonstration projects, award-winning soil and water conservation plans, and typical small watershed comprehensive management projects); laws, regulations and administrative rules and regulations related to soil and water conservation promulgated by administrative departments at all levels (including national laws and regulations and technical standards, provincial local regulations and implementation rules, industry norms and technical guidelines, relevant policy documents and notices); and scientific research results such as disciplinary basic textbooks, research papers, and postgraduate theses (including professional textbooks published by key universities, research papers published in core journals, excellent master's and doctoral theses, and research reports of key laboratories).

[0007] S2. Data parsing step: Use the text parsing interface to parse and uniformly convert various formatted documents collected into text data in markdown format, and construct a soil and water conservation text database; specifically including: document preprocessing (unifying file naming rules, optimizing OCR recognition, processing image resolution, and standardizing format conversion); content structuring (identifying chapter levels, extracting table data, annotating chart relationships, and processing formula symbols); data quality assessment (checking text integrity, proofreading professional terms, verifying format consistency, and automatically correcting errors).

[0008] S3. Q&A pair generation step: For nine application scenarios, call the large language model and use specific prompt clusters to generate professional Q&A pairs; the design of the prompt clusters includes: specific task descriptions set for different application scenarios (scene background description, task objective definition, technical requirement specification, expected output format); professional identity positions set for different roles (engineer perspective, researcher perspective, manager perspective, learner perspective); answer format requirements set for different question types (technical solution format, analysis report format, operation guide format, evaluation report format); and answer level requirements set for different professional depths (entry level, application level, professional level, research level).

[0009] S4. Dataset update step: Use vector matching technology to detect the overlap between new data and existing data, and achieve automatic and intelligent update of the dataset; specifically including: inputting new text data and parsing it into markdown format (automatically identifying document types, extracting key information, and standardizing processing); using vector matching technology to detect data overlap (text vectorization processing, similarity calculation, overlap assessment); update process control (ending the update when the overlap > 90%, continuing to process when the overlap < 90%, and handling boundary cases); generating and integrating new Q&A pairs (extracting different content, generating Q&A pairs, data quality assessment, data classification storage, and dataset integration).

[0010] Furthermore, the nine major application scenarios include: soil and water loss prediction, layout of soil and water conservation measures, evaluation of the benefits of soil and water conservation measures, generation of calculation code scripts related to soil and water conservation based on Python, consultation on soil and water conservation laws and regulations, assistance in the preparation of soil and water conservation plans, assistance in soil and water conservation scientific research, assistance in the study of soil and water conservation for undergraduates and postgraduate students, and popular science Q&A on soil and water conservation.

[0011] Furthermore, the text parsing interface in the data parsing step is the PDF and image-based text data parsing API provided by ByteDance, and the API is used to build an automated text parsing process.

[0012] Furthermore, the large language model used in the Q&A pair generation step is ByteDance's doubao-pro-128k large language model, and the large language model generates Q&A pairs with the content of the text database as the context.

[0013] Furthermore, the use of specific prompt clusters in the Q&A pair generation step to generate professional Q&A pairs can generate professional Q&A pairs for different application scenarios.

[0014] Furthermore, the dataset update step specifically includes:

[0015] S41: Input new text data and parse it into markdown format;

[0016] S42: Use vector matching technology to detect data overlap;

[0017] S43: End the update process when the overlap is greater than 90%;

[0018] S44: When the overlap is less than 90%, call the large language model API to generate new Q&A pairs;

[0019] S45: Summarize the newly generated Q&A pairs;

[0020] S46: Update the dataset.

[0021] Furthermore, the prompt clusters include:

[0022] S71: Specific task descriptions set for different application scenarios;

[0023] S72: Professional identity positioning set for different roles;

[0024] S73: Answer format requirements set for different question types;

[0025] S74: Requirements for the level of answers set for different professional depths.

[0026] Furthermore, the dataset update step includes:

[0027] S81. Summary of the main content of the newly added data;

[0028] S82. Distribution of application scenarios of the newly added question-and-answer pairs;

[0029] S83. Explanation of the supplementary value of the newly added question-and-answer pairs to the existing data set;

[0030] S84. Statistics on the scale and coverage of the updated data set.

[0031] Furthermore, the data quality assessment includes the following steps:

[0032] S91. Evaluate the professional accuracy of the generated question-and-answer pairs;

[0033] S92. Evaluate the logical integrity of the question-and-answer pairs;

[0034] S93. Evaluate the practicality of the question-and-answer pairs;

[0035] S94. Screen and optimize the question-and-answer pairs based on the evaluation results.

[0036] Furthermore, the data classification and storage includes the following steps:

[0037] S101. Classify and store the question-and-answer pairs according to application scenarios;

[0038] S102. Establish the association relationship between the question-and-answer pairs;

[0039] S103. Use a knowledge graph to construct a retrieval index for the question-and-answer pairs;

[0040] S104. Implement a multi-dimensional retrieval function by combining vector features.

[0041] By adopting the above technical solutions, the beneficial effects of the present invention are as follows:

[0042] 1. Achieved the automatic parsing and unified management of soil and water conservation professional literature, increased the data processing efficiency by more than 80%, reduced the manual sorting cost by more than 60%, ensured the consistency of data quality, and supported the processing of multi-format documents.

[0043] 2. Automatically generate high-quality professional Q&A pairs through large language models, covering nine application scenarios such as soil and water loss prediction (calculation of rainfall erosivity, assessment of soil erodibility, analysis of topographic factors, application of prediction models), layout of soil and water conservation measures (selection of measure types, optimization design of layout, determination of engineering parameters, guidance on construction key points), evaluation of the benefits of soil and water conservation measures (calculation of ecological benefits, analysis of economic benefits, assessment of social benefits, evaluation of comprehensive benefits), generation of Python-based calculation code scripts related to soil and water conservation (data processing scripts, model calculation scripts, visualization scripts, evaluation and analysis scripts), consultation on soil and water conservation laws and regulations (policy interpretation, application of standards and specifications, case analysis, compliance assessment), assistance in the preparation of soil and water conservation plans (generation of preparation outlines, design of content frameworks, recommendation of technical parameters, standardization of document formats), assistance in soil and water conservation scientific research (generation of literature reviews, recommendation of research methods, guidance on data analysis, suggestions for paper writing), assistance in the study of undergraduates and postgraduates in soil and water conservation (explanation of knowledge points, solution of exercises, experimental guidance, exam review), and popular science Q&A on soil and water conservation (explanation of basic concepts, introduction of typical cases, sharing of practical experience, answers to common questions).

[0044] 3. An intelligent dataset update mechanism is established to achieve automatic update and maintenance, ensure data timeliness, guarantee knowledge integrity, and support incremental updates.

[0045] 4. Classification management and multi-dimensional retrieval of the dataset are realized, including a multi-dimensional classification system, a precise retrieval function, and an associated recommendation mechanism, significantly improving the usage efficiency.

[0046] Through the above technical solutions, the present invention realizes the intelligent management and application of soil and water conservation professional knowledge, providing strong technical support for soil and water conservation work. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the overall flowchart of the present invention;

[0048] Figure 2 is the flowchart of data collection and parsing in the present invention;

[0049] Figure 3 is the flowchart of Q&A pair generation in the present invention;

[0050] Figure 4 is the flowchart of dataset update in the present invention;

[0051] Figure 5 is the system architecture diagram of the present invention; Figure 6 is the diagram of multi-dimensional data organization method in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0053] The present invention provides a method for constructing a large language model dataset for soil and water conservation. As Figure 1 shown, it mainly includes four main steps: data collection, data parsing, Q&A pair generation, and dataset update. At the same time, it also includes two auxiliary steps: data quality control and dataset classification management.

[0054] In an embodiment of the present invention, the data collection step mainly includes collecting soil and water conservation-related text data in the recent 10 years. The following uses specific examples to illustrate the collection content of three major categories of data:

[0055] 1. Example of industry-related case data:

[0056]

[0057] 2. Example of laws and regulations documents:

[0058]

[0059]

[0060] 3. Example of scientific research and teaching materials:

[0061]

[0062] The above examples show the basic structure and key information of the three major categories of data, which are stored and managed in a unified JSON format, facilitating subsequent data processing and application. Each category of data contains basic information such as time, source, and content, ensuring the traceability and availability of the data.

[0063] In the Q&A pair generation step, the present invention designs a complete set of prompt word cluster systems. The following uses specific examples to illustrate the structure and application of the prompt word clusters:

[0064]

[0065]

[0066] The above examples show the core components of the prompt word cluster system, including basic role settings, scenario-specific prompt words, and quality control parameters. Each scenario has a unique role positioning and task requirements, ensuring that the generated Q&A pairs are both professional and practical. Through strict quality control parameters, the reliability of the generated content is guaranteed.

[0067] This invention uses the ByteDance doubao-pro-128k large language model API, combined with the above prompt clusters, to generate professional Q&A pairs for nine application scenarios. As Figure 4 shown, these nine scenarios include: soil erosion prediction, layout of soil and water conservation measures, benefit evaluation of soil and water conservation measures, generation of Python-based calculation code scripts related to soil and water conservation, consultation on soil and water conservation laws and regulations, assistance in the preparation of soil and water conservation plans, assistance in soil and water conservation scientific research, assistance in the study of undergraduates and postgraduate students majoring in soil and water conservation, and popular science Q&A on soil and water conservation.

[0068] The following uses specific examples to illustrate the generation process and output format of Q&A pairs for the nine application scenarios:

[0069]

[0070]

[0071] The above examples show the generation results of Q&A pairs for three typical application scenarios. Each Q&A pair contains specific context information, professional questions, structured answers, and quality scores. Among them:

[0072] 1. The soil erosion prediction scenario emphasizes the standardization of calculation methods and the integrity of steps;

[0073] 2. The Python calculation script generation scenario focuses on the practicality of the code and the clarity of annotations;

[0074] 3. The scenario of assisting in plan preparation emphasizes the integrity and operability of the measure system.

[0075] All Q&A pairs are stored in JSON format for easy subsequent processing and application. The quality score is based on dimensions such as professionalism, accuracy, and practicality, with a full score of 5 points. The generation of Q&A pairs for other scenarios follows a similar structure and standard, adjusting the professional depth and expression of the content according to specific application scenarios.

[0076] In the dataset update step, this invention uses vector matching technology to achieve intelligent update. The following uses specific examples to illustrate the implementation process of this update process:

[0077] When the system receives a new soil and water conservation plan document, it first parses it into a standardized markdown format, as shown in the following example:

[0078]

[0079] Subsequently, the system uses the sentence-transformers model to calculate text vectors and calculate the cosine similarity with the documents in the existing database. The implementation code example:

[0080]

[0081]

[0082] When it is determined that an update is needed, the system will generate an update summary report, and the example format is as follows:

[0083]

[0084]

[0085] This example demonstrates the specific implementation methods of three key steps in the dataset update process, namely document parsing, similarity calculation, and update summary generation. Through standardized data formats and automated processing flows, the continuous update and quality control of the dataset are ensured.

[0086] In the data quality control step, the present invention establishes an evaluation system, as shown in Table 1:

[0087] Table 1 Quality Evaluation Index System for the Soil and Water Conservation Large Language Model Dataset

[0088]

[0089] Example evaluation records:

[0090]

[0091]

[0092] The above evaluation system ensures the quality of the dataset through multi-dimensional indicators, and specific evaluation methods and target values are set for each indicator. The evaluation results generate reports through an automated program, providing a basis for the continuous optimization of the dataset. The implementation of this evaluation system has significantly improved the professionalism, integrity, and practicality of the dataset, ensuring the training effect of the soil and water conservation large language model.

[0093] In the dataset classification and management step, the present invention realizes a multi-dimensional data organization method, such as Figure 6 shown, including classifying and storing question-and-answer pairs according to application scenarios, establishing the association relationship between question-and-answer pairs, constructing a retrieval index, and realizing a multi-dimensional retrieval function.

[0094] The following illustrates the data organizational structure through a specific example:

[0095]

[0096]

[0097]

[0098] This structure realizes the multi-dimensional organization and flexible retrieval of question-answer pairs, supports storage by scenario classification, association relationship management, and multi-dimensional retrieval. Different question-answer pairs are associated through a unique identifier (qa_id) to establish a knowledge-graph-like connection relationship, and similarity retrieval is supported by combining vector features.

[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a large language model dataset for soil and water conservation, characterized in that, It includes the following steps: S1. Data collection step: Collect the text data related to soil and water conservation in the recent 10 years, specifically including: industry-related cases such as soil and water conservation plans, soil and water conservation programs, and small watershed treatment designs (including national key soil and water conservation projects, provincial soil and water conservation demonstration projects, award-winning soil and water conservation plans, and typical small watershed comprehensive treatment projects); laws, regulations, and administrative rules and regulations related to soil and water conservation promulgated by administrative departments at all levels (including national laws, regulations, and technical standards, provincial local regulations and implementation rules, industry norms and technical guidelines, relevant policy documents and notices); and scientific research results such as discipline basic textbooks, research papers, and postgraduate theses (including professional textbooks published by key universities, research papers published in core journals, excellent master's and doctoral theses, and research reports of key laboratories). S2. Data parsing step: Use the text parsing interface to parse and uniformly convert various format documents collected into text data in markdown format, and construct a soil and water conservation text database; specifically including: document preprocessing (unifying file naming rules, optimizing OCR recognition, processing image resolution, and standardizing format conversion); content structuring (identifying chapter levels, extracting table data, annotating chart relationships, and processing formula symbols); data quality assessment (checking text integrity, proofreading professional terms, verifying format consistency, and automatically correcting errors). S3. Q&A pair generation step: For nine application scenarios, call the large language model and use specific prompt clusters to generate professional Q&A pairs; the design of the prompt clusters includes: specific task descriptions set for different application scenarios (scene background description, task objective definition, technical requirement specification, expected output format); professional identity positioning set for different roles (engineer perspective, researcher perspective, manager perspective, learner perspective); answer format requirements set for different question types (technical solution format, analysis report format, operation guide format, evaluation report format); and answer level requirements set for different professional depths (entry level, application level, professional level, research level). S4. Dataset update step: Use vector matching technology to detect the overlap degree between new data and existing data, and realize the automated intelligent update of the dataset; specifically including: inputting new text data and parsing it into markdown format (automatically identifying document types, extracting key information, and standardizing processing); using vector matching technology to detect data overlap degree (text vectorization processing, similarity calculation, overlap degree evaluation); update process control (ending the update when the overlap degree > 90%, continuing to process when the overlap degree < 90%, and handling boundary cases); generating and integrating new Q&A pairs (extracting different content, generating Q&A pairs, data quality assessment, data classification storage, and dataset integration).

2. The method for constructing a large language model dataset for soil and water conservation according to claim 1, wherein The nine application scenarios include: soil and water loss prediction, layout of soil and water conservation measures, benefit evaluation of soil and water conservation measures, generation of calculation code scripts related to soil and water conservation based on Python, consultation on soil and water conservation laws and regulations, assistance in the preparation of soil and water conservation plans, assistance in soil and water conservation scientific research, assistance in the study of undergraduates and postgraduate students majoring in soil and water conservation, and popular science Q&A on soil and water conservation.

3. A method for constructing a large language model dataset for soil and water conservation according to claim 1, characterized in that The text parsing interface in the data parsing step is the PDF and picture text data parsing API provided by ByteDance, and the API is used to build an automated text parsing process.

4. A method for constructing a large language model dataset for soil and water conservation according to claim 1, characterized in that, The large language model used in the Q&A pair generation step is ByteDance's doubao-pro-128k large language model, and the large language model generates Q&A pairs based on the content of the text database.

5. A method for constructing a large language model dataset for soil and water conservation according to claim 1, characterized in that, In the Q&A pair generation step, the use of specific prompt clusters to generate professional Q&A pairs can generate professional Q&A pairs for different application scenarios.

6. A method for constructing a large language model dataset for soil and water conservation according to claim 1, characterized in that, The dataset update step specifically includes: S41. Input new text data and parse it into markdown format; S42. Use vector matching technology to detect data overlap; S43. When the overlap is greater than 90%, end the update process; S44. When the overlap is less than 90%, call the large language model API to generate new Q&A pairs; S45. Summarize the newly generated Q&A pairs; S46. Update the dataset.

7. A method for constructing a large language model dataset for soil and water conservation according to claim 1, characterized in that, The prompt clusters include: S71. Specific task descriptions set for different application scenarios; S72. Professional identity positioning set for different roles; S73. Answer format requirements set for different question types; S74. Answer level requirements set for different professional depths.

8. The method for constructing a large language model dataset for soil and water conservation according to claim 1, wherein, The dataset update step includes: S81. Overview of the main content of the new data; S82. Distribution of application scenarios of the newly added Q&A pairs; S83. Explanation of the supplementary value of the newly added Q&A pairs to the existing dataset; S84. Statistics on the scale and coverage of the dataset after update.

9. The method for constructing a large language model dataset for soil and water conservation according to claim 1, wherein, The data quality assessment includes the following steps: S91. Evaluate the professional accuracy of the generated Q&A pairs; S92. Evaluate the logical integrity of the Q&A pairs; S93. Evaluate the practicality of the Q&A pairs; S94. Screen and optimize the Q&A pairs based on the evaluation results.

10. The method for constructing a large language model dataset for soil and water conservation according to claim 1, wherein, The data classification and storage includes the following steps: S101. Classify and store Q&A pairs by application scenario; S102. Establish the association relationship between Q&A pairs; S103. Use a knowledge graph to construct a Q&A pair retrieval index; S104. Combine vector features to achieve multi-dimensional retrieval functions.