Intelligent water conservancy system corpus enhancement processing method and storage medium
Through the intelligent water conservancy system corpus enhancement treatment method, the problem that traditional corpus treatment methods are difficult to accurately analyze and extract water conservancy corpus is solved, and high-quality water conservancy corpus generation is achieved, providing reliable data support for water conservancy natural language processing tasks.
Patent Information
- Application Number
- CN202510360816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional corpus processing methods are difficult to meet the needs of accurate analysis and knowledge extraction of water conservancy corpus, especially in document tiling, vectorization, similarity calculation and knowledge extraction.
A method for corpus enhancement treatment of intelligent water conservancy systems is proposed, including multi-channel document collection and preprocessing, semantic document tiling, water conservancy professional vocabulary library construction and vectorization, multi-method similarity calculation, document block combination based on knowledge system and big model knowledge extraction.
Through this method, water conservancy text resources can be effectively integrated and standardized, ensure that document blocks have independent and complete semantic information, accurately capture the semantic relationships of professional vocabulary, improve the accuracy of similarity calculations and the rationality of document block combinations, generate high-quality water conservancy corpus, and provide reliable support for natural language processing tasks.
Smart Images

Figure CN120218220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of natural language processing and water conservancy technology, in particular to a method for enhancing the processing of corpus of an intelligent water conservancy system and a storage medium. Background Art
[0002] With the wide application of information technology in the water conservancy industry, it has become increasingly important to effectively analyze and utilize the text data in the water conservancy field. However, the documents in the water conservancy field have the characteristics of wide sources, diverse formats, strong content professionalism and complex structures. Traditional corpus processing methods are difficult to meet the needs of accurate analysis and knowledge extraction of water conservancy corpus. For example, during document chunking, it is impossible to reasonably split according to the specific logical structure of water conservancy documents; during the vectorization process, the semantic relationships between water conservancy professional terms cannot be accurately captured; it is difficult to consider the complex semantic similarity factors in the water conservancy field during similarity calculation; unreasonable combinations are likely to occur in document block combination; there are limitations in the understanding and application of water conservancy professional knowledge by large models during knowledge extraction, resulting in low - quality generated corpus sets and unable to provide reliable support for natural language processing tasks related to water conservancy. Summary of the Invention
[0003] The present invention proposes a method for enhancing the processing of corpus of an intelligent water conservancy system, which solves the problem that traditional corpus processing methods in the prior art are difficult to meet the needs of accurate analysis and knowledge extraction of water conservancy corpus. The technical solution of the present invention is realized as follows:
[0004] A method for enhancing the processing of corpus of an intelligent water conservancy system, characterized by including the following steps:
[0005] Collect relevant documents in the water conservancy field from multiple channels and pre - process them to uniformly convert them into standardized documents in pure text format; chunk the above - mentioned documents semantically to form independent corpus units; construct a water conservancy professional vocabulary library, and use large - scale text data in the water conservancy field to train a word - vector model based on the corpus to vectorize the above - mentioned corpus units; calculate the similarity of the vectorized corpus units based on multiple methods; establish a combination rule library based on the knowledge system in the water conservancy field in combination with the similarity, and combine the document blocks of water conservancy engineering design; use the above - mentioned document materials as the input of a large - language model to achieve knowledge extraction and corpus set generation in the water conservancy system field.
[0006] As a preferred technical solution, collect relevant documents in the water conservancy field from multiple channels such as water conservancy professional databases, scientific research institutions and enterprise internal document libraries, and the document types include academic papers, engineering reports, policy documents, and monitoring data records.
[0007] As a preferred technical solution, the preprocessing process includes: using a format conversion tool to uniformly convert documents in different formats into plain text format, and using a text cleaning tool to remove special symbols, garbled characters, and irrelevant information such as headers and footers. At the same time, standardize the abbreviations of water conservancy terms.
[0008] As a preferred technical solution, the document chunking is carried out according to the following rules: 1) For academic papers, based on the standard structure of introduction, method, results and discussion, and conclusion, combined with the paragraph level, each main part is subdivided into sub-chunks centered around a clear theme; 2) Engineering report documents are chunked according to the project process, modules, and specific engineering facilities or technical links in each stage, and the key content chunks are accurately located through keyword matching and title recognition technologies; 3) Monitoring data record documents are chunked according to time periods, monitoring stations, or monitoring indicators, and the data is preliminarily sorted and labeled for correlation analysis with text information.
[0009] As a preferred technical solution, professional vocabulary should be combined with manual annotation and word segmentation to avoid mis-segmentation or splitting that may occur with conventional word segmentation tools.
[0010] As a preferred technical solution, for the word vector model, large-scale text data in the water conservancy field is used for training to learn the vector representation of professional vocabulary. At the same time, through manually annotated synonym pairs or using context semantic relationships, the model learns the relevance of words with similar meanings in the vector space; for longer document chunks, sentence-level vectorization is performed based on a pre-trained language model, and the sentence vectors are weighted and summed or fused through a self-attention mechanism to obtain the vector representation of the document chunk.
[0011] As a preferred technical solution, similarity calculation uses semantic extension technology based on a knowledge graph to associate key water conservancy entities in the document chunk with the water conservancy field knowledge graph, obtain relevant attribute, relationship, and concept information, and incorporate it into the similarity calculation to more comprehensively and accurately measure the semantic similarity degree between document chunks; similarity calculation can also introduce additional features and weights based on methods such as cosine similarity to calculate vector similarity, combined with water conservancy field knowledge and semantic understanding.
[0012] As a preferred technical solution, for water conservancy engineering design document chunks, document chunks that are relevant or complementary in terms of project type, design standards, geographical environment conditions, etc. are preferentially combined; for water resource management document chunks, reasonable combination is carried out according to the classification and association relationship of management objects and measures.
[0013] As a preferred technical solution, knowledge extraction and corpus generation include the following steps:
[0014] Step S1: Fine-tune and train the LLM_HB model using data such as authoritative textbooks, standard specifications, the latest research results, and expert knowledge in the water conservancy field, so that it can better understand and master the professional knowledge in the water conservancy field;
[0015] Step S2: Design a knowledge extraction template and prompting strategy for the water conservancy field to guide the large model to extract knowledge according to specific formats and requirements, and generate a corpus set LOCAL_QA in the form of question-answer pairs, such as extracting key parameters in water conservancy project design and their basis for value selection, knowledge points in aspects such as the current situation and problems of water resources management;
[0016] Step S3: Establish a regular update mechanism for the corpus set, closely follow the latest developments in the water conservancy field, collect and sort out new knowledge sources, and integrate new knowledge into the existing corpus set according to the above process to continuously update and improve LOCAL_QA.
[0017] A non-temporary storage medium is used to store a program for executing the above-mentioned intelligent water conservancy system corpus enhancement processing method.
[0018] Compared with the prior art, the present solution has the following beneficial effects:
[0019] 1) Through targeted document collection and preprocessing, various text resources in the water conservancy field can be effectively integrated and converted into a unified and standardized format, providing high-quality raw data for subsequent processing.
[0020] 2) In the document chunking link, reasonable segmentation is carried out according to the unique structure and content characteristics of water conservancy documents to ensure that each document chunk has independent and complete semantic information, which is beneficial to improving the accuracy of subsequent processing.
[0021] 3) In the vectorization process, special training and optimization are carried out for water conservancy professional vocabulary, which can accurately capture the semantic relationships between professional terms, thereby improving the accuracy of vector representation and laying a good foundation for similarity calculation and knowledge extraction.
[0022] 4) The similarity calculation method fully considers the complex semantic factors in the water conservancy field, combines the knowledge graph and domain characteristics for comprehensive evaluation, and can more accurately discover semantically similar document chunks, improving the rationality of document chunk combination.
[0023] 5) The document chunk combination is based on the rule library and review mechanism of the water conservancy field knowledge system, avoiding unreasonable text collocations and ensuring that the combined text conforms to the water conservancy professional logic and actual application scenarios.
[0024] 6) The knowledge extraction process can efficiently and accurately extract valuable water conservancy knowledge from the text through enhanced training of the large model in the water conservancy field and customized prompting strategies, generate a high-quality corpus set LOCAL_QA, provide reliable data support for natural language processing tasks related to water conservancy, and promote the informatization development of the water conservancy industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a method flow chart of a corpus enhancement processing method for an intelligent water conservancy system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0028] Referring to Figure 1 , the present invention provides a corpus enhancement processing method for an intelligent water conservancy system, including the following steps: collecting relevant documents in the water conservancy field from multiple channels and preprocessing them to uniformly convert them into standardized documents in pure text format; cutting the above documents into independent corpus units according to semantics; constructing a water conservancy professional vocabulary library, and training the word vector model based on the corpus using large-scale water conservancy field text data to vectorize the above corpus units; calculating the similarity of the vectorized corpus units based on multiple methods; establishing a combination rule library based on the water conservancy field knowledge system in combination with the similarity, and combining the water conservancy engineering design document blocks; using the above document materials as the input of the large language model to realize knowledge extraction and corpus set generation in the water conservancy system field.
[0029] Specifically, it includes the following steps:
[0030] 1. Document collection and preprocessing
[0031] 1) Diversified collection channels
[0032] a) Obtain documents such as academic research results and cutting-edge technology discussions from professional databases, such as the water conservancy engineering subject database of CNKI and the water conservancy-related document database of Web of Science. The database contains high-quality academic papers from around the world, covering various sub-fields of water conservancy engineering, such as theoretical research and practical analysis articles in hydraulics, hydrology and water resources, water conservancy and hydropower engineering, etc.
[0033] b) Project reports, experimental data records, technical solutions and other documents within water conservancy research institutions are also important sources. For example, the detailed reports produced by the various research institutes of the China Institute of Water Resources and Hydropower Research after completing major national water conservancy research projects contain undisclosed but extremely valuable first-hand research materials, in-depth analysis of specific water conservancy problems and exploration of solutions.
[0034] c) Internal document library of enterprises, including design drawings of water conservancy engineering design companies, engineering construction records of construction companies, product technical manuals of water conservancy equipment manufacturing companies, etc. For example, the construction logs and quality inspection reports accumulated by China Gezhouba Group during the construction of many large-scale water conservancy hub projects reflect the specific conditions and technical application details of the actual projects.
[0035] 2) Format unification and cleaning
[0036] a) Use software such as Adobe Acrobat Pro to convert PDF documents into plain text, use the "Save As" function of Microsoft Word to convert complex Word documents into simple text format, and for water conservancy data records in Excel, use the data export function to convert them into CSV format and then organize them into text form to ensure that all documents are in a processable plain text format.
[0037] b) Use the regular expression library (re) in Python to write a text cleaning program to remove special symbols (such as mathematical symbols, non-standard punctuation, etc.), garbled characters (which may be caused by document format conversion errors or original input errors), headers and footers (including irrelevant information such as document page numbers, author units, dates, etc.). For example, use `re.sub(r'[^\w\s]',",text)` to remove most non-alphanumeric and space characters, and delete them by identifying the specific format of headers and footers (such as text in a fixed position, a specific font or font size), so as to obtain clean and tidy text content for subsequent processing.
[0038] c) For abbreviations in water conservancy terms, establish a dictionary of water conservancy term abbreviations, such as {"COD": "Chemical Oxygen Demand", "DO": "Dissolved Oxygen", "TMDL": "Total Maximum Daily Load", "SWMM": "Storm Water Management Model"}. Through text replacement operations, replace all abbreviations in the document with the complete terms to ensure the consistency and standardization of the terms, so that subsequent natural language processing steps can accurately understand and process these professional vocabulary. (The cases of polysemy and synonymy also need to be considered, as simple replacement may lead to errors).
[0039] 2. Document chunking
[0040] 1) Academic paper chunking
[0041] a) First, based on the standard structure of the paper, use the title and paragraph formats to identify parts such as the introduction, methods, results and discussion, and conclusion. For example, usually the introduction part will elaborate on the research background and purpose, the methods part will describe in detail the specific steps and technical means of the experiment or research, the results and discussion part will present the data and the analysis of the results, and the conclusion part will summarize the research results and look ahead to future research directions.
[0042] b) Within each main part, further subdivide according to the semantic coherence of the paragraphs. For example, in the "methods" part, if there are multiple experimental methods, each experimental method and its related data collection and processing steps are taken as a sub-block. By identifying the key sentences (such as sentences containing conjunctions like "first", "second", "then", etc., and sentences elaborating on the core steps) and keywords (such as experimental technique names, equipment names, etc.) in the paragraphs, consecutive relevant paragraphs are divided into a document block with a clear theme, ensuring that each block can independently describe a complete sub-process of the experiment or research.
[0043] 2) Engineering report chunking
[0044] a) Make a preliminary division according to the project life cycle, namely the stages of planning and design, construction, operation and maintenance, etc. For the planning and design stage, use keyword matching (such as "site selection analysis", "scheme comparison", "engineering layout", etc.) (if the format unification and cleaning operations are carried out first, the title level recognition cannot be performed. Whether a conventional report can be converted into html / xml format, otherwise the h1 / h2 tags cannot be recognized through the title level recognition technology), and divide the complete design scheme of a certain water conservancy facility (such as a dam, reservoir, irrigation canal, etc.), including design parameters (such as dam height, reservoir capacity, canal slope, etc.), design basis (such as geological exploration report, hydrological calculation results, etc.), design drawing descriptions, etc. into a document block.
[0045] b) During the construction stage, taking each construction section and construction process (such as foundation excavation, concrete pouring, earth filling, etc.) as a unit, combining information such as the date in the construction log and the record of the construction team, organize relevant construction records, quality inspection reports, safety accident reports, etc. into a document block to fully reflect the actual situation of this construction link.
[0046] c) During the operation and maintenance stage, cut according to different systems of water conservancy facilities (such as the monitoring system, flood discharge system, power generation system, etc. of the dam) and maintenance tasks (such as regular inspections, equipment maintenance, troubleshooting, etc.), and group the corresponding operation data records, maintenance operation manuals, fault diagnosis reports, etc. together to facilitate the subsequent extraction and analysis of the operation and maintenance knowledge of specific facilities.
[0047] 3) Cutting of monitoring data records (There seems to be redundancy among the three cuts a, b, and c in this part. Consider mainly cutting along the time period as the main line, with stations / indicators only as auxiliary, such as using database queries).
[0048] a) Cut according to the time period. For example, for the monitoring data of daily water level, flow rate, water quality indicators, etc., summarize the data of the same day into a document block. The record contains information such as the name of the monitoring station, monitoring time, the values of each indicator, and the status of the data acquisition equipment, etc., to facilitate the subsequent analysis of the data change trend according to the time series.
[0049] b) Cut according to the monitoring stations, and integrate various types of monitoring data of the same station at different times. This can deeply study the long-term change characteristics of water conservancy elements in the area represented by this station, and is of great significance for analyzing the water resources status and the impact of water conservancy projects in local areas.
[0050] c) Cut according to the monitoring indicators, extract all the monitoring data regarding a specific indicator (such as ammonia nitrogen content, dissolved oxygen concentration, etc. in water quality) from the data sets of different stations and times to form a document block, which helps to specifically study the distribution and change law of this indicator in the entire study area, and provides detailed data support for water resources quality assessment and pollution control. At the same time, conduct preliminary data collation on each cut monitoring data document block, such as removing obviously incorrect data (through methods such as data range judgment and abnormal data fluctuation detection), supplementing missing data (using data interpolation method or estimating according to the trend of adjacent data), and adding necessary annotations (such as data source, data processing method, etc.) to effectively associate and analyze with other text information.
[0051] 3. Vectorization processing
[0052] 1) Vectorization of professional vocabulary
[0053] a) Build a vocabulary library for water conservancy. Extract professional words and terms from standards and specifications in the water conservancy field (such as "Classification and Flood Standards for Water Resources and Hydropower Projects", "Design Code for Irrigation and Drainage Engineering", etc.), professional dictionaries (such as "Comprehensive Dictionary of Water Conservancy"), authoritative textbooks (such as "Hydraulics", "Fundamentals of Hydrogeology", etc.), and a large number of documents in the water conservancy field to establish a vocabulary library containing thousands of water conservancy professional words. (Ensure that there is no obvious semantic overlap between the words.)
[0054] b) Use the Word2Vec or GloVe model to train on a large-scale text data in the water conservancy field (such as the text collection after preprocessing of various types of documents collected above, with a quantity reaching millions of words or more). During the training process, set appropriate parameters, such as window size (5 - 10), word vector dimension (100 - 300), minimum word frequency (5 - 10, filtering out words with too low occurrence frequencies, which may be spelling mistakes or overly rare professional terms and have little impact on the overall semantic understanding), etc.
[0055] c) Manually annotate some word pairs with similar meanings but different expressions in the water conservancy field (such as "sluice" and "dam", "irrigation water" and "farmland irrigation water", etc.). When training the model, increase the probability of co-occurrence of these word pairs in the context, so that the model can learn the semantic similarity between them, and thus make the vector distances of these synonyms close in the vector space. At the same time, using the context semantic relationship of words, for words that frequently appear in the same or similar contexts (such as during the description of the construction process of water conservancy projects, "concrete pouring" and "vibrating" often appear simultaneously), the model can automatically learn the close connection between them, and this correlation can also be reflected in the vector representation to more accurately capture the semantic connotations and interrelationships of water conservancy professional words.
[0056] 2) Vectorization of document blocks
[0057] a) For longer document blocks, use a pre-trained language model (such as BERT or its variants, such as RoBERTa, ALBERT, etc.) for sentence-level vectorization. First, input each sentence in the document block into the pre-trained model. The model will convert each sentence into a vector representation with a fixed dimension (such as 768 dimensions or 1024 dimensions) according to its internal word embedding layer, multi-head attention mechanism, and feed-forward neural network and other structures. This vector can capture features such as the semantic information, grammatical structure, and context of the sentence.
[0058] b) Then, obtain the vector representation of the entire document block through weighted summation or other fusion strategies. In the weighted summation method, different weights can be assigned to the vectors of each sentence according to factors such as the positional importance of the sentence in the document block (for example, sentences at the beginning and end are usually more summary and critical, and are given higher weights), the number and importance of keywords in the sentence (calculating keyword weights using methods such as TF-IDF), etc. Then, add and average the weighted sentence vectors to obtain the final vector representation of the document block. Or adopt a fusion method based on the attention mechanism, allowing the model to automatically learn the contribution degree of each sentence vector to the semantic representation of the entire document block, thereby generating a document block vector that comprehensively considers the information of each sentence, in order to better capture the overall semantic information and context of the document block, and provide a more effective data representation for subsequent similarity calculation and knowledge extraction.
[0059] 4. Similarity calculation
[0060] 1) Vector distance adjustment based on domain knowledge
[0061] a) When calculating the vector distance, in addition to using common methods such as cosine similarity, combine water conservancy domain knowledge and semantic understanding to introduce some additional features and weights. First, classify and label the document blocks. For example, label them according to dimensions such as water conservancy project type (divided into categories such as dam projects, reservoir projects, irrigation projects, flood control projects, etc.), geographical region (divided by river basin, such as the Yangtze River Basin, Yellow River Basin, Pearl River Basin, etc., or divided by administrative region, such as water conservancy projects in a certain province, water resource management in a certain city, etc.), hydrological and meteorological conditions (such as arid regions, humid regions, flood-prone regions, etc.).
[0062] b) When calculating the similarity between two document blocks, if they belong to the same water conservancy project type, on the basis of cosine similarity, give a certain weight increase (such as increasing the weight coefficient by 0.1 - 0.3, and the specific value can be adjusted according to the actual situation and experimental results), because water conservancy projects of the same type have more similarities and correlations in design principles, construction methods, operation management, etc. Such weight adjustment can highlight the importance of such domain-specific factors in similarity judgment. Similarly, for document blocks from the same geographical region or with similar hydrological and meteorological conditions, also give appropriate weight additions according to their degree of relevance, so that the similarity calculation results are more in line with the actual situation and knowledge logic of the water conservancy field.
[0063] 2) Semantic expansion assisted by knowledge graph
[0064] a) Construct a knowledge graph in the water conservancy field to structurally represent entities such as water conservancy engineering facilities (such as dams, sluices, pumping stations, etc.), water conservancy geographical elements (such as rivers, lakes, reservoirs, etc.), water conservancy technical methods (such as concrete pouring technology, hydrological monitoring technology, water resource allocation methods, etc.), water conservancy events (such as flood disasters, drought events, water conservancy project construction events, etc.) and the relationships between them (such as inclusion relationship, location relationship, function relationship, causal relationship, etc.). For example, in the knowledge graph, the "dam" entity is connected to the "concrete pouring" technology through the "construction method" relationship, to the "river" through the "interception" relationship, and to the "reservoir" through the "component" relationship, etc.
[0065] b) Use the semantic extension technology based on the knowledge graph to associate the key water conservancy entities in the document block with the water conservancy field knowledge graph, obtain their relevant attribute, relationship and concept information, and integrate this extended semantic information into the similarity calculation. For example, if a document block mentions the "Three Gorges Dam", through the knowledge graph, the attribute information such as the dam height, reservoir capacity, installed power generation capacity, basin where it is located, construction year, etc. of the Three Gorges Dam can be obtained, as well as the relationship information with the upstream and downstream water conservancy facilities and the surrounding geographical environment. Compare and match this information with another document block that also mentions relevant dam content, and evaluate the similarity between them from a more extensive semantic level, rather than just being limited to a simple match at the lexical level. This can more comprehensively and deeply measure the semantic similarity between document blocks, improve the accuracy and reliability of the similarity calculation, and provide a more valuable basis for subsequent document block combination and knowledge extraction.
[0066] 5. Document block combination
[0067] 1) Construction of a rule base based on the knowledge system
[0068] a) For water conservancy engineering design document blocks, establish a detailed combination rule base. For example, in terms of project type, if a document block is about the design of a gravity dam, preferentially combine other document blocks related to the gravity dam, such as document blocks comparing different design schemes (including comparisons of different dam types, different material selections, different structural layouts), document blocks on the foundation treatment technology of the gravity dam (such as dam foundation excavation, reinforcement methods, etc.), and document blocks on the key technologies and problem-solving during the construction of the gravity dam (such as concrete temperature control measures, prevention and control of construction period cracks, etc.). These document blocks are directly relevant in terms of project type, and combining them can provide more comprehensive and in-depth knowledge of gravity dam design.
[0069] b) In terms of design standards, according to different flood control standards, seismic standards, durability standards, etc., combine document blocks that meet the same or similar design standards. For example, for the design document block of a water conservancy project with a once-in-a-century flood control standard, combine it with other project cases, interpretations of design specifications, flood calculation methods, etc. that also adopt this flood control standard, so as to compare and learn from the experiences and technical measures of different projects under the same design standard.
[0070] c) In terms of geographical environmental conditions, consider factors such as topography and geomorphology (such as mountains, plains, hills, etc.), geological conditions (such as rock types, geological structures, soil characteristics, etc.), and climate conditions (such as temperature, precipitation, wind speed, etc.), and combine the design document blocks of water conservancy projects in similar geographical environments. For example, for the design document block of a water conservancy project built in a mountainous canyon area, combine it with the document blocks of other mountain water conservancy projects in aspects such as site selection analysis, dam type selection, and construction transportation plans. Because similar geographical environments will face some common problems and challenges, the combined text can provide more ideas and methods for solving these problems.
[0071] d) For water resource management document blocks, also combine them according to the classification and correlation of management objects and measures. In terms of management objects, combine different types of document blocks such as urban water use management, agricultural water use management, and industrial water use management together, and at the same time consider the combination of management document blocks for different water sources (such as surface water, groundwater, rainwater utilization, etc.). For example, combine the document blocks of urban sewage treatment and reuse with the document blocks of urban water conservation measures and urban water resource planning to form a text set on the comprehensive management of urban water resources, so as to comprehensively analyze the management strategies and technical methods in the whole process of urban water resource acquisition, utilization, and treatment and reuse. In terms of management measures, combine relevant document blocks such as water resource monitoring and assessment, water resource allocation, and water price policy formulation according to their internal logical relationships. For example, combine the document blocks using the same monitoring technology and index system to analyze how to formulate effective allocation plans and policy measures based on water resource monitoring data in different regions, so as to provide a more systematic and comprehensive reference basis for water resource management decisions.
[0072] 2) Combination review mechanism
[0073] a) Introduce a manual review process, where experts in the water conservancy field or professionals with rich engineering practice experience check and evaluate the preliminarily combined text. Based on their professional knowledge and practical experience, the experts judge whether there are problems in aspects such as logical coherence, rationality of professional knowledge, and compliance with actual application scenarios of the combined text. For example, when checking the combined water conservancy project design text, check whether the parameters of different design schemes are contradictory, whether the application of construction technologies conforms to the actual project conditions, and whether the design process is complete and reasonable; for the water resources management text, review whether the water resources allocation plan takes into account the local water use demand and water resources distribution characteristics, and whether the water price policy is feasible and reasonable, etc.
[0074] b) At the same time, use a rule-based intelligent review system as an auxiliary means. This system automatically checks the combined text according to the pre-set knowledge rules and logical relationships in the water conservancy field. For example, by writing rule codes, check whether there are problems such as incorrect use of professional terms, missing or abnormal key data, and incorrect order of technical processes in the text. If problems are found in the combined text, such as technical parameter conflicts (e.g., different and unreasonable numerical values for the design flow of the same water conservancy facility are given in two document blocks), management measures not conforming to the actual situation (e.g., an unrealistic high-water-consuming agricultural irrigation plan is proposed in a water-scarce area), etc., adjust the combination method of the document blocks (replace one or several of the document blocks) or re-select appropriate document blocks for combination in a timely manner to ensure that the quality of the finally combined text is reliable, meets the professional requirements and actual application scenarios in the water conservancy field, and provides high-quality input text for subsequent knowledge extraction.
[0075] 6. Knowledge extraction and corpus generation
[0076] 1) Knowledge enhancement training of the large model in the water conservancy field
[0077] a) Collect data such as authoritative textbooks in the water conservancy field (such as undergraduate and postgraduate textbooks of water conservancy majors from well-known domestic universities), standard specifications (such as water conservancy project design specifications, construction specifications, and water resources management standards issued by the state and localities), the latest research results (such as screening out the latest 3 - 5 years of frontier research papers from top academic journals and conference proceedings in the water conservancy field), and expert knowledge (by inviting senior experts in the water conservancy industry to write technical summaries, experience sharing and other documents, or conducting interviews with experts and organizing them into text materials), etc., to form a rich water conservancy field knowledge dataset, with a total word count reaching millions of words or more.
[0078] b) Use this dataset to fine-tune and train the large model LLM_HB. First, divide the dataset into a training set, a validation set, and a test set according to a certain ratio (e.g., 80% for training, 10% for validation, and 10% for testing). Then, according to the training framework and interface of the large model, input the training set into the model to let the model learn knowledge such as professional terms, concept systems, technical principles, and engineering cases in the water conservancy field. During the training process, set appropriate training parameters, such as the learning rate (generally initially set to 1e -5 ~5e -5 , and adjust it according to the convergence situation and training effect of the model), the number of training epochs (usually set to 10 - 20 epochs to ensure that the model can fully learn the knowledge in the water conservancy field but avoid overfitting), the Batch_Size batch size (set to 16 - 64 according to the hardware resources and dataset scale), etc. By continuously adjusting these parameters, use the validation set to verify the training effect of the model. When the loss function value of the model on the validation set no longer decreases and the accuracy reaches a certain level (e.g., above 85%), stop the training to obtain a large model enhanced by water conservancy field knowledge training, enabling it to better understand and master the water conservancy field.
[0079] Beneficial effects achieved:
[0080] 1. Through targeted document collection and preprocessing, various text resources in the water conservancy field can be effectively integrated and transformed into a unified and standardized format, providing high-quality raw data for subsequent processing.
[0081] 2. In the document chunking process, reasonable segmentation is carried out based on the unique structure and content characteristics of water conservancy documents to ensure that each document chunk has independent and complete semantic information, which is conducive to improving the accuracy of subsequent processing.
[0082] 3. During the vectorization process, specialized training and optimization are carried out for water conservancy professional vocabulary, which can accurately capture the semantic relationships between professional terms, thereby improving the accuracy of vector representation and laying a good foundation for similarity calculation and knowledge extraction.
[0083] 4. The similarity calculation method fully considers the complex semantic factors in the water conservancy field and conducts comprehensive evaluation by combining the knowledge graph and domain characteristics, which can more accurately discover semantically similar document chunks and improve the rationality of document chunk combination.
[0084] 5. The document chunk combination is based on the rule base and review mechanism of the water conservancy field knowledge system, avoiding unreasonable text collocations and ensuring that the combined text conforms to the professional logic and actual application scenarios of water conservancy.
[0085] 6. Through the enhanced training of the large model in the water conservancy field and customized prompting strategies in the knowledge extraction process, valuable water conservancy knowledge can be efficiently and accurately extracted from the text, generating a high-quality corpus LOCAL_QA, providing reliable data support for natural language processing tasks related to water conservancy, and promoting the informatization development of the water conservancy industry.
[0086] Based on the above innovative points, this application provides a method for enhancing the processing of water conservancy corpora to solve the problems existing in the prior art, improve the quality and usability of water conservancy corpora, and provide strong data support for the informatization development of the water conservancy industry.
[0087] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for corpus enhancement processing of intelligent water conservancy system, characterized in that: The following steps are involved: Collect relevant documents in the field of water conservancy from various channels and convert them into standardized documents in plain text format after preprocessing; segment the above documents into independent corpus units according to semantics; construct a water conservancy professional vocabulary database, and use large-scale water conservancy field text data to train the corpus-based word vector model to vectorize the above corpus units; Calculate the similarity of corpus units after vectorization based on various methods; Combined with the similarity, a combination rule base based on the knowledge system in the water conservancy field is established to combine the water conservancy project design document blocks; the above document data is used as the input of the large language model to realize knowledge extraction and corpus generation in the field of water conservancy system.
2. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: Relevant documents in the field of water conservancy are collected through multiple channels, including water conservancy professional databases, scientific research institutions, and internal corporate document libraries. The document types include academic papers, engineering reports, policy documents, and monitoring data records.
3. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: The preprocessing process includes: using a format conversion tool to convert documents of different formats into a plain text format, and using a text cleaning tool to remove special symbols, garbled characters, and irrelevant information in headers and footers, and standardizing the abbreviations of water conservancy terms.
4. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: Document segmentation is carried out according to the following rules: 1) For academic papers, each main part is subdivided into sub-blocks around a clear theme based on the standard structure of introduction, methods, results and discussion, and conclusion, combined with the paragraph level; 2) Engineering report documents are segmented according to project processes and modules as well as specific engineering facilities or technical links in each stage, and key content blocks are accurately located through keyword matching and title recognition technology; 3) Monitoring data record documents are segmented according to time periods, monitoring sites or monitoring indicators, and the data are preliminarily organized and labeled for correlation analysis with text information.
5. The method for corpus enhancement processing of intelligent water conservancy system according to claim 4, characterized in that: Professional vocabulary should be manually marked and segmented to avoid possible mis-segmentation or fragmentation by conventional segmentation tools.
6. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: For the word vector model, large-scale text data in the water conservancy field is used for training to learn the vector representation of professional vocabulary. At the same time, the model learns the correlation between words with similar meanings in the vector space by manually annotating synonym pairs or using contextual semantic relationships. For longer document blocks, sentence-level vectorization based on a pre-trained language model is adopted, and the sentence vectors are weighted summed or fused through a self-attention mechanism to obtain the document block vector representation.
7. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: Similarity calculation uses semantic extension technology based on knowledge graph to associate key water conservancy entities in document blocks with water conservancy field knowledge graph, obtain relevant attributes, relationships and concept information, and integrate them into similarity calculation to measure the semantic similarity between document blocks more comprehensively and accurately; Similarity calculation can also introduce additional features and weights based on the calculation of vector similarity using methods such as cosine similarity, combined with water conservancy domain knowledge and semantic understanding.
8. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: For water conservancy project design document blocks, priority is given to combining document blocks that are relevant or complementary in terms of project type, design standards, geographical environmental conditions, etc.; for water resources management document blocks, reasonable combinations are made based on the classification and correlation of management objects and measures.
9. The method for corpus enhancement processing of intelligent water conservancy system according to claim 1, characterized in that: Knowledge extraction and corpus generation include the following steps: Step S1: Use authoritative textbooks, standards and specifications, the latest research results, and expert knowledge in the field of water conservancy to fine-tune the LLM_HB model so that it can better understand and master the professional knowledge in the field of water conservancy; Step S2: Design a knowledge extraction template and prompt strategy for the water conservancy field, guide the large model to extract knowledge according to specific formats and requirements, and generate a corpus LOCAL_QA in the form of question-answer pairs, such as extracting key parameters of water conservancy project design and their value basis, the current status and problems of water resources management, and other knowledge points; Step S3: Establish a regular update mechanism for the corpus, pay close attention to the latest developments in the water conservancy field, collect and organize new sources of knowledge, integrate new knowledge into the existing corpus according to the above process, and continuously update and improve LOCAL_QA.
10. A non-temporary storage medium, characterized in that: It is used to store a program, which is used to execute a method for corpus enhancement processing of an intelligent water conservancy system as described in any one of claims 1 to 9 above.