Method and system for constructing hydrological survey knowledge model based on large model

By constructing a hydrological knowledge graph and using a large model to generate confused samples for adversarial training, the problem of knowledge transformation and updating in hydrological surveying education was solved. This enabled dynamic updating of trainees' knowledge status and generation of personalized learning plans, thereby improving the intelligence and practicality of hydrological surveying education.

CN120851181AActive Publication Date: 2025-10-28BUREAU OF HYDROLOGY CHANGJIANG WATER RESOURCES COMMISSION

Patent Information

Application Number
CN202511364262.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-10-28
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing technologies struggle to transform hydrological expertise into structured inputs that are understandable to large models, and they also fail to enable long-term memory and dynamic updates of learners' knowledge status, resulting in insufficient intelligence and practicality in hydrological surveying education.

Method used

We construct a hydrological knowledge graph, use a large model to generate confused samples and conduct adversarial training, update the knowledge status through bidirectional mapping of test questions and student answer data, and generate personalized learning plans.

Benefits of technology

It has enhanced the intelligence and practicality of hydrological survey education, realized the structured input of hydrological professional knowledge and the dynamic updating of students' knowledge status, and enhanced the adaptability and generalization ability of knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851181A_ABST
    Figure CN120851181A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hydrological survey skill training, in particular to a construction method and system of a hydrological survey knowledge model based on a large model. The method comprises the following steps: acquiring multi-source hydrological data, and constructing a hydrological knowledge graph; generating a confusion sample corresponding to the hydrological knowledge graph, and performing adversarial training on the confusion sample; according to the confrontation training result and historical test questions in the multi-source hydrological data, obtaining bidirectional mapping test questions; inputting the bidirectional mapping test questions into a question setting interface of a hydrological survey learning platform; collecting answer data and student physiological data corresponding to the bidirectional mapping test questions in real time, and updating a student knowledge state according to the answer data and the student physiological data; constructing a hydrological survey knowledge model based on the updated student knowledge state; and generating a student recommendation learning scheme by using the hydrological survey knowledge model. According to the invention, the intellectualization and practicability of hydrological survey education can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological survey skills training technology, and in particular to a method and system for constructing a hydrological survey knowledge model based on a large model. Background Technology

[0002] Hydrological surveying technology plays a vital role in water resource allocation, flood control and disaster reduction, and ecological environment monitoring. Traditional hydrological surveying mainly relies on field measurements, remote sensing image analysis, and numerical simulation methods based on physical laws to observe and analyze hydrological elements such as precipitation, runoff, evaporation, and groundwater. However, while existing technologies include rule-based intelligent test paper generation systems and simple machine learning recommendation models, they struggle to transform hydrological expertise (including standards, typical cases, and instrument operation procedures) into structured inputs that large models can understand, and they also cannot achieve long-term memory and dynamic updating of learners' knowledge status. Summary of the Invention

[0003] Therefore, the present invention needs to provide a method and system for constructing a hydrological survey knowledge model based on a large model, in order to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a method for constructing a hydrological survey knowledge model based on a large model includes the following steps: Step S1: Obtain multi-source hydrological data, define the entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph; Step S2: Generate confused samples corresponding to the hydrological knowledge graph using the pre-set large model, and perform adversarial training on the confused samples; obtain bidirectional mapping questions based on the adversarial training results and historical questions in multi-source hydrological data; Step S3: Input the bidirectional mapping test questions into the question generation interface of the hydrological survey learning platform; collect the answer data and student physiological data corresponding to the bidirectional mapping test questions in real time, and update the student's knowledge status based on the answer data and student physiological data; Step S4: Construct a hydrological survey knowledge model based on the updated knowledge status of trainees; use the hydrological survey knowledge model to generate recommended learning plans for trainees.

[0005] This application performs structured processing and semantic modeling on multi-source hydrological data to construct a hydrological knowledge graph. Then, it uses a large model to generate confused samples and combines them with historical test questions to form bidirectional mapping test questions. Through adversarial training, it improves the model's ability to identify and reason about knowledge points. Furthermore, it collects students' answer data and physiological data such as EEG in real time, using these as observations to dynamically update the students' mastery of each knowledge point in the knowledge graph, and adaptively generates personalized learning plans based on the updated results. Compared with existing technologies, this application not only transforms hydrological professional knowledge such as standards, typical cases, and instrument operation procedures into structured inputs that the large model can understand, but also enhances the adaptability and generalization ability of hydrological knowledge by introducing confused samples and bidirectional mapping test questions. Simultaneously, by combining multimodal student data to achieve dynamic updating and long-term memory of knowledge status, it effectively solves the problems of traditional learning systems such as difficulty in personalized recommendations and lagging knowledge updates, thereby significantly improving the intelligence and practicality of hydrological surveying education.

[0006] Optionally, this application also provides a system for constructing a hydrological survey knowledge model based on a large model, used to execute the method for constructing a hydrological survey knowledge model based on a large model as described above. The system for constructing a hydrological survey knowledge model based on a large model includes: The data acquisition module is used to acquire multi-source hydrological data, define entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph. The adversarial training module is used to generate confused samples corresponding to the hydrological knowledge graph using a pre-set large model, and to perform adversarial training on the confused samples; based on the adversarial training results and historical questions in multi-source hydrological data, bidirectional mapping questions are obtained. The knowledge status update module is used to input bidirectional mapping questions into the question generation interface of the hydrological survey learning platform; it collects the answer data and student physiological data corresponding to the bidirectional mapping questions in real time, and updates the student's knowledge status based on the answer data and student physiological data; The learning plan recommendation module is used to construct a hydrological survey knowledge model based on the updated knowledge status of trainees; and to generate recommended learning plans for trainees using the hydrological survey knowledge model.

[0007] The present invention relates to a system for constructing a hydrological survey knowledge model based on a large model. This system can implement any of the methods for constructing a hydrological survey knowledge model based on a large model according to the present invention. It is used to combine the operation and signal transmission media between various modules to complete the construction method of the hydrological survey knowledge model based on a large model. The modules within the system cooperate with each other, thereby improving the intelligence and practicality of hydrological survey education. Attached Figure Description

[0008] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the steps of the method for constructing a hydrological survey knowledge model based on a large model according to the present invention. Figure 2 In the embodiments of the present invention, the electroencephalogram (EEG) signals are... A schematic diagram of wave frequency variation data; Figure 3 This is a block diagram of the hydrological survey knowledge model construction system based on a large model in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0009] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0010] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0011] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0012] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a method for constructing a hydrological survey knowledge model based on a large model, the method comprising the following steps: Step S1: Obtain multi-source hydrological data, define the entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph; In some embodiments, structured hydrological data, including rainfall, river levels, runoff velocity, and evapotranspiration, can be obtained through interfaces connecting hydrological monitoring sensors and historical databases. Simultaneously, unstructured data such as hydrological research reports, geological survey archives, and remote sensing image text are collected. The data is first uniformly timestamped, and a distributed data cleaning algorithm is used to remove noisy data and missing value entries. Then, based on a hydrological dictionary, the unstructured text is segmented and labeled with terms, defining hydrological knowledge entities such as rivers, lakes, rainfall, evaporation, and aquifers. Semantic relationships between entities are established by calculating entity pairs with a co-occurrence frequency greater than 0.6 and entity pairs with a dependency syntactic relation strength greater than 0.7. Finally, the knowledge is organized into a graph structure using triples (entity-relation-entity) to construct a hydrological knowledge graph. This knowledge graph contains no fewer than 20,000 entity nodes and 50,000 semantic edges, and an indexing mechanism is implemented to support efficient subsequent retrieval and modeling.

[0013] In another embodiment, multi-source hydrological data and the defined entities and semantic relationships within the data can also be stored in a database. This database uses OCR to recognize the operational step text in the document, and LayoutLMv3 to parse equipment parameter tables (such as transmission frequency and beam angle) and calibration flowcharts. The labeled entities and relationships can be: labeled entities: ADCP (instrument type), calibration steps (instance type); relationships: ADCP - equipment used - current meter, calibration steps - compliance standard - T / CHES 61-2021. A segmentation example could be: dividing the document into semantic blocks such as equipment installation, parameter settings, and data correction, generating a 768-dimensional vector for each block; an index is created in Milvus, and when a student queries "ADCP data abnormality," it returns associated knowledge blocks such as parameter setting errors and equipment calibration procedures.

[0014] Step S2: Generate confused samples corresponding to the hydrological knowledge graph using the pre-set large model, and perform adversarial training on the confused samples; obtain bidirectional mapping questions based on the adversarial training results and historical questions in multi-source hydrological data; In another embodiment, when generating obfuscated samples using a pre-defined large model, a Transformer-based text generation model is selected. This model contains a 12-layer encoder and a 12-layer decoder, each layer including a self-attention module and a feedforward neural network module. It is pre-trained on a hydrological corpus to obtain domain-adaptive capabilities. During the construction of obfuscated samples, independent knowledge points and their adjacent semantic units are first selected based on the hydrological knowledge graph to form a context set. Dependency parsing is then performed on the semantic text of the independent knowledge points to extract core terms, which are then aligned with the context set. When the lexical similarity exceeds 0.75 or the semantic embedding cosine similarity exceeds 0.8, the adjacent semantic units are identified as obfuscation candidates. Subsequently, three types of obfuscated samples are generated based on these candidate points: replacing core terms with candidate point terms results in term substitution obfuscated samples; exchanging context relationships results in relation inversion obfuscated samples; and introducing principal component features of candidate points into the semantic text results in semantic interference obfuscated samples. The confused samples are then input into an adversarial training framework, which employs an alternating optimization approach between a generator and a discriminator. The generator rewrites semantics based on a large model, while the discriminator determines the degree of confusion of the samples using a convolutional neural network. The adversarial training iterates for 20 rounds to enhance model robustness. A confused semantic feature library is generated based on the adversarial training results, and dependency syntax annotation is performed using historical test questions to extract the distribution of knowledge points and determine the semantic framework of the test questions. Subsequently, confused semantic features are injected into the semantic framework to construct three types of test questions: discrimination tests, logical approximations, and reasoning rewriting, which are then merged into a confused test question set in a 4:3:3 ratio. Finally, a bidirectional mapping relationship is established between the confused test question set and the knowledge graph link, forming bidirectional mapped test questions.

[0015] In other embodiments, the large-scale model can employ the Qwen-14B model, trained using 50,000 hydrological documents, focusing on the semantic representation of terms in areas such as water level-discharge relationships and sediment particle analysis. Historical test questions can be input to briefly describe the main error sources in ADCP flow measurement. The model outputs test questions containing a knowledge graph path (flow measurement → current meter method → ​​error analysis) and a capability mapping matrix (equipment operation capability 0.3, data analysis capability 0.4), which are then incorporated into the training set after expert review.

[0016] Step S3: Input the bidirectional mapping test questions into the question generation interface of the hydrological survey learning platform; collect the answer data and student physiological data corresponding to the bidirectional mapping test questions in real time, and update the student's knowledge status based on the answer data and student physiological data; In another embodiment, after the bidirectional mapped test questions are input into the question-generating interface of the hydrological survey learning platform, student answer data and physiological data are collected in real time. Answer data includes accuracy rate, time spent on a single question, and eye-tracking trajectory, with a sampling frequency of 60 Hz; physiological data consists of electroencephalogram (EEG) signals collected via a portable EEG device, from which key extraction is performed. Waves and The power spectral density of the wave is calculated, and the frequency change is calculated using a sliding time window of 500ms. Then, an extended Kalman filter algorithm is used to combine the answer data and physiological data as observation vectors with the knowledge state of the previous moment to update the student mastery level of each knowledge point node in the knowledge graph. Mastery level is represented by a probability value of 0-1. For knowledge point nodes that have not had any answer records for more than 30 days, the mastery level is decayed according to a decay factor of 0.95. The system also statistically analyzes answer error patterns, categorizing question types with an accuracy rate below 70% into common error patterns and updating the error distribution in the knowledge graph to characterize the student's weaknesses.

[0017] Step S4: Construct a hydrological survey knowledge model based on the updated knowledge status of trainees; use the hydrological survey knowledge model to generate recommended learning plans for trainees.

[0018] In another embodiment, when constructing the learning model based on the updated learner knowledge state, a graph neural network (GNN)-based inference model is employed. This model can learn node representations within the knowledge graph, thereby capturing the upstream and downstream dependencies between knowledge points. The model outputs a mastery vector for each knowledge point. When generating a recommended learning plan for learners using this model, a difficulty threshold for each knowledge point is first defined based on the difficulty distribution of bidirectional mapping questions, with the threshold set to 0.6. Then, comparing the learner's mastery vector, if the accuracy rate for three consecutive questions is higher than 85%, the mastery level of the corresponding knowledge point is increased by 0.1 levels; if the time taken for a single question exceeds 25 seconds, the mastery level of that knowledge point is decreased by 0.1 levels. The subsequent learning path is dynamically adjusted based on the upgraded and downgraded mastery levels, for example, prioritizing knowledge points with a mastery level below 0.5, or increasing the proportion of related questions to 40% of the total questions. The final recommended learning plan includes not only the order of knowledge point learning but also the number of questions, difficulty distribution, and learning time arrangement to achieve personalized adaptive training for learners.

[0019] Optionally, the process of generating a recommended learning plan for students in step S4 includes: The difficulty threshold of knowledge points is defined based on the difficulty level of the two-way mapping test questions; In one embodiment, the system first defines the difficulty threshold of knowledge points based on the existing difficulty parameters in the bidirectional mapping questions. Specifically, the bidirectional mapping questions are constructed with initial difficulty values ​​set according to the question discrimination and semantic complexity of confusion, and these difficulty values ​​are distributed in the range of 0 to 1. Using a quantile-based approach, questions with difficulty values ​​below 0.3 are classified as low difficulty, and the corresponding knowledge point difficulty threshold is set to the basic level; questions with difficulty values ​​between 0.3 and 0.7 are classified as medium difficulty, and the corresponding knowledge point difficulty threshold is set to the advanced level; questions with difficulty values ​​above 0.7 are classified as high difficulty, and the corresponding knowledge point difficulty threshold is set to the challenge level. This ensures that the difficulty threshold of knowledge points is derived from the statistical characteristics of the questions, rather than being set arbitrarily, which can form a clear hierarchical structure in the knowledge graph.

[0020] Based on the knowledge point difficulty threshold, and using the hydrological survey knowledge model to generate the updated knowledge status of trainees, the corresponding knowledge point mastery level is determined. In another embodiment, a hydrological survey knowledge model based on graph neural networks is used to generate the updated learner knowledge state. This model contains three graph convolutional layers and two fully connected layers, enabling it to learn representations of nodes in the hydrological knowledge graph. It also combines learner performance with upstream and downstream dependencies of knowledge points, outputting a mastery vector for each knowledge point, with values ​​ranging from 0 to 1. A difficulty threshold is used as a threshold value; when a learner's mastery level is below the threshold, the knowledge point is marked as a node requiring reinforcement; when the learner's mastery level is above the threshold, it is marked as a mastered node.

[0021] In other embodiments, a hydrological survey knowledge model can be used to determine whether the difficulty level of the questions is suitable for the learner based on their updated knowledge status. For example, if a learner answers a question correctly but... A high cognitive load indicates that the knowledge point has not been mastered; conversely, a low cognitive load and correct answers indicate that the learner has a strong grasp of the knowledge point.

[0022] In another embodiment, combined The fluctuations in the wave pattern can dynamically determine whether the difficulty of the questions is suitable for the learners: for example, high cognitive load + long-term fixation may indicate that the difficulty of the questions is too high, and the complexity of the questions may be reduced or the steps may be increased in subsequent training; low cognitive load + high accuracy may indicate that the difficulty of the questions is moderate or too low, and the training difficulty may be increased or reasoning questions may be added.

[0023] The difficulty of the corresponding knowledge point is adjusted based on the accuracy rate of the answer data. If the accuracy rate of three consecutive questions in the answer data is greater than 85%, the student's mastery of the corresponding knowledge point will be upgraded by one level, resulting in an upgraded knowledge point mastery. If the time taken for any question in the answer data is greater than 25 seconds, the student's mastery of the corresponding knowledge point will be downgraded by one level, resulting in a downgraded knowledge point mastery. In another embodiment, the difficulty of knowledge points is dynamically adjusted based on the student's answer data. If the accuracy rate of answering three consecutive questions exceeds 85%, the student is considered to have a high level of stable mastery over that knowledge point, and the mastery level for that knowledge point is increased by 0.1, with a maximum of 1.0. If a student spends more than 25 seconds on a question related to a particular knowledge point, the student is considered to have a high cognitive load or a problem with their comprehension path, and the mastery level for that knowledge point is decreased by 0.1, with a minimum of 0.0. These adjustments are performed at the knowledge point level, not at the question level, thus avoiding over-correction due to accidental errors or external interference. The adjustment history is also recorded; if the mastery level for the same knowledge point decreases more than three times within a week, it is automatically marked as a key or difficult point, and the frequency of its recurrence is increased in subsequent learning plans.

[0024] Based on the mastery of upgraded and downgraded knowledge points, the bidirectional mapping test questions adaptively adjust the learning order and test question ratio of each student in the hydrological survey learning platform to obtain a recommended learning plan for each student.

[0025] In another embodiment, the bidirectional mapping questions are adaptively adjusted based on the upgraded and downgraded mastery levels. Specifically, in the hydrological surveying learning platform, each student's question bank is dynamically generated, and the learning order and question percentage for each knowledge point are determined by their latest mastery vector. When the mastery level of a knowledge point increases to above 0.8, the percentage of subsequent questions for that knowledge point will automatically decrease to 15% of the total questions to avoid redundant training; conversely, when the mastery level of a knowledge point decreases to below 0.4, the percentage of questions for that knowledge point will increase to 35%, and the percentage of related upstream and downstream knowledge points will also increase by 5% to strengthen associative learning. The final recommended learning plan for students includes: knowledge point ranking, the distribution of questions for each knowledge point, the distribution of question difficulty (approximately 30% for low difficulty, approximately 40% for medium difficulty, and approximately 30% for high difficulty), and the estimated completion time. In this way, the learning platform can provide personalized hydrological surveying training paths for different students, achieving continuous optimization of knowledge point mastery.

[0026] Optionally, step S1 includes: Step S11: Obtain multi-source hydrological data and define knowledge entities based on the structured multi-source hydrological data; In one embodiment, raw data is first obtained from multi-source hydrological data, including real-time flow data from hydrological monitoring stations, geological survey reports, remote sensing images, and meteorological observation data. To ensure the data can be used for subsequent knowledge extraction, an ETL (Extract-Transform-Load) process is used to structure the data. This includes preprocessing text data by sentence segmentation and denoising, normalizing time-series numerical data, and performing semantic segmentation on remote sensing images to extract water body boundary information. After the structured processing, terminology recognition is performed on the text corpus based on a hydrological dictionary, and information such as monitoring station numbers, flow indicators, and geological structure types are defined as knowledge entities. Each entity is stored in the Neo4j graph database with a unique ID, serving as a basic node in the knowledge graph.

[0027] Step S12: Based on the co-occurrence frequency and semantic dependency relationship of each knowledge entity in the multi-source hydrological data, define the semantic association relationship of the knowledge entities, and construct an initial hydrological knowledge graph based on the semantic association relationship between knowledge entities. In another embodiment, semantic dependencies between entities are identified by calculating the frequency of co-occurrence of knowledge entities within the same or cross-data sources and combining this with dependency parsing. For example, if a semantic dependency exists in a hydrological monitoring report indicating an increase in rainfall and runoff, an edge with a causal attribute is created in the knowledge graph. The point mutual information (PMI) method is used to calculate a co-occurrence frequency threshold; only when the PMI values ​​of two entities are greater than 0.5 are they considered to have a statistically significant association. Simultaneously, a dependency parsing tree is introduced, using the verbs "cause" and "affect" as semantic labels for the edges, thereby establishing semantic associations between entities. Ultimately, the entities and their semantic associations together constitute the initial hydrological knowledge graph, where nodes represent hydrological entities and edges represent different types of semantic dependencies.

[0028] Step S13: Divide the multi-source hydrological data into blocks according to the granularity of each knowledge point, and optimize the block division results by combining them with a preset hydrological professional dictionary to obtain a set of hydrological knowledge blocks; In another embodiment, the structured multi-source hydrological data is segmented according to the granularity of knowledge points. First, the text corpus is segmented into sentences and paragraphs, and a hydrological dictionary is used to detect the number of hydrological terms contained in each sentence and their semantic dependency path coverage. When the coverage reaches 60% or more, the sentence is considered a candidate knowledge point. Subsequently, the similarity value between adjacent sentences is calculated based on a word vector model. If the similarity is less than 0.4, the candidate sentence is classified as an independent knowledge point to ensure the distinguishability and independence of the segmented knowledge points. Further, the number of terms in each independent knowledge point is used as a complexity indicator and combined with the similarity value to generate a knowledge point granularity score. This score ranges from 1 to 5 points, with higher scores indicating finer granularity of the knowledge point.

[0029] Step S14: Construct a hydrological knowledge graph using the initial hydrological knowledge graph as the index structure and the set of hydrological knowledge blocks as the content carrier to be indexed.

[0030] In another embodiment, an initial hydrological knowledge graph is used as the index structure, and a set of hydrological knowledge blocks is used as the indexed content carrier to construct a complete hydrological knowledge graph. Specifically, each knowledge block is bound to a corresponding entity node in the graph. For example, the knowledge block "river flow change pattern" will be associated with three entity nodes: river, flow, and change. An inverted index mechanism is used to ensure a bidirectional mapping between knowledge blocks and entities, ensuring that when users query the knowledge graph, they can retrieve relevant knowledge blocks through entity nodes and quickly locate relevant entity nodes through knowledge blocks. Simultaneously, weight values ​​are assigned to the mapping relationship between knowledge blocks and entities, with the weight determined by the semantic similarity between the knowledge block and the entity; the higher the similarity, the greater the weight. The resulting hydrological knowledge graph not only contains entities and relationships but also includes an extended semantic layer with knowledge blocks as content, achieving a unified approach to knowledge indexing and content creation.

[0031] Optionally, the calculation method for the granularity of knowledge points in step S13 includes: Text segmentation and paragraph splitting were performed on multi-source hydrological data to obtain several initial semantic units; Text segmentation and paragraph splitting were performed on multi-source hydrological data. The data sources included daily reports from hydrological observation stations, hydrogeological survey reports, and water resource utilization planning documents. A sentence segmenter combining rule-based and statistical methods was employed. The rule-based part relied on punctuation marks (such as periods and semicolons) as delimiters, while the statistical part used a BiLSTM-CRF model to predict potential semantic breakpoints. In the experimental setup, sentence lengths were controlled between 15 and 30 words, and the maximum paragraph length was set to 150 words. The segmentation results were several initial semantic units, which were stored in a structured data table as input for subsequent knowledge point extraction.

[0032] Based on the detection of the number of hydrological terms and the semantic dependency relationship between hydrological terms in each initial semantic unit using a hydrological professional dictionary, if the dependency path coverage rate of the semantic dependency relationship between hydrological terms in the semantic unit reaches a preset coverage threshold, then the semantic unit is regarded as a candidate knowledge point. In another embodiment, the number of hydrological terms and their semantic dependencies in each initial semantic unit are detected based on a hydrological dictionary. This dictionary contains approximately 15,000 manually verified terms, covering areas such as hydrological processes, geological structures, and monitoring indicators. First, terms are labeled using dictionary matching, and then semantic dependency paths are extracted using a dependency parser (a modified version of Stanford Parser). The dependency path coverage rate is calculated as: Coverage = (Number of dependency relationships between terms ÷ Total number of term pairs) × 100%. A coverage threshold of 60% is set; that is, when the dependency path coverage rate between terms in a semantic unit exceeds 60%, the unit is identified as a candidate knowledge point. For example, in the sentence "Heavy rainfall leads to a rapid increase in surface runoff," the terms "rainfall" and "surface runoff" establish a causal dependency through "cause," exceeding the coverage threshold; therefore, this sentence is marked as a candidate knowledge point.

[0033] Calculate the similarity value representation between the candidate knowledge point and its adjacent initial semantic units. If the similarity is lower than the preset similarity threshold, then the candidate knowledge point is treated as an independent knowledge point. In another embodiment, the similarity value between the candidate knowledge point and its adjacent initial semantic units is calculated to determine whether the knowledge point is independent. The similarity calculation is based on a dual-channel model: on the one hand, Word2Vec is used to generate semantic vectors and calculate cosine similarity; on the other hand, the Jaccard coefficient is used to measure the overlap of the term sets. The two indicators are fused with weights of 0.7 and 0.3 to obtain the final similarity score. In practical applications, the similarity threshold is set to 0.4. When the similarity between a candidate knowledge point and its adjacent semantic units is less than 0.4, the candidate knowledge point is identified as an independent knowledge point. For example, the fused similarity between the candidate point "river erosion" and its adjacent semantic unit "soil loss" is only 0.32, therefore it is identified as an independent knowledge point.

[0034] The number of hydrological terms for independent knowledge points is used as a knowledge point complexity index, and the knowledge point granularity is generated by combining the knowledge point complexity index with the similarity value.

[0035] In another embodiment, the number of hydrological terms in an independent knowledge point is used as the knowledge point complexity index, and its granularity is generated by combining it with its similarity value. The complexity index is directly quantified by the number of terms, ranging from 1 to 10, with a higher value indicating more specialized concepts involved in the knowledge point. Based on this, a weighted scoring function is constructed: Granularity = a × Complexity Index + b × (1 - Similarity Value), where a = 0.6 and b = 0.4. This function considers both the specialized complexity within the knowledge point and its semantic differences from adjacent units. Finally, the granularity is divided into 5 levels. For example, the rainfall intensity classification contains only 2 terms and has a high similarity to adjacent units, so its granularity is level 1; while the watershed runoff generation process simulation contains 8 terms and has a similarity of only 0.25 with adjacent units, so its granularity is level 5.

[0036] Optionally, the confused samples generated in step S2 corresponding to the hydrological knowledge graph include: Based on the hydrological knowledge graph, independent knowledge points and their adjacent initial semantic units are selected to construct a semantic context set for independent knowledge points; In one embodiment, independent knowledge points and their adjacent initial semantic units are selected based on a hydrological knowledge graph to construct a semantic context set for each independent knowledge point. Specifically, the knowledge graph consists of three parts: knowledge entities, semantic relations, and knowledge block indexes. Entity nodes represent hydrological professional concepts, and semantic edges represent dependency or causal relationships. Centered on an independent knowledge point, its first-order adjacent nodes are retrieved, and the context window size is limited to three semantic units to obtain the semantic context set. For example, the adjacent units of the independent knowledge point surface runoff in the graph include rainfall and runoff processes, which together with surface runoff form the semantic context set.

[0037] Using a large model, dependency parsing is performed on the semantic description text corresponding to independent knowledge points in multi-source hydrological data, and core term extraction is performed based on the syntactic analysis results to obtain the core terms of independent knowledge points. In another embodiment, a pre-defined large model is used to perform dependency parsing on the semantic description text corresponding to independent knowledge points in multi-source hydrological data, and core terms are extracted based on the analysis results. The large model is a Transformer-based Chinese RoBERTa-wwm-ext structure, and its input is a hydrological description sentence, such as heavy rainfall leading to increased surface runoff. The large model first generates a context vector representation of the input text, and then constructs a syntactic dependency tree using syntactic analysis tools (such as LTP and Stanza). Sub-paths centered on causal and modification relationships are selected on the dependency tree, and nouns and verbs directly related to independent knowledge points are extracted as candidate core terms. These are then further validated using a hydrological dictionary to obtain the core terms for each independent knowledge point.

[0038] Align the semantic context set of independent knowledge points with the core terms of independent knowledge points. If the lexical similarity or semantic similarity of the alignment result exceeds the preset corresponding threshold, then the adjacent initial semantic unit is regarded as a candidate knowledge point for confusion of independent knowledge points. In another embodiment, the semantic context set of independent knowledge points is aligned with their core terms, and lexical similarity and semantic similarity are calculated to filter out candidate knowledge points for confusion. Lexical similarity uses the Jaccard coefficient with a threshold of 0.5; semantic similarity is based on vector representations generated by the Sentence-BERT model, with a cosine similarity threshold of 0.7. When any adjacent semantic unit in the context set meets the similarity requirement with the core term at the lexical or semantic level, that adjacent unit is identified as a candidate knowledge point for confusion. For example, the semantic similarity between the core term "runoff" and the adjacent semantic unit "flow measurement" is 0.76, exceeding the threshold; therefore, "flow measurement" is marked as a candidate knowledge point for confusion.

[0039] Based on the semantic description text of the confused candidate knowledge points in multi-source hydrological data, multiple types of confused samples are generated.

[0040] In another embodiment, multiple types of confused samples are generated based on the semantic description text of confused candidate knowledge points in multi-source hydrological data. Specifically, sample generation includes three strategies: first, a differentiation test strategy, which replaces core terms in the description text with confused candidate knowledge points to form test samples that are semantically similar but have different answers; second, a logical approximation strategy, which adjusts the causal relationships or constraints in the text, for example, rewriting "increased rainfall intensity leads to increased runoff" as "increased rainfall duration leads to increased runoff"; and third, a reasoning rewriting strategy, which replaces upstream or downstream nodes in the knowledge point chain, for example, rewriting "changes in runoff volume" as "changes in sediment content," thereby introducing reasoning differences. The ratio of the three types of sample generation can be set to 4:3:3, and the final confused sample set is approximately 120% the size of the original knowledge point samples.

[0041] Optionally, the multiple types of obfuscated samples include: By replacing the core terms in independent knowledge points with the core terms of the confused candidate knowledge points, term substitution-type confusion samples are obtained; In one embodiment, terminology substitution-type obfuscation samples are obtained by replacing the core terms in independent knowledge points with the core terms in obfuscated candidate knowledge points. The core terms of independent knowledge points are extracted using Transformer combined with dependency parsing, while the core terms of obfuscated candidate knowledge points are similarly obtained through correction using a hydrological dictionary. During the generation process, a core term alignment table is first established, and the similarity between terms is calculated. Word vectors are generated using the Word2Vec model. If the cosine similarity value is greater than 0.6, substitution is allowed. For example, the core term corresponding to the independent knowledge point "runoff" is "flow rate," and the core term corresponding to the obfuscated candidate knowledge point "flow velocity" is "water flow velocity." The similarity is 0.67, exceeding the threshold. Therefore, in the text describing "increased rainfall intensity leads to increased flow rate," "flow rate" is replaced with "water flow velocity," thus forming a terminology substitution-type obfuscation sample.

[0042] Based on the semantic description text of confused candidate knowledge points and independent knowledge points in multi-source hydrological data, the contextual relationship between confused candidate knowledge points and independent knowledge points is exchanged to obtain a reversed confusion sample. In another embodiment, based on the semantic description texts of confused candidate knowledge points and independent knowledge points in multi-source hydrological data, a relation-reversed confused sample is obtained by exchanging their contextual relationships. In implementation, the Sentence-BERT model is used to generate contextual semantic vectors for independent knowledge points and confused candidate knowledge points, and the directionality of contextual dependency edges is detected. If their directions are opposite on core semantic paths such as causal relationships and modification relationships, the exchange is allowed. For example, the dependency path between the independent knowledge point "increased rainfall" and the confused candidate knowledge point "increased runoff" in the original text is "increased rainfall → increased runoff". Exchanging it to "increased runoff → increased rainfall" yields a relation-reversed confused sample. To prevent the generation of semantically incomprehensible samples, the readability score (calculated using the perplexity level (PPL) of the language model) after the exchange is required to be no more than 1.2 times that of the original text.

[0043] By introducing principal component semantic features of confusing candidate knowledge points into the semantic description text of independent knowledge points, semantic interference-type confusing samples are obtained.

[0044] In another embodiment, semantically obfuscated samples are obtained by introducing principal component semantic features of obfuscated candidate knowledge points into the semantic description text of independent knowledge points. This process first uses PCA (Principal Component Analysis) to reduce the dimensionality of the description text vectors of obfuscated candidate knowledge points, extracts the first two principal components as their semantic principal features, and integrates them into the semantic description of the independent knowledge points through vector interpolation. Operationally, weighted... Linear interpolation is performed with a value of 0.3, using the following formula: ,in For each knowledge point, a semantic vector. To obfuscate candidate knowledge points, the principal component semantic vector is used. Taking "evapotranspiration intensity" as an independent knowledge point and "soil moisture" as an obfuscated candidate knowledge point, the principal component semantic related to "soil moisture" is injected into the descriptive text "evapotranspiration intensity depends on temperature and wind speed". This causes the text to generate "evapotranspiration intensity depends on temperature, wind speed and potential soil moisture conditions", thus constructing a semantic interference obfuscated sample.

[0045] Of particular importance is that among the various types of obfuscated samples, the proportions of term substitution obfuscated samples, relation inversion obfuscated samples, and semantic interference obfuscated samples are 40%, 30%, and 30% of the total obfuscated samples, respectively.

[0046] Optionally, obtaining the bidirectional mapping questions in step S2 includes: Construct a confusing semantic feature library based on the adversarial training results; In one embodiment, a confusion semantic feature library is first constructed based on the results of adversarial training. Specifically, the adversarial training employs a combination of the adversarial sample generation framework FGSM (Fast Gradient Sign Method) and PGD (Projected Gradient Descent) to model semantic perturbations in the already generated confusion samples in the hydrological knowledge graph. The RoBERTa-large model is used as the semantic encoder, inputting the original hydrological terms and their contextual descriptions, and the perturbation direction in the word vector space is obtained through adversarial gradient calculation. During iterative updates, the perturbation strength parameter ε is set to 0.15 to ensure semantic readability. After multiple rounds of adversarial training, frequently occurring perturbation semantic features are extracted, such as the inversion of quantitative indicators and conditional factors, and the replacement of physical quantity boundary conditions, and stored as the confusion semantic feature library. This feature library serves as a perturbation factor for subsequent question rewriting, providing systematic confusion strategy support for question generation.

[0047] Historical test texts were extracted from multi-source hydrological data, and dependency parsing was performed on the test texts. The distribution of test knowledge points was marked based on the results of the dependency parsing. The semantic framework of the test questions was determined based on the distribution of test knowledge points. In another embodiment, historical test text is extracted from multi-source hydrological data and parsed at the syntactic and semantic levels. Specifically, dependency parsing is first performed on the historical test text. A dependency parser based on the biaffine attention mechanism is used to generate a dependency tree, identifying subject-verb-object structures, modification relations, and causal links within the dependency tree. The dependency tree nodes are then labeled using a hydrological dictionary, marking terms related to "hydrological processes," "hydrological factors," and "calculation parameters" as candidate knowledge points, and storing the distribution of test knowledge points in a graph structure. For example, in the test question "Calculate the runoff under a certain rainfall intensity in a watershed," "rainfall intensity" is labeled as an upstream knowledge point, and "runoff" as a downstream knowledge point, forming a clear knowledge point link distribution, which serves as the basic structure of the test question's semantic framework.

[0048] Using a confused semantic feature library as a perturbation factor, the semantic framework of the test questions is rewritten at multiple levels to generate bidirectional mapping test questions.

[0049] In another embodiment, a confusion semantic feature library is used as a perturbation factor to perform multi-layered rewriting of the test question semantic framework to generate bidirectional mapping test questions. Specifically, this includes three rewriting strategies: First, terminology-level rewriting, utilizing replacement terms in the confusion semantic feature library with a semantic similarity exceeding 0.65 to the core terms to perform term substitution, thereby generating terminology-substitution type test questions; second, logic-level rewriting, adjusting constraints in dependency syntax, such as adjusting numerical boundaries or causal order, to ensure that the PPL (Perplexity Points) of the rewritten text is less than 1.2 times that of the original test questions, thereby generating logic-approximation type test questions; third, reasoning chain rewriting, inserting or replacing nodes upstream or downstream of the test question knowledge point chain to construct reasoning-expanded type test questions. Finally, these are merged into a bidirectional mapping test question set according to a preset ratio (40% terminology substitution type, 30% logic-approximation type, and 30% reasoning-expanded type), and a bidirectional mapping relationship is established with the original knowledge point chain.

[0050] Optionally, performing multi-layered rewriting of the semantic framework of the test questions includes: Replace or insert confusing semantic terms from the confusion semantic feature library into the terms in the semantic framework of the test questions to construct distinguishing test questions; In one embodiment, core terms in the semantic framework of the test question are replaced or inserted into confusing semantic terms in the confusing semantic feature library to construct a distinguishing test question. Specifically, the SimCSE semantic coding model is used to calculate the semantic similarity between the core terms of the test question and the terms in the confusing semantic feature library. If the similarity score is between 0.65 and 0.80, the confusing term is used as a replacement candidate. To ensure the solvability of the test question, a beam search strategy is used to select the optimal term combination from the candidate set, and perplexity detection (PPL≤120) is performed on the rewritten result to control the fluency of the language. Through the above replacement operation, the river flow in the original test question can be rewritten as the river section water volume, forming a distinguishing test question that is semantically distracting but logically feasible.

[0051] Adjust the constraints in the question stem of the semantic framework of the test questions to construct logically approximate test questions; In another embodiment, based on the question stem structure of the question semantic framework, the constraints in the question stem are adjusted to generate a logically approximate question. Specifically, the condition constraint nodes in the syntactic dependency tree of the question stem are identified, such as "rainfall intensity = 30 mm / h" and "boundary condition = infiltration rate 0.25". Combined with the perturbation factor in the obfuscation semantic feature library, these constraints are adjusted, with the adjustment range controlled within ±15% to maintain the reasonableness of the question's solution range. Simultaneously, a numerical verification module checks whether a unique solution exists in the solution space of the rewritten question; if multiple feasible solutions exist, a backtracking adjustment is performed. Thus, the original question "Calculate the watershed runoff under a rainfall intensity of 30 mm / h" can be rewritten as "Calculate the watershed runoff under a rainfall intensity of 35 mm / h", forming a logically approximate question.

[0052] Rewrite the upstream or downstream knowledge points in the knowledge point chain within the semantic framework of the test question, and construct reasoning to rewrite the test question; In another embodiment, the knowledge point links in the semantic framework of the test questions are rewritten to construct reasoning-rewritten test questions. Specifically, the upstream and downstream knowledge points corresponding to the target test question are identified using the knowledge point link graph as an index. For example, when the original knowledge point link is "rainfall intensity → runoff", a new adjacent knowledge point "soil moisture content" can be introduced as the upstream knowledge point based on the confusion semantic feature library, or the downstream knowledge point can be replaced with "peak flow". During the rewriting process, the knowledge point dependency consistency check module ensures that the rewritten link maintains the rationality of causal logic, and the completeness of the rewritten knowledge points is verified by link coverage calculation (coverage ≥ 0.85). This method can effectively improve the test questions' ability to assess students' reasoning ability and generate reasoning-rewritten test questions.

[0053] It is worth noting that the knowledge point dependency consistency verification module is an independent computing unit built on the hydrological knowledge graph, used to ensure the rationality of the causal logic of the knowledge point links in the test questions after rewriting or expansion. The construction process of this module includes the following steps: First, all knowledge point nodes and their directed dependency edges are extracted from the hydrological knowledge graph to construct a knowledge point dependency matrix. Rows in the matrix represent upstream knowledge points, columns represent downstream knowledge points, and matrix elements are set to dependency strength, ranging from 0 to 1. Larger values ​​indicate stronger causal relationships between knowledge points. Second, the module defines a dependency consistency threshold (e.g., 0.85) to determine whether the rewritten link maintains the original causal logic. Then, when question rewriting involves adjustments to upstream or downstream knowledge points, the module calculates the coverage rate of the rewritten link in the original knowledge graph using matrix multiplication. Coverage rate = sum of edge weights in the rewritten link that conform to the original dependency direction / total edge weight of the link. If the coverage rate is lower than the threshold, link rollback or reselection of adjacent knowledge points is triggered. Finally, the module provides a configurable interface that allows adjustment of dependency strength weights for different disciplines or hydrological scenarios to adapt to knowledge point networks of varying complexity.

[0054] The test questions, logic approximation questions, and reasoning rewriting questions are merged into a confusing question set according to a preset question allocation ratio. The confusing question set is then mapped to the corresponding knowledge point links to obtain two-way mapped questions.

[0055] In another embodiment, differentiation test questions, logical approximation questions, and reasoning rewriting questions are merged into a confused question set according to a preset question allocation ratio, and a mapping relationship with knowledge point links is established to obtain bidirectional mapped questions. Specifically, according to the set ratio parameters, differentiation test questions account for 40%, logical approximation questions account for 30%, and reasoning rewriting questions account for 30%, and each type of question is stored in a hierarchical index in the question bank. To ensure effective alignment between the confused question set and the knowledge point links, a graph matching algorithm is used to compare the knowledge point distribution of each question with the original links. If the link matching degree exceeds 0.9, a formal mapping relationship is established; otherwise, a fine-tuning rewriting is performed. The final bidirectional mapped questions can maintain the correspondence with the original knowledge point links and have diverse confusion features, which can be used for subsequent student knowledge status assessment and dynamic learning recommendations.

[0056] Optionally, updating the student's knowledge status based on the answer data and the student's physiological data in step S3 includes: Real-time collection of answer data corresponding to bidirectional mapping test questions, including answer accuracy, time taken, and student eye movement trajectory; In one embodiment, a hydrological survey learning platform collects students' answer data for bidirectional mapping questions in real time. The answer data includes accuracy, time taken, and eye-tracking patterns. Accuracy is recorded as a percentage and stored in the answer data matrix immediately after each question is completed. Time taken is measured in seconds, and outliers are processed using a moving average window (window length of 3 questions) to form a time series vector. Students' eye-tracking patterns are collected by a wearable eye-tracking device at a sampling frequency of 120Hz. After noise filtering and coordinate standardization, a fixation point matrix and fixation duration matrix are generated for analyzing attention distribution and answering strategies.

[0057] Collect physiological data from trainees, including electroencephalogram (EEG) signals. Wave frequency variation data; In another embodiment, physiological data of the trainees is collected, primarily consisting of electroencephalogram (EEG) signals. The relative energy changes in the wave (8-12Hz) frequency band are used to reflect the cognitive load of trainees. EEG signals are acquired using a head-mounted multi-channel acquisition device at a sampling frequency of 250Hz. The data is first filtered by a 0.5Hz high-pass filter and a 50Hz low-pass filter to remove power frequency noise and electromyographic interference, and then... The wave power is subjected to a short-time Fourier transform to obtain a time series of cognitive load indicators per second, which is then synchronized and matched with the answer data according to timestamps.

[0058] Using students' physiological data and answer data as observations, and combining them with the students' knowledge status at the previous moment, the students' mastery of each knowledge point node in the hydrological knowledge graph is updated. In another embodiment, the collected answer data and trainee EEG data are used as observations input into the extended Kalman filter update model. Simultaneously, the trainee's knowledge state from the previous time step is combined to update the trainee's mastery of each knowledge point node in the hydrological knowledge graph. The mastery of each knowledge point node is represented by a floating-point number between 0 and 1, and each node is updated independently. The model sets the state noise parameter Q=0.01 and the observation noise parameter R=0.05, and calculates the node mastery through a prediction-update loop, outputting an updated knowledge state matrix to reflect the trainee's current mastery of each knowledge point.

[0059] For knowledge point nodes in the hydrological knowledge graph that have not had corresponding answer data for more than 30 days, the student's mastery of the knowledge point node is reduced according to the preset decay factor; the error pattern distribution of knowledge point nodes is updated by statistically analyzing the question types with low answer accuracy; and the student's knowledge status is updated based on the update status of each knowledge point node in the hydrological knowledge graph.

[0060] Most importantly, the update status of each knowledge point node includes the update status of students' mastery of the knowledge point node and the update status of the error pattern distribution of the knowledge point node.

[0061] In another embodiment, for knowledge point nodes in the hydrological knowledge graph that have not had corresponding question data for more than 30 days, the system performs exponential decay using a decay factor of 0.85. The decay formula is as follows: ,in The cumulative number of days beyond 30 days ensures a reasonable decline in the mastery of knowledge points that have not been reviewed for a long time, while also providing an initial estimate for the next test and simulating the knowledge forgetting process.

[0062] In another embodiment, the types of questions with low accuracy rates are statistically analyzed, and the error pattern distribution of relevant knowledge point nodes in the hydrological knowledge graph is updated. Error patterns can include "conceptual confusion," "calculation errors," and "misunderstanding biases," etc., and are recorded as percentages. The proportion of each type is dynamically adjusted based on the answer records of the most recent 30 days, forming an error pattern set, which is jointly stored with the knowledge point node mastery matrix to provide a basis for subsequent personalized learning strategies for students.

[0063] In another embodiment, the overall knowledge status of trainees is updated synchronously based on the update status of each knowledge point node in the hydrological knowledge graph. The update operation integrates the mastery level of each knowledge point node with the corresponding error pattern set, calculates the trainees' overall mastery level of hydrology through weighted averaging, and records the trend of mastery changes to provide input for the generation of subsequent recommended learning plans.

[0064] Optionally, this application also provides a system for constructing a hydrological survey knowledge model based on a large model, used to execute the method for constructing a hydrological survey knowledge model based on a large model as described above. The system for constructing a hydrological survey knowledge model based on a large model includes: The data acquisition module is used to acquire multi-source hydrological data, define entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph. The adversarial training module is used to generate confused samples corresponding to the hydrological knowledge graph using a pre-set large model, and to perform adversarial training on the confused samples; based on the adversarial training results and historical questions in multi-source hydrological data, bidirectional mapping questions are obtained. The knowledge status update module is used to input bidirectional mapping questions into the question generation interface of the hydrological survey learning platform; it collects the answer data and student physiological data corresponding to the bidirectional mapping questions in real time, and updates the student's knowledge status based on the answer data and student physiological data; The learning plan recommendation module is used to construct a hydrological survey knowledge model based on the updated knowledge status of trainees; and to generate recommended learning plans for trainees using the hydrological survey knowledge model.

[0065] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0066] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for constructing a hydrological survey knowledge model based on a large model, characterized in that, Includes the following steps: Step S1: Obtain multi-source hydrological data, define the entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph; Step S1 includes: Step S11: Obtain multi-source hydrological data and define knowledge entities based on the structured multi-source hydrological data; Step S12: Based on the co-occurrence frequency and semantic dependency relationship of each knowledge entity in the multi-source hydrological data, define the semantic association relationship of the knowledge entities, and construct an initial hydrological knowledge graph based on the semantic association relationship between knowledge entities. Step S13: Divide the multi-source hydrological data into blocks according to the granularity of each knowledge point, and optimize the block division results by combining them with a preset hydrological professional dictionary to obtain a set of hydrological knowledge blocks; the calculation method for the granularity of knowledge points includes: Text segmentation and paragraph splitting were performed on multi-source hydrological data to obtain several initial semantic units; Based on the hydrological professional dictionary, the number of hydrological terms and the semantic dependency relationship between hydrological terms in each initial semantic unit are detected. If the dependency path coverage rate of the semantic dependency relationship between hydrological terms in the semantic unit reaches the preset coverage rate threshold, the semantic unit is selected as a candidate knowledge point. Calculate the similarity value representation between the candidate knowledge point and its adjacent initial semantic units. If the similarity is lower than the preset similarity threshold, then the candidate knowledge point is treated as an independent knowledge point. The number of hydrological terms for independent knowledge points is used as a knowledge point complexity index, and the knowledge point granularity is generated by combining the knowledge point complexity index with the similarity value. Step S14: Construct a hydrological knowledge graph using the initial hydrological knowledge graph as the index structure and the set of hydrological knowledge blocks as the content carrier to be indexed. Step S2: Generate confused samples corresponding to the hydrological knowledge graph using the pre-set large model, and perform adversarial training on the confused samples; obtain bidirectional mapping questions based on the adversarial training results and historical questions in multi-source hydrological data; Step S3: Input the bidirectional mapping test questions into the question generation interface of the hydrological survey learning platform; collect the answer data and student physiological data corresponding to the bidirectional mapping test questions in real time, and update the student's knowledge status based on the answer data and student physiological data; Step S4: Construct a hydrological survey knowledge model based on the updated knowledge status of trainees; use the hydrological survey knowledge model to generate recommended learning plans for trainees.

2. The method for constructing a hydrological survey knowledge model based on a large model according to claim 1, characterized in that, Step S4 generates a recommended learning plan for the student, including: The difficulty threshold of knowledge points is defined based on the difficulty level of the two-way mapping test questions; Based on the knowledge point difficulty threshold, and using the hydrological survey knowledge model to generate the updated knowledge status of trainees, the corresponding knowledge point mastery level is determined. The difficulty of the corresponding knowledge point is adjusted based on the accuracy rate of the answer data. If the accuracy rate of three consecutive questions in the answer data is greater than 85%, the student's mastery of the corresponding knowledge point will be upgraded by one level, resulting in an upgraded knowledge point mastery. If the time taken for any question in the answer data is greater than 25 seconds, the student's mastery of the corresponding knowledge point will be downgraded by one level, resulting in a downgraded knowledge point mastery. Based on the mastery of upgraded and downgraded knowledge points, the bidirectional mapping test questions adaptively adjust the learning order and test question ratio of each student in the hydrological survey learning platform to obtain a recommended learning plan for each student.

3. The method for constructing a hydrological survey knowledge model based on a large model according to claim 1, characterized in that, The confused samples generated in step S2 corresponding to the hydrological knowledge graph include: Based on the hydrological knowledge graph, independent knowledge points and their adjacent initial semantic units are selected to construct a semantic context set for independent knowledge points; Using a large model, dependency parsing is performed on the semantic description text corresponding to independent knowledge points in multi-source hydrological data, and core term extraction is performed based on the syntactic analysis results to obtain the core terms of independent knowledge points. Align the semantic context set of independent knowledge points with the core terms of independent knowledge points. If the lexical similarity or semantic similarity of the alignment result exceeds the preset corresponding threshold, then the adjacent initial semantic unit is regarded as a candidate knowledge point for confusion of independent knowledge points. Based on the semantic description text of the confused candidate knowledge points in multi-source hydrological data, multiple types of confused samples are generated.

4. The method for constructing a hydrological survey knowledge model based on a large model according to claim 3, characterized in that, Multiple types of obfuscated samples include: By replacing the core terms in independent knowledge points with the core terms of the confused candidate knowledge points, term substitution-type confusion samples are obtained; Based on the semantic description text of confused candidate knowledge points and independent knowledge points in multi-source hydrological data, the contextual relationship between confused candidate knowledge points and independent knowledge points is exchanged to obtain a reversed confusion sample. By introducing principal component semantic features of confusing candidate knowledge points into the semantic description text of independent knowledge points, semantic interference-type confusing samples are obtained.

5. The method for constructing a hydrological survey knowledge model based on a large model according to claim 1, characterized in that, Step S2 involves obtaining bidirectional mapping questions, including: Construct a confusing semantic feature library based on the adversarial training results; Historical test texts were extracted from multi-source hydrological data, and dependency parsing was performed on the test texts. The distribution of test knowledge points was marked based on the results of the dependency parsing. The semantic framework of the test questions was determined based on the distribution of test knowledge points. Using a confused semantic feature library as a perturbation factor, the semantic framework of the test questions is rewritten at multiple levels to generate bidirectional mapping test questions.

6. The method for constructing a hydrological survey knowledge model based on a large model according to claim 5, characterized in that, Performing multi-layered rewriting of the semantic framework of the test questions includes: Replace or insert confusing semantic terms from the confusion semantic feature library into the terms in the semantic framework of the test questions to construct distinguishing test questions; Adjust the constraints in the question stem of the semantic framework of the test questions to construct logically approximate test questions; Rewrite the upstream or downstream knowledge points in the knowledge point chain within the semantic framework of the test question, and construct reasoning to rewrite the test question; The test questions, logic approximation questions, and reasoning rewriting questions are merged into a confusing question set according to a preset question allocation ratio. The confusing question set is then mapped to the corresponding knowledge point links to obtain two-way mapped questions.

7. The method for constructing a hydrological survey knowledge model based on a large model according to claim 1, characterized in that, Step S3, which updates the student's knowledge status based on the answer data and the student's physiological data, includes: Real-time collection of answer data corresponding to bidirectional mapping test questions, including answer accuracy, time taken, and student eye movement trajectory; Collect physiological data from trainees, including electroencephalogram (EEG) signals. Wave frequency variation data; Using students' physiological data and answer data as observations, and combining them with the students' knowledge status at the previous moment, the students' mastery of each knowledge point node in the hydrological knowledge graph is updated. For knowledge point nodes in the hydrological knowledge graph that have not had corresponding answer data for more than 30 days, the student's mastery of the knowledge point node is reduced according to the preset decay factor; the error pattern distribution of knowledge point nodes is updated by statistically analyzing the question types with low answer accuracy; and the student's knowledge status is updated based on the update status of each knowledge point node in the hydrological knowledge graph.

8. A system for constructing a hydrological survey knowledge model based on a large model, characterized in that, The system for constructing a hydrological survey knowledge model based on a large model as described in claim 1 includes: The data acquisition module is used to acquire multi-source hydrological data, define entities and semantic relationships in the multi-source hydrological data, and thus construct a hydrological knowledge graph. The adversarial training module is used to generate confused samples corresponding to the hydrological knowledge graph using a pre-set large model, and to perform adversarial training on the confused samples; based on the adversarial training results and historical questions in multi-source hydrological data, bidirectional mapping questions are obtained. The knowledge status update module is used to input bidirectional mapping questions into the question generation interface of the hydrological survey learning platform; it collects the answer data and student physiological data corresponding to the bidirectional mapping questions in real time, and updates the student's knowledge status based on the answer data and student physiological data; The learning plan recommendation module is used to construct a hydrological survey knowledge model based on the updated knowledge status of trainees; and to generate recommended learning plans for trainees using the hydrological survey knowledge model.

Citation Information

Patent Citations

  • Water conservancy knowledge service system based on knowledge graph technology

    CN118152525A

  • Knowledge question and answer method and device, electronic equipment and storage medium

    CN118228824A

  • Hydrological forecasting model recommendation method and device based on knowledge graph

    CN118861323A

  • Personalized learning recommendation method based on personalized knowledge graph

    CN119808919A

  • Method and device for text-enhanced knowledge graph joint representation learning

    US20220147836A1

Cited By

  • Hydrological data accurate recommendation method and system

    CN121524445A

  • Hydrological data accurate recommendation method and system

    CN121524445B

  • Method for obtaining training data of answer generation model and training method of answer generation model

    CN122432678A

  • Method for obtaining training data of answer generation model and training method of answer generation model

    CN122432678B