A runoff derivation system and method for small watersheds with scarce data based on deep learning

Through deep learning algorithms, semantic embedding encoding and high-dimensional semantic fuzzy matching of geographical data in small watersheds is solved, and the uncertainty problem of runoff evaluation in small watersheds is achieved, achieving higher accuracy and applicability.

CN119988463BActive Publication Date: 2025-08-05NANJING HYDRAULIC RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510481797.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-05
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Runoff assessment in small watersheds faces uncertainty in the context of global climate change, the accuracy and applicability of traditional methods are limited, making it difficult to obtain direct and continuous runoff data.

Method used

By extracting the runoff data set of known basins from the prior database, deep learning algorithms are used to encode the geographic data and each reference data samples of the target data-lost watershed by semantic embedding, and high-dimensional semantic fuzzy matching encoding, so as to use the runoff data and global precipitation data in the known basin samples with geographical similarity to deduce the runoff situation of the target data-lost watershed.

Benefits of technology

The accuracy and applicability of runoff in small watersheds with short data were improved, and global information was made full use of the results to achieve more accurate runoff prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988463B_ABST
    Figure CN119988463B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of runoff inference, and specifically discloses a deep learning-based system and method for inferring runoff in small watersheds with insufficient data. The system extracts runoff datasets of known watersheds from a priori databases as reference samples, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample, thereby extracting semantic embedding representations of the geographic data of the target small watershed with insufficient data and each reference data sample. Furthermore, using the geographic information of the target small watershed with insufficient data as a query, the system performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, thereby utilizing runoff data from known watershed samples with geographical similarity and global precipitation data to infer the runoff of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of runoff inference, and more specifically, to a system and method for inferring runoff in a small watershed with insufficient data based on deep learning. Background Art

[0002] Accurately assessing river basin runoff is crucial for global water resources management and hydrological research, including flood prevention and mitigation, water resources planning, and environmental protection. However, obtaining direct and continuous runoff data is often difficult for many small or remote river basins due to a lack of monitoring facilities, incomplete data records, or technical limitations. Accurately assessing the hydrological characteristics of these "data-deficient small watersheds" has become a major challenge in hydrological research.

[0003] Traditionally, approaches to assessing runoff in data-deficient watersheds have relied on empirical formulas, regionalized approaches, or analogies based on neighboring watersheds. While these methods can compensate for data deficiencies to a certain extent, they are often limited by regional specificity, the impact of climate change, and the complexity of topography, limiting the accuracy and applicability of predictions. In particular, in the context of global climate change, assessing the hydrological characteristics of small watersheds with limited data faces even greater uncertainty.

[0004] In recent years, with the rapid development of geographic information systems (GIS), remote sensing technology, and big data processing capabilities, it has become possible to use global precipitation data and geospatial information to infer the hydrological characteristics of data-deficient watersheds. Therefore, a deep learning-based system and method for inferring runoff in small watersheds with limited data is desired. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a system and method for inferring runoff in small watersheds with insufficient data based on deep learning, which extracts the runoff data set of known watersheds from a priori database as a reference sample, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample, so as to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample, and then uses the geographic information of the target small watershed with insufficient data as a query, and performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, so as to utilize the runoff data in the known watershed samples with geographical similarity and global precipitation data to infer the runoff situation of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data.

[0006] Accordingly, according to one aspect of the present application, a method for estimating runoff in a small watershed with insufficient data based on deep learning is provided, which includes:

[0007] Obtain geographic data of target small watersheds with insufficient data;

[0008] Extracting a runoff dataset of a known watershed from a priori database, wherein each data sample in the runoff dataset of the known watershed includes geographic data, runoff data, and global precipitation data;

[0009] Performing vector coding on each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with missing data, respectively, to obtain a set of embedding coding vectors of the known watershed sample data and an embedding coding vector of the geographic data of the target small watershed with missing data;

[0010] Using the target data-deficient small watershed geographic data embedded coding vector as a query vector, performing a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of the known watershed sample data embedded coding vectors to obtain a dynamic response coding vector for the semantic query of the target data-deficient small watershed;

[0011] Based on the dynamic response coding vector of the semantic query of the target small watershed with missing data, the runoff estimated value of the target small watershed with missing data is determined.

[0012] According to another aspect of the present application, a deep learning-based system for estimating runoff in small watersheds with insufficient data is provided, comprising:

[0013] The geographic data acquisition module for small watersheds with missing data is used to obtain the geographic data of the target small watersheds with missing data;

[0014] A known watershed runoff data extraction module is used to extract a runoff dataset of a known watershed from a priori database, wherein each data sample in the runoff dataset of the known watershed includes geographic data, runoff data and global precipitation data;

[0015] A vectorization encoding module is used to vectorize and encode each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with missing data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data;

[0016] A semantic fuzzy query encoding module is configured to use the target data-deficient small watershed geographic data embedded coding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain a dynamic response coding vector for the semantic query of the target data-deficient small watershed;

[0017] The runoff inference module is used to determine the runoff inference value of the target small watershed with insufficient data based on the semantic query dynamic response coding vector of the target small watershed with insufficient data.

[0018] Compared with the existing technology, the deep learning-based runoff inference system and method for small watersheds with insufficient data provided by this application extracts runoff datasets of known watersheds from a priori databases as reference samples, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample. Then, using the geographic information of the target small watershed with insufficient data as a query, the system performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, and utilizes the runoff data from the known watershed samples with geographical similarity and global precipitation data to infer the runoff situation of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference for small watersheds with insufficient data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 This is a flowchart of a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.

[0021] Figure 2 Schematic diagram of data flow for a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.

[0022] Figure 3 This is a flowchart of step S4 in the method for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.

[0023] Figure 4 This is a block diagram of a deep learning-based system for inferring runoff in small watersheds with insufficient data according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0025] Figure 1This is a flowchart of a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 2 Schematic diagram of data flow for the method of estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 1 and Figure 2 As shown, according to an embodiment of the present application, the runoff inference method for a small watershed with insufficient data based on deep learning includes the following steps: S1, obtaining the geographic data of the target small watershed with insufficient data; S2, extracting the runoff data set of a known watershed from a priori database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data and global precipitation data; S3, vectorizing and encoding each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with insufficient data to obtain a set of known watershed sample data embedded coding vectors and a target small watershed with insufficient data geographic data embedded coding vector; S4, using the target small watershed with insufficient data geographic data embedded coding vector as the query vector, performing a high-dimensional semantic fuzzy query of the query vector and the set of the known watershed sample data embedded coding vector based on adaptive decision anchoring to obtain a target small watershed with insufficient data semantic query dynamic response coding vector; S5, determining the runoff inference value of the target small watershed with insufficient data based on the semantic query dynamic response coding vector of the target small watershed with insufficient data.

[0026] In the above-mentioned method for inferring runoff of a small watershed with insufficient data based on deep learning, step S1 obtains the geographical data of the target small watershed with insufficient data. It should be understood that the geographical characteristics of the target small watershed with insufficient data are one of the key factors affecting the formation of its runoff. Different geographical locations, topography and water system distribution, etc., will lead to significant differences in the runoff generation and confluence process. This application collects geographical data of the target small watershed with insufficient data, such as latitude and longitude coordinates, altitude, water area, soil type and vegetation coverage, etc., to fully grasp its geographical characteristics, thereby facilitating a more accurate prediction of the runoff situation of the target watershed based on geographical similarity.

[0027] Specifically, when estimating runoff in data-deficient small watersheds, geographic data for these target watersheds can be obtained through a variety of methods. First, satellite remote sensing technology is a key method for obtaining geographic data for these data-deficient watersheds. Satellites, equipped with a variety of sensors, can observe the Earth's surface over large areas from high altitudes, acquiring a wealth of geographic information. For example, optical remote sensing satellites can collect surface reflectance spectral information from target watersheds using sensors in different wavelengths. Using this information, specialized image processing software and algorithms can be used to classify land features and identify vegetation cover, water distribution, and land use types within the watershed. For example, green vegetation has a high reflectivity in the near-infrared band. By analyzing data in this band, accurate vegetation cover maps can be constructed to understand the distribution and density of vegetation. Furthermore, radar remote sensing satellites are not restricted by weather and lighting conditions and can penetrate clouds to obtain surface information. They can measure surface roughness, topographic relief, and other information, providing data support for the construction of high-precision digital elevation models (DEMs). The DEM generated by radar remote sensing data can clearly show the topographic features of the small watershed, including the direction of mountains, valleys, rivers and the changes in terrain height.

[0028] Geographic Information System (GIS) technology also plays an indispensable role in acquiring geographic data for target watersheds where data is scarce. GIS is a computer system specifically designed for collecting, storing, managing, analyzing, and displaying geographic spatial data. It can integrate geographic data from various sources, centrally managing and processing satellite remote sensing imagery, terrain data, soil data, and other data. GIS's spatial analysis capabilities allow for operations such as overlay and buffer analysis to extract more valuable geographic information. For example, by overlaying a land-use type map with a river distribution layer, different land-use types around a river can be identified; buffer analysis can be used to determine topographic variations within different distances within a watershed. GIS can also be integrated with the Global Positioning System (GPS). GPS can accurately measure the three-dimensional coordinates of ground points. By setting up multiple GPS measurement points within a target watershed, acquiring these coordinates, and importing them into a GIS system, the boundaries of the watershed can be precisely defined. This can also be used to calibrate and verify geographic data from other sources.

[0029] Field surveying is a direct and accurate method for obtaining geographic data for small watersheds where data is scarce. Professional surveyors, equipped with various measuring instruments, conduct field observations within the watershed. Total stations are a commonly used instrument for topographic surveying. They can measure angles and distances. By setting up survey stations at various locations, they measure surrounding topographic features and obtain their three-dimensional coordinates. By organizing and processing this large amount of coordinate data, it is possible to accurately map the watershed's topography, detailing the topographical variations. When measuring soil data, soil samples are collected at various depths and locations and brought back to the laboratory for analysis. Indicators such as soil texture, pH, porosity, and nutrient content are measured to understand the distribution of physical and chemical properties within the watershed. For vegetation data, in addition to macroscopic monitoring using satellite remote sensing, field surveys are also essential. Field counts and measurements of parameters such as plant height and diameter at breast height provide a more accurate understanding of vegetation species, abundance, and growth status. For example, for forest vegetation, measuring the diameter at breast height and height of trees can calculate the forest volume, which is important for assessing the impact of vegetation on runoff.

[0030] Furthermore, geographic data for target, data-deficient small watersheds can be obtained through data sharing with relevant departments and institutions. Water conservancy departments typically possess hydrological monitoring data within small watersheds, including information on river flows and water levels. While not directly geographic data, these data are closely related to geographic data and are of great value for understanding hydrological processes and runoff formation mechanisms in small watersheds. Meteorological departments can provide meteorological data within and surrounding areas of small watersheds, such as precipitation, temperature, and wind speed. These meteorological factors are closely related to the generation and variability of runoff. Natural resources departments may possess more detailed geographic data resources, such as current land use maps and soil type distribution maps. By sharing data and collaborating with these departments, more comprehensive and accurate geographic data can be obtained, providing richer information support for runoff estimation in target, data-deficient small watersheds.

[0031] When acquiring geographic data for target small watersheds with insufficient data, data quality control is crucial. Satellite remote sensing data requires preprocessing operations such as radiometric and geometric correction to eliminate data biases caused by sensor errors, atmospheric influences, and other factors, thereby improving data accuracy. Field measurement data must be strictly operated in accordance with measurement specifications, and measuring instruments must be regularly calibrated and maintained to ensure data reliability. Furthermore, during data integration, data from different data sources must be checked for consistency and quality assessed to eliminate erroneous data and outliers, ensuring that the resulting geographic data truly and accurately reflects the geographic characteristics of the target small watersheds with insufficient data.

[0032] In the above-mentioned deep learning-based method for inferring runoff in data-deficient small watersheds, step S2 extracts a runoff dataset for a known watershed from a priori database, wherein each data sample in the runoff dataset for the known watershed includes geographic data, runoff data, and global precipitation data. It should be understood that geographic data determines the underlying surface conditions of the watershed, such as topography, soil type, and vegetation cover. These factors affect processes such as precipitation interception, infiltration, and confluence, and thus affect the generation and size of runoff. Global precipitation data is the direct source of runoff generation. The intensity, frequency, and total amount of precipitation directly determine how much water can be converted into runoff. Since geographically similar watersheds have similar runoff formation mechanisms, the target data-deficient small watershed often lacks precipitation and runoff observation conditions, making it impossible to directly obtain runoff through hydrological observations or infer runoff through the establishment of hydrological models. Therefore, this application extracts the runoff dataset of known basins from the prior database, uses the known basin runoff data, geographical data and global precipitation data as reference, and indirectly infers the runoff conditions of the target small basin with insufficient data by understanding the runoff characteristics under different geographical conditions and precipitation conditions.

[0033] Specifically, a priori databases serve as information hubs storing vast amounts of known watershed data. These databases are typically constructed through long-term data collection, organization, and accumulation by hydrological survey departments, research and design institutions, or related enterprises. These databases encompass a wealth of information on numerous watersheds over varying timeframes, including geographic data, runoff data, and global precipitation data. These data provide valuable reference samples for estimating runoff in data-deficient small watersheds.

[0034] In practical applications, before extracting data, a thorough understanding of the database's structure and storage methods is required. Databases may use relational database management systems (such as MySQL and Oracle) or non-relational database management systems (such as MongoDB) for data storage. Relational databases organize data in tables and perform data operations using Structured Query Language (SQL). Non-relational databases prioritize data flexibility and scalability, employing specific query syntax and data models. Develop a data extraction strategy tailored to the database type and structure.

[0035] During the data extraction process, data integrity and consistency must also be considered. Because data in a priori databases may come from multiple different data sources, data formats and quality may vary. Therefore, after data extraction, data cleaning and preprocessing are necessary. Check the data for missing values, outliers, or duplicate records. Missing values can be filled based on the data's characteristics and statistical patterns, such as using the mean, median, or interpolation. Outliers require analysis to determine if they represent erroneous data. If so, correct or remove them. Duplicate records require deduplication to ensure the accuracy of the extracted data.

[0036] Furthermore, with the continuous growth of data volumes and the need for data updates, the data extraction process should be automated and scalable. Scripts can be written or data processing tools (such as ETL tools, or Extract-Transform-Load) can be used to implement regular data extraction and updates. By configuring appropriate parameters and task scheduling rules, the data extraction process can be automatically executed at predetermined intervals, promptly obtaining the latest known watershed data and ensuring data timeliness. This provides more real-time and reliable data support for runoff estimation in target small watersheds with insufficient data.

[0037] In the above-mentioned method for inferring runoff in a small watershed with insufficient data based on deep learning, the step S3 is to vectorize and encode each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with insufficient data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with insufficient data. It should be understood that the specific and diverse data structures of the watershed's geographic data, runoff data and global precipitation data are difficult to be directly used for model reasoning. Therefore, in order to achieve efficient data processing and at the same time mine the deep semantic information in the data, the present application adopts a semantic embedding coding technology based on deep learning to vectorize and encode each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with insufficient data, and to represent the semantic information of each data sample and the geographic data of the target small watershed with insufficient data in a unified and quantifiable manner, thereby obtaining a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with insufficient data. In the embodiments of this application, the BERT model is used as the core algorithm for semantic embedding coding. Through a multi-layer Transformer structure, it effectively captures the contextual semantic dependencies in the data and generates an embedded coding vector representation that accurately reflects the characteristics of the watershed. In this way, based on the characteristics of semantic embedding coding, it helps to perform efficient fuzzy matching based on vector similarity in a high-dimensional semantic space, thereby achieving rapid retrieval and matching of geographically similar watershed samples.

[0038] In the above-mentioned deep learning-based method for inferring runoff in small watersheds with insufficient data, step S4 uses the target small watershed with insufficient data geographic data embedded coding vector as the query vector, and performs a high-dimensional semantic fuzzy query of the watershed data based on adaptive decision anchoring on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain the target small watershed with insufficient data semantic query dynamic response coding vector. That is, in order to further explore the semantic association between the target small watershed and the known watershed in terms of geographical features, and to use the runoff data of the known watershed to more accurately infer the runoff conditions of the target small watershed, the present application further uses the target small watershed with inadequate data geographical data embedded coding vector as the query vector, and performs high-dimensional semantic fuzzy query on the set of query vectors and known watershed sample data embedded coding vectors to adaptively strengthen the aggregation of runoff data features of known watershed samples that are most similar to the geographical features of the target small watershed with inadequate data, and generate a dynamic response coding vector for the semantic query of the target small watershed with inadequate data, so as to facilitate the deduction of the runoff conditions of the target small watershed with inadequate data based on the correlation between the geographical features and precipitation data and runoff data in the queried known watershed sample data.

[0039] Figure 3 This is a flowchart of step S4 in the method for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 3 As shown, it includes: S41, performing deep implicit feature extraction on the query vector and each known watershed sample data embedded coding vector in the set of known watershed sample data embedded coding vectors to obtain a set of target data-deficient small watershed semantic query feature deep implicit coding vector and known watershed sample data semantic feature deep implicit coding vector; S42, performing semantic response decision anchor coding on the target data-deficient small watershed semantic query feature deep implicit coding vector and each known watershed sample data semantic feature deep implicit coding vector in the set of known watershed sample data semantic feature deep implicit coding vector to obtain a set of target watershed-known watershed sample data semantic response anchor coding matrices; S43, based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrix, performing dynamic aggregation coding on the set of target watershed-known watershed sample data semantic response anchor coding matrices to obtain the target data-deficient small watershed semantic query dynamic response coding vector.

[0040] Specifically, step S41 includes: using a deep implicit feature extraction module based on a fully connected coding network to process the query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors to obtain the target data-deficient small watershed semantic query feature deep implicit coding vector and the set of the known watershed sample data semantic feature deep implicit coding vector, which is expressed as follows:

[0041]

[0042]

[0043]

[0044] in, represents the set of embedding encoding vectors of known basin sample data, 、 、 and They represent the first, second, and third embedding vectors of the known watershed sample data. and The known watershed sample data is embedded in the coding vector, is the number of vectors in the set of embedding encoding vectors of the known watershed sample data, represents the query feature weight matrix, Represents the semantic feature weight matrix of known basin sample data, represents the query feature bias term, Represents the semantic feature bias of known basin sample data, represents the query vector, represents the sigmoid activation function, Represents the deep implicit encoding vector of the semantic query feature of the target small watershed with missing data, express The corresponding semantic feature deep implicit encoding vector of known watershed sample data.

[0045] Here, in order to enhance the semantic query information of the target small watershed with insufficient data and the semantic expression ability of each known watershed sample data, this application first uses a fully connected encoding network to perform deep implicit feature extraction on the query vector and the embedded encoding vector of each known watershed sample data, so as to utilize the powerful nonlinear mapping ability of the fully connected encoding network, learn the global nonlinear interaction relationship within each data, and extract deeper semantic feature representations, thereby obtaining a set of deep implicit encoding vectors of semantic query features of the target small watershed with insufficient data and deep implicit encoding vectors of semantic features of known watershed sample data.

[0046] Specifically, the step S42 is expressed as follows:

[0047]

[0048] in, represents the transpose of a vector, represents vector multiplication, represents the feature scale scaling factor, express and The semantic response anchor encoding matrix between the target basin and the known basin sample data.

[0049] That is, the semantic query feature deep implicit coding vector of the target small watershed with inadequate data and the semantic feature deep implicit coding vector of each known watershed sample data are further subjected to semantic response anchor coding respectively, and semantic interaction is performed through multiplication operation between vectors to capture the potential semantic association pattern between the geographical features of the target small watershed with inadequate data and the sample data of each known watershed, thereby generating a set of target watershed-known watershed sample data semantic response anchor coding matrices.

[0050] Specifically, the step S43 includes: first, calculating the decision anchor adaptive splicing factor of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices to obtain a set of target watershed-known watershed sample data decision anchor adaptive splicing factors. Here, in order to more accurately measure the semantic similarity between each known watershed sample data and the target small watershed with insufficient data in terms of geographical features, the present application introduces a dynamic response aggregation coding method based on the attention mechanism to process the set of target watershed-known watershed sample data semantic response anchor coding matrices. By calculating the decision anchor adaptive splicing factor of each target watershed-known watershed sample data semantic response anchor coding matrix, the method adaptively focuses on the known watershed sample data that is most similar to the geographical features of the target small watershed with insufficient data, so as to generate the most responsive feature representation by enhancing the aggregation of similar sample runoff data features.

[0051] In a specific example of the present application, based on the statistical eigenvalues of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factors of each target watershed-known watershed sample data semantic response anchor coding matrix are calculated to obtain the set of target watershed-known watershed sample data decision anchor adaptive splicing factors, wherein the statistical eigenvalues include the maximum eigenvalue, the characteristic variance, the characteristic mean, and the number of eigenvalues. More specifically, the sum of the characteristic variance and the drift coefficient of the target watershed-known watershed sample data semantic response anchor coding matrix is used as the numerator, and the square of the difference between the maximum eigenvalue and the characteristic mean of the target watershed-known watershed sample data semantic response anchor coding matrix is calculated and multiplied by the number of its eigenvalues, and the drift coefficient and twice the characteristic variance are added as the denominator to obtain the target watershed-known watershed sample data decision anchor adaptive splicing factor, which is expressed as follows:

[0052]

[0053]

[0054] in, Indicates the number of elements in the calculation matrix, represents the difference amplification coefficient, that is, the number of eigenvalues of the semantic response anchor coding matrix of the target basin-known basin sample data, represents the characteristic variance of the semantic response anchor coding matrix of the target basin-known basin sample data, represents the drift coefficient of the semantic response anchor coding matrix of the target basin-known basin sample data, represents the feature mean of the semantic response anchor coding matrix of the target basin-known basin sample data, To obtain the maximum value function, express The corresponding target basin-known basin sample data decision anchor adaptive splicing factor.

[0055] Here, the target basin-known basin sample data decision anchor adaptive splicing factor is used to reflect the semantic similarity between the geographical features of each known basin sample data and the target data-deficient small basin, as well as the dominance of the semantic interaction information between the two in the global context. It is the weight basis for the subsequent feature dynamic response aggregation coding.

[0056] In particular, in the process of calculating the adaptive splicing factor of the target basin-known basin sample data decision anchor, the drift coefficient is used to smooth the characteristic distribution fluctuation of the target basin-known basin sample data semantic response anchor coding matrix.

[0057] In a preferred example of the present application, for the target basin-known basin sample data decision anchor adaptive splicing factor, the drift coefficient , for the semantic response of the target basin-known basin sample data anchor encoding matrix, the distribution of the eigenvalue set changes from the weakly interpretable mean value to the strong interpretable maximum value. This application introduces the drift coefficient The weak to strong interpretable generalization is used to enhance the global dominance of the semantic response anchor encoding matrix of the target basin-known basin sample data. Accordingly, the calculation process of the drift coefficient can be expressed as follows:

[0058]

[0059]

[0060] in, is the intermediate transition representation value of the feature distribution equilibrium state of the semantic response anchor coding matrix of the target basin-known basin sample data, is a matrix No. eigenvalues, is a natural constant.

[0061] Here, As an intermediate state transition representation from weak interpretability to strong interpretability, each eigenvalue of the target basin-known basin sample data semantic response anchor encoding matrix , which is used as the importance score of the target basin-known basin sample data semantic response anchor encoding matrix for the global smooth state transition to analyze the intermediate state transition The importance score weights relative to the global state transition are globally controlled to achieve interpretable generalized inference of the weight-dependent adaptive splicing factors of the target basin-known basin sample data decision anchors.

[0062] Next, the set of target basin-known basin sample data decision anchor adaptive splicing factors is weighted based on the Softmax function to obtain a set of target basin-known basin sample data decision anchor adaptive splicing weight factors; based on the set of target basin-known basin sample data decision anchor adaptive splicing weight factors, the set of target basin-known basin sample data semantic response anchor encoding matrices is weighted fused and feature reshaped to obtain the target data-deficient small basin semantic query dynamic response encoding vector, which is expressed as follows:

[0063]

[0064]

[0065] in, represents the normalized exponential function, Representation matrix The target watershed-known watershed sample data decision anchor adaptive splicing weight factor, Represents the target basin-known basin sample data semantic response anchor encoding fusion matrix, represents the characteristic shape reshaping function, Represents the dynamic response encoding vector of the semantic query of the target small watershed with missing data.

[0066] Specifically, the Softmax function is used to normalize the set of adaptive splicing factors of the target basin-known basin sample data decision anchors into a probability distribution with a value range in the interval [0,1]. This serves as a weighting factor for the subsequent dynamic response aggregation coding of features. This enhances the distinguishability and expressiveness of features by amplifying the significant differences between the semantic response anchor coding matrices of the target basin-known basin sample data. Finally, based on the generated weighting factors, a weighted fusion is performed on the set of semantic response anchor coding matrices of the target basin-known basin sample data. This strengthens the integration of runoff data features of the known basin sample data that most closely resembles the geographic characteristics of the target small basin with missing data, generating a dynamic response coding vector for the semantic query of the target small basin with missing data.

[0067] In the above-mentioned deep learning-based method for estimating runoff in small watersheds with insufficient data, step S5 determines the estimated runoff value of the target small watershed with insufficient data based on the semantic query dynamic response encoding vector of the target small watershed with insufficient data. In a specific example of the present application, step S5 includes: inputting the semantic query dynamic response encoding vector of the target small watershed with insufficient data into a runoff estimator based on an RNN model to obtain the estimated runoff value of the target small watershed with insufficient data. It should be understood that the RNN model is a recursive neural network that, through its internal recurrent connection structure, can effectively capture potential long-range correlation dependency patterns in the input data. In the present application, the dynamic response coding vector for semantic query of the target small watershed with missing data integrates the runoff data features of known watershed samples with similar geographical features to the target small watershed. After the RNN model receives the dynamic response coding vector for semantic query of the target small watershed with missing data as input, it can adaptively learn and mine the potential correlation patterns between the geographical features and precipitation data and runoff data in the known watershed sample data within a long range by performing step-by-step feature analysis and long-distance dependency mining, thereby realizing the intelligent inference of the runoff of the target small watershed with missing data and outputting the runoff inference value of the target small watershed with missing data.

[0068] In a preferred embodiment, the dynamic response encoding vector of the semantic query of the target small watershed with missing data is input into a runoff estimator based on an RNN model to obtain the runoff estimated value of the target small watershed with missing data, including:

[0069] Based on the probability encoder of the Sigmoid-softmax hybrid function, the multi-scale probability distribution deduction input is performed on the dynamic response encoding vector of the semantic query of the target small watershed with insufficient data to obtain the multi-scale probability distribution encoding vector of the watershed semantic query, which is expressed as:

[0070]

[0071] in, represents the normalized exponential function, represents the sigmoid activation function, represents the dynamic response encoding vector of the semantic query of the target small watershed with missing data, Represents a cascade function, Represents the multi-scale probability distribution encoding vector of the watershed semantic query;

[0072] Calculating the Hadamard product between the multi-scale probability distribution encoding vector of the watershed semantic query and its transposed vector to obtain a global attribute correlation coupling matrix of the watershed semantic query;

[0073] Based on the energy distribution benchmark value of the global attribute correlation coupling matrix of the watershed semantic query, the global attribute correlation coupling matrix of the watershed semantic query is normalized and modulated to obtain the normalized attribute response matrix of the watershed semantic query, which is expressed as:

[0074]

[0075]

[0076] in, Represents the domain attribute correlation coupling matrix of the watershed semantic query The value of the position, represents the energy distribution benchmark value of the global attribute correlation coupling matrix of the watershed semantic query, represents matrix multiplication, represents the transposed vector of the multi-scale probability distribution encoding vector of the watershed semantic query, represents the normalized attribute response matrix of the watershed semantic query;

[0077] The watershed semantic query normalized attribute response matrix is input into the multi-scale gated mask generator to obtain the watershed semantic query sparse attribute attention mapping matrix, which is expressed as:

[0078]

[0079]

[0080] in, represents the eigenvalues of each position of the normalized attribute response matrix of the watershed semantic query, represents the predetermined hyperparameters, represents a predetermined threshold, represents the mask function, Represents the sparse attribute attention mapping matrix of the watershed semantic query;

[0081] The dynamic response encoding vector of the target data-deficient small watershed semantic query is used as the benchmark feature vector, which is mapped to the mask space of the watershed semantic query sparse attribute attention mapping matrix to obtain the optimized dynamic response encoding vector of the target data-deficient small watershed semantic query, which is expressed as:

[0082]

[0083] in, and represents different weight hyperparameters, Represents the optimized target data-deficient small watershed semantic query dynamic response encoding vector;

[0084] The optimized target data-deficient small watershed semantic query dynamic response encoding vector is input into the runoff estimator based on the RNN model to obtain the runoff estimated value of the target data-deficient small watershed.

[0085] Accordingly, in this preferred embodiment, a dynamic noise explicit modeling factor based on multi-scale distribution perception dimensions is introduced into the dynamic response encoding vector for semantic queries of target watersheds with insufficient data. This strengthens the attribute-related interactive coupling between different variables in the dynamic response encoding vector for semantic queries of target watersheds with insufficient data, thereby enabling adaptive scaling optimization of the source domain feature vector based on probabilistic feedback. This more effectively preserves key intrinsic mode features in the source domain features and selectively suppresses noise components in the feature map. Ultimately, the decoding smoothness of the manifold representation of the dynamic response encoding vector for semantic queries of target watersheds with insufficient data during the sequence decoding process is improved, thereby enhancing the decoding inference accuracy of the runoff inferred value of the target watersheds with insufficient data.

[0086] In summary, according to the embodiment of the present application, a method for inferring runoff in small watersheds with insufficient data based on deep learning is explained. It extracts the runoff dataset of a known watershed from a priori database as a reference sample, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample. Then, using the geographic information of the target small watershed with insufficient data as a query, the sample data of the known watershed is subjected to high-dimensional semantic fuzzy matching encoding to utilize the runoff data in the known watershed samples with geographical similarity and global precipitation data to infer the runoff situation of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data.

[0087] Furthermore, the present application also provides a system for estimating runoff in small watersheds with insufficient data based on deep learning.

[0088] Figure 4 FIG. 1 is a block diagram of a system for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 4 As shown, the runoff inference system 100 for a small watershed with insufficient data based on deep learning according to an embodiment of the present application includes: a geographic data acquisition module 110 for a small watershed with insufficient data, for acquiring geographic data of a target small watershed with insufficient data; a runoff data extraction module 120 for a known watershed, for extracting a runoff dataset of a known watershed from a priori database, wherein each data sample in the runoff dataset of the known watershed includes geographic data, runoff data and global precipitation data; a vector encoding module 130, for vector encoding each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with insufficient data, respectively. A set of known watershed sample data embedding coding vectors and a target data-deficient small watershed geographic data embedding coding vector are obtained; a semantic fuzzy query coding module 140 is used to use the target data-deficient small watershed geographic data embedding coding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedding coding vectors to obtain a target data-deficient small watershed semantic query dynamic response coding vector; a runoff inference module 150 is used to determine the runoff inference value of the target data-deficient small watershed based on the target data-deficient small watershed semantic query dynamic response coding vector.

[0089] Here, those skilled in the art will understand that the specific operations of each module in the above-mentioned deep learning-based small watershed runoff estimation system have been referenced above. Figures 1 to 3 The description of the deep learning-based method for estimating runoff in small watersheds with insufficient data has been introduced in detail, and therefore, its repeated description will be omitted.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for estimating runoff in small watersheds with insufficient data based on deep learning, characterized by: include: Obtain geographic data of target small watersheds with insufficient data; Extracting a runoff dataset of a known watershed from a priori database, wherein each data sample in the runoff dataset of the known watershed includes geographic data, runoff data, and global precipitation data; Performing vector coding on each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with missing data, respectively, to obtain a set of embedding coding vectors of the known watershed sample data and an embedding coding vector of the geographic data of the target small watershed with missing data; Using the target data-deficient small watershed geographic data embedded coding vector as a query vector, performing a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of the known watershed sample data embedded coding vectors to obtain a dynamic response coding vector for the semantic query of the target data-deficient small watershed; Based on the dynamic response coding vector of the semantic query of the target small watershed with missing data, the runoff estimated value of the target small watershed with missing data is determined.

2. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 1 is characterized in that: Each data sample in the runoff dataset of the known watershed includes geographic data, runoff data and global precipitation data.

3. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 2 is characterized in that: Using the target data-deficient small watershed geographic data embedded coding vector as a query vector, a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring is performed on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain a dynamic response coding vector for the semantic query of the target data-deficient small watershed, including: Performing deep latent feature extraction on the query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors to obtain a set of deep latent coding vectors of semantic query features of target small watersheds with missing data and deep latent coding vectors of semantic features of known watershed sample data; Performing semantic response decision anchor coding on the semantic query feature depth implicit coding vector of the target small watershed with missing data and each semantic feature depth implicit coding vector of the known watershed sample data in the set of semantic feature depth implicit coding vectors of the known watershed sample data to obtain a set of target watershed-known watershed sample data semantic response anchor coding matrices; Based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the set of target watershed-known watershed sample data semantic response anchor coding matrices is dynamically aggregated and encoded to obtain the dynamic response coding vector of the target data-deficient small watershed semantic query.

4. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 3 is characterized in that: Performing deep implicit feature extraction on the query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors to obtain a set of deep implicit coding vectors of semantic query features of a target small watershed with insufficient data and deep implicit coding vectors of semantic features of known watershed sample data, including: A deep implicit feature extraction module based on a fully connected coding network is used to process the query vector and each known watershed sample data embedded coding vector in the set of known watershed sample data embedded coding vectors respectively to obtain the set of the deep implicit coding vector of the semantic query feature of the target small watershed with insufficient information and the deep implicit coding vector of the semantic feature of the known watershed sample data.

5. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 4 is characterized in that: Based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the set of target watershed-known watershed sample data semantic response anchor coding matrices is dynamically aggregated and coded to obtain the target data-deficient small watershed semantic query dynamic response coding vector, including: Calculating the decision anchor adaptive splicing factor of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices to obtain a set of target watershed-known watershed sample data decision anchor adaptive splicing factors; Performing a weighting process based on a Softmax function on the set of target watershed-known watershed sample data decision anchor adaptive splicing factors to obtain a set of target watershed-known watershed sample data decision anchor adaptive splicing weight factors; Based on the set of adaptive splicing weight factors of the target watershed-known watershed sample data decision anchor, the set of semantic response anchor coding matrices of the target watershed-known watershed sample data is weighted fused and feature-reshaped to obtain the dynamic response coding vector of the semantic query of the target small watershed with insufficient data.

6. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 5 is characterized in that: Calculating the decision anchor adaptive splicing factor of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices to obtain a set of target watershed-known watershed sample data decision anchor adaptive splicing factors, including: Based on the statistical eigenvalues of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factors of each target watershed-known watershed sample data semantic response anchor coding matrix are calculated to obtain the set of target watershed-known watershed sample data decision anchor adaptive splicing factors, wherein the statistical eigenvalues include the maximum eigenvalue, eigenvariance, eigenmean and number of eigenvalues.

7. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 6 is characterized in that: Based on the statistical eigenvalues of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factor of each target watershed-known watershed sample data semantic response anchor coding matrix is calculated to obtain the set of target watershed-known watershed sample data decision anchor adaptive splicing factors, including: The sum of the characteristic variance and the drift coefficient of the target basin-known basin sample data semantic response anchor coding matrix is used as the numerator, and the square of the difference between the maximum eigenvalue and the characteristic mean of the target basin-known basin sample data semantic response anchor coding matrix is calculated multiplied by the number of its eigenvalues, and then the drift coefficient and twice the characteristic variance are added as the denominator to obtain the target basin-known basin sample data decision anchor adaptive splicing factor, wherein the drift coefficient is used to smooth the characteristic distribution fluctuation of the target basin-known basin sample data semantic response anchor coding matrix.

8. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 7 is characterized in that: Determining a runoff estimation value of the target data-deficient small watershed based on the semantic query dynamic response encoding vector of the target data-deficient small watershed includes: The dynamic response encoding vector of the semantic query of the target small watershed with missing data is input into the runoff estimator based on the RNN model to obtain the runoff estimated value of the target small watershed with missing data.

9. A deep learning-based system for estimating runoff in small watersheds with insufficient data, characterized by: include: The geographic data acquisition module for small watersheds with missing data is used to obtain the geographic data of the target small watersheds with missing data; A known watershed runoff data extraction module is used to extract a runoff dataset of a known watershed from a priori database, wherein each data sample in the runoff dataset of the known watershed includes geographic data, runoff data and global precipitation data; A vectorization encoding module is used to vectorize and encode each data sample in the runoff dataset of the known watershed and the geographic data of the target small watershed with missing data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data; A semantic fuzzy query encoding module is configured to use the target data-deficient small watershed geographic data embedded coding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain a dynamic response coding vector for the semantic query of the target data-deficient small watershed; The runoff inference module is used to determine the runoff inference value of the target small watershed with insufficient data based on the semantic query dynamic response coding vector of the target small watershed with insufficient data.

Citation Information

Patent Citations

  • Method for estimating runoff in non-data area based on ensemble kalman filter

    CN106971034A

  • Small sample learning and LSTM (Long Short Term Memory)-based runoff prediction method for areas lacking data

    CN114372631A