Data-deficient small watershed runoff deduction system and method based on deep learning
Through deep learning algorithms, the geographical data of small watersheds with missing data are encoded semantic embeddings, and high-dimensional semantic fuzzy matching encoding is used for high-dimensional semantic fuzzy matching encoding in known watershed samples with geographical similarity, which solves the problem that it is difficult to obtain runoff data in small watersheds with missing data, and improves the accuracy and applicability of runoff inference.
Patent Information
- Application Number
- CN202510481797.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Data-deficient small watersheds have difficulty obtaining direct and continuous runoff data, resulting in limited accuracy and applicability of hydrological properties assessments, especially in the context of global climate change.
By extracting the runoff data set of known basins from the prior database as reference samples, deep learning algorithms are used to semantically embed and encode the geographical data and reference data samples of target data-lost watershed, and high-dimensional semantic fuzzy matching encoding are used to deduce the runoff situation of target data-lost watershed.
It has improved the accuracy and applicability of the runoff in small watersheds with insufficient data, made full use of global information, and enhanced the ability to assess hydrological characteristics in the context of global climate change.
Smart Images

Figure CN119988463A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of runoff estimation, and more specifically, to a system and method for estimating runoff in a small watershed with insufficient data based on deep learning. Background Art
[0002] In global water resources management and hydrological research, accurate assessment of river basin runoff is of great significance for flood prevention and disaster reduction, water resources planning and environmental protection. However, for many small or remote river basins, it is often difficult to obtain direct and continuous runoff data due to factors such as lack of monitoring facilities, incomplete data records or technical limitations. Accurate assessment of the hydrological characteristics of such areas, known as "data-deficient small watersheds", has become an important challenge in hydrological science research.
[0003] Traditionally, methods to solve the problem of runoff assessment in data-deficient watersheds mostly rely on empirical formulas, regionalization methods, or analogies based on neighboring watersheds. Although these methods can make up for the lack of data to a certain extent, they are often limited by regional specificity, climate change impacts, and the complexity of topography and landforms, resulting in limited accuracy and applicability of prediction results. Especially in the context of global climate change, the assessment of hydrological characteristics of small watersheds with insufficient data faces greater uncertainty.
[0004] In recent years, with the rapid development of geographic information system (GIS) technology, remote sensing technology and big data processing capabilities, it has become possible to use global precipitation data and geospatial information to infer the hydrological characteristics of data-deficient basins. Therefore, a system and method for inferring runoff in small basins with insufficient data based on deep learning is expected. Summary of the invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a system and method for inferring runoff in small watersheds with insufficient data based on deep learning, which extracts the runoff data set of known watersheds from the prior database as reference samples, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample, so as to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample, and then uses the geographic information of the target small watershed with insufficient data as a query, and performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, so as to utilize the runoff data in the known watershed samples with geographical similarity and the global precipitation data to infer the runoff of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data.
[0006] Accordingly, according to one aspect of the present application, a method for estimating runoff in a small watershed with insufficient data based on deep learning is provided, which includes: Obtain geographic data of target small watersheds with insufficient data; Extracting a runoff data set of a known watershed from a priori database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data, and global precipitation data; Vectorizing and encoding each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with missing data respectively to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data; Taking the target data-deficient small watershed geographic data embedded coding vector as the query vector, performing a watershed data high-dimensional semantic fuzzy query based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedded coding vectors to obtain a target data-deficient small watershed semantic query dynamic response coding vector; Based on the dynamic response coding vector of the semantic query of the target small watershed with missing data, the runoff estimated value of the target small watershed with missing data is determined.
[0007] According to another aspect of the present application, a system for estimating runoff in small watersheds with insufficient data based on deep learning is provided, which comprises: A module for acquiring geographic data of small watersheds with missing data, used to acquire geographic data of target small watersheds with missing data; A known basin runoff data extraction module is used to extract a runoff data set of a known basin from a priori database, wherein each data sample in the runoff data set of the known basin includes geographic data, runoff data and global precipitation data; A vectorization encoding module is used to vectorize and encode each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with missing data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data; A semantic fuzzy query encoding module is used to use the target data-deficient small watershed geographic data embedded encoding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedded encoding vectors to obtain a dynamic response encoding vector for semantic query of the target data-deficient small watershed; The runoff inference module is used to determine the runoff inference value of the target small watershed with missing data based on the dynamic response coding vector of the semantic query of the target small watershed with missing data.
[0008] Compared with the prior art, the system and method for inferring runoff in small watersheds with insufficient data based on deep learning provided by the present application extracts the runoff data set of known watersheds from the prior database as reference samples, and uses the deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample, so as to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample, and then uses the geographic information of the target small watershed with insufficient data as a query, and performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, so as to use the runoff data in the known watershed samples with geographical similarity and the global precipitation data to infer the runoff of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 This is a flowchart of a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.
[0011] Figure 2 This is a data flow diagram of a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.
[0012] Figure 3 This is a flowchart of step S4 in the method for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application.
[0013] Figure 4 This is a block diagram of a system for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0015] Figure 1 This is a flowchart of a method for inferring runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 2Schematic diagram of data flow of a method for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 1 and Figure 2 As shown, according to the embodiment of the present application, the method for inferring runoff in a small watershed with insufficient data based on deep learning includes the following steps: S1, acquiring the geographic data of the target small watershed with insufficient data; S2, extracting the runoff data set of the known watershed from the prior database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data and global precipitation data; S3, vectorizing and encoding each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with insufficient data respectively to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with insufficient data; S4, using the embedded coding vector of the geographic data of the target small watershed with insufficient data as the query vector, performing a high-dimensional semantic fuzzy query of the watershed data based on adaptive decision anchoring on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain a semantic query dynamic response coding vector of the target small watershed with insufficient data; S5, determining the runoff inferred value of the target small watershed with insufficient data based on the semantic query dynamic response coding vector of the target small watershed with insufficient data.
[0016] In the above-mentioned method for inferring runoff in small watersheds with insufficient data based on deep learning, the step S1 obtains the geographic data of the target small watershed with insufficient data. It should be understood that the geographical characteristics of the target small watershed with insufficient data are one of the key factors affecting the formation of its runoff. Different geographical locations, topography and water system distribution, etc., will lead to significant differences in runoff generation and confluence processes. This application collects geographical data of the target small watershed with insufficient data, such as latitude and longitude coordinates, altitude, water area, soil type and vegetation coverage, to fully grasp its geographical characteristics, thereby facilitating a more accurate prediction of the runoff situation of the target watershed based on geographical similarities.
[0017] Specifically, in the process of inferring the runoff of small watersheds with insufficient data, the geographic data of the target small watershed with insufficient data can be obtained in a variety of ways. First of all, satellite remote sensing technology is one of the important ways to obtain the geographic data of the target small watershed with insufficient data. Satellites are equipped with various sensors, which can observe the surface over a large area from high altitude and obtain rich geographic information. For example, optical remote sensing satellites can collect surface reflectance spectrum information of the target small watershed through sensors of different bands. Using this information, through professional image processing software and algorithms, it is possible to classify land objects and identify vegetation coverage areas, water body distribution, land use types, etc. in the small watershed. For example, green vegetation has a high reflectivity in the near-infrared band. By analyzing the data of this band, the vegetation coverage map can be accurately drawn to understand the distribution range and density of vegetation. At the same time, radar remote sensing satellites are not restricted by weather and lighting conditions and can penetrate clouds to obtain surface information. It can measure information such as surface roughness and terrain undulation, providing data support for the construction of high-precision digital elevation models (DEM). The DEM generated by radar remote sensing data can clearly display the topographic features of the small watershed, including the direction of mountains, valleys, rivers and the changes in terrain height.
[0018] Geographic Information System (GIS) technology also plays an indispensable role in the acquisition of geographic data of target small watersheds with insufficient data. GIS is a computer system specially used for collecting, storing, managing, analyzing and displaying geographic spatial data. On the one hand, it can integrate geographic data from different data sources, and uniformly manage and process satellite remote sensing images, terrain data, soil data, etc. Through the spatial analysis function of GIS, these data can be overlaid and analyzed, so as to obtain more valuable geographic information. For example, by overlaying and analyzing the land use type map and the river distribution layer, different land use types around the river can be determined; using buffer analysis, the terrain changes within different distances in the small watershed can be determined. On the other hand, GIS can also be combined with the Global Positioning System (GPS). GPS can accurately measure the three-dimensional coordinates of ground points. By setting multiple GPS measurement points in the target small watershed, obtaining the coordinate information of these points, and importing them into the GIS system, the boundaries of the small watershed can be accurately defined, and it can also be used to correct and verify geographic data from other sources.
[0019] At the same time, field measurement is a direct and accurate method to obtain geographic data of small watersheds with insufficient data. Professional surveyors carry various measuring instruments to conduct field observations in small watersheds. For topographic measurement, total station is one of the commonly used instruments. Total station can measure angles and distances. By setting up measuring stations at different locations, the surrounding terrain feature points are measured to obtain the three-dimensional coordinate information of these points. By sorting and processing a large amount of coordinate data of terrain feature points, it is possible to accurately draw a topographic map of the small watershed and show the ups and downs of the terrain in detail. When measuring soil data, soil samples at different depths and locations are collected and brought back to the laboratory for analysis to determine indicators such as soil texture, pH, porosity, and nutrient content, so as to understand the distribution of physical and chemical properties of soil in the small watershed. For the acquisition of vegetation data, in addition to macroscopic monitoring using satellite remote sensing, field investigation is also essential. By counting on the spot and measuring parameters such as the height and diameter at breast height of vegetation, the types, quantity, and growth conditions of vegetation can be more accurately understood. For example, for forest vegetation, measuring the diameter at breast height and height of trees can calculate the forest volume, which is of great significance for assessing the impact of vegetation on runoff.
[0020] In addition, geographic data of target small watersheds with insufficient data can also be obtained through data sharing with relevant departments and institutions. Water conservancy departments usually have hydrological monitoring data in small watersheds, including river flow, water level and other information. Although these data are not direct geographic data, they are closely related to geographic data and are of great value for understanding the hydrological process and runoff formation mechanism of small watersheds. Meteorological departments can provide meteorological data in small watersheds and their surrounding areas, such as precipitation, temperature, wind speed, etc. These meteorological factors are closely related to the generation and change of runoff. Natural resources departments may have more detailed land use status maps, soil type distribution maps and other geographic data resources. Through data sharing and cooperation with these departments, more comprehensive and accurate geographic data can be obtained, providing richer information support for the runoff derivation of target small watersheds with insufficient data.
[0021] In the process of obtaining geographic data of the target small watershed with insufficient data, data quality control is crucial. For satellite remote sensing data, it is necessary to perform preprocessing operations such as radiation correction and geometric correction to eliminate data deviations caused by sensor errors, atmospheric influences and other factors, and improve data accuracy. For field measurement data, operations must be carried out in strict accordance with measurement specifications, and measurement instruments must be calibrated and maintained regularly to ensure the reliability of measurement data. At the same time, in the process of data integration, data from different data sources must be checked for consistency and quality assessed, and erroneous data and outliers must be eliminated to ensure that the final geographic data obtained can truly and accurately reflect the geographical characteristics of the target small watershed with insufficient data.
[0022] In the above-mentioned method for inferring runoff in small watersheds with insufficient data based on deep learning, the step S2 extracts the runoff data set of the known watershed from the prior database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data and global precipitation data. It should be understood that geographic data determines the underlying surface conditions such as the topography, soil type, and vegetation coverage of the watershed, which will affect the interception, infiltration, and confluence of precipitation, and thus affect the generation and size of runoff. Global precipitation data is the direct source of runoff generation, and the intensity, frequency, and total amount of precipitation directly determine how much water can be converted into runoff. Since geographically similar watersheds have similarities in runoff formation mechanisms, and the target small watersheds with insufficient data often lack precipitation and runoff observation conditions, it is impossible to directly obtain runoff through hydrological observations or infer runoff by establishing hydrological models. Therefore, this application extracts the runoff data set of known basins from the prior database, uses the known basin runoff data, geographical data and global precipitation data as reference, and indirectly infers the runoff conditions of the target small basin with insufficient data by understanding the runoff characteristics under different geographical conditions and precipitation conditions.
[0023] Specifically, the priori database, as an information center for storing a large amount of known watershed data, is usually constructed by hydrological survey departments, scientific research and design institutions or related enterprises through long-term data collection, collation and accumulation. These data cover rich information of many watersheds in different time spans, including geographic data, runoff data and global precipitation data, etc., providing valuable reference samples for the runoff derivation of target small watersheds with insufficient data.
[0024] In practical applications, before extracting data, you need to have a deep understanding of the structure and storage method of the prior database. The database may use a relational database management system (such as MySQL, Oracle) or a non-relational database management system (such as MongoDB) for data storage. Relational databases organize data in a table form and perform data operations through structured query language (SQL); non-relational databases focus more on data flexibility and scalability, and use specific query syntax and data models. According to the type and structure of the database, formulate a corresponding data extraction strategy.
[0025] During the data extraction process, the integrity and consistency of the data must also be considered. Since the data in the priori database may come from multiple different data sources, there are differences in data format and quality. Therefore, after extracting the data, data cleaning and preprocessing are required. Check whether the data has missing values, outliers, or duplicate records. For missing values, they can be filled according to the characteristics and statistical laws of the data, such as using the mean, median, or interpolation method for processing; for outliers, the cause of their generation should be analyzed to determine whether they are erroneous data. If so, they should be corrected or eliminated; for duplicate records, deduplication operations should be performed to ensure that the extracted data is accurate.
[0026] In addition, with the continuous growth of data volume and the need for data update, the data extraction process should have a certain degree of automation and scalability. You can write scripts or use data processing tools (such as ETL tools, Extract-Transform-Load, that is, data extraction, transformation and loading tools) to achieve regular data extraction and update. By configuring the corresponding parameters and task scheduling rules, the data extraction process can be automatically executed at predetermined time intervals, and the latest known basin data can be obtained in a timely manner to ensure the timeliness of the data, providing more real-time and reliable data support for the runoff derivation of the target small basin with insufficient data.
[0027] In the above-mentioned method for inferring runoff in a small watershed with insufficient data based on deep learning, the step S3 is to vectorize and encode each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with insufficient data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the target small watershed with insufficient data geographic data. It should be understood that the specific and diverse data structures of the watershed geographic data, runoff data and global precipitation data are difficult to be directly used for model reasoning. Therefore, in order to achieve efficient data processing and at the same time mine the deep semantic information in the data, the present application adopts a semantic embedding coding technology based on deep learning to vectorize and encode each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with insufficient data, respectively, to represent the semantic information of each data sample and the geographic data of the target small watershed with insufficient data in a unified and quantifiable manner, thereby obtaining a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with insufficient data. In the embodiment of the present application, the BERT model is used as the core algorithm of semantic embedding coding, and the contextual semantic dependencies in the data are effectively captured through a multi-layer Transformer structure, and an embedded coding vector representation that can accurately reflect the characteristics of the watershed is generated. In this way, it is helpful to perform efficient fuzzy matching based on the similarity of vectors in a high-dimensional semantic space based on the characteristics of semantic embedding coding, thereby realizing rapid retrieval and matching of geographically similar watershed samples.
[0028] In the above-mentioned deep learning-based method for inferring runoff in small watersheds with insufficient data, the step S4 uses the target small watershed with insufficient data geographic data embedded coding vector as the query vector, and performs a high-dimensional semantic fuzzy query of the watershed data based on adaptive decision anchoring on the query vector and the set of embedded coding vectors of the known watershed sample data to obtain the target small watershed with insufficient data semantic query dynamic response coding vector. That is, in order to further explore the semantic association between the target small watershed and the known watershed in terms of geographical features, so as to use the runoff data of the known watershed to more accurately infer the runoff conditions of the target small watershed, the present application further uses the target small watershed with missing data geographical data embedded coding vector as the query vector, and performs high-dimensional semantic fuzzy query on the query vector and the set of known watershed sample data embedded coding vectors to adaptively strengthen the aggregation of runoff data features of known watershed samples that are most similar to the geographical features of the target small watershed with missing data, and generate a dynamic response coding vector for semantic queries of the target small watershed with missing data, so as to facilitate the deduction of the runoff conditions of the target small watershed with missing data based on the correlation between the geographical features and precipitation data and runoff data in the queried known watershed sample data.
[0029] Figure 3 FIG. 4 is a flowchart of step S4 in the method for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 3 As shown, it includes: S41, performing deep implicit feature extraction on the query vector and each known watershed sample data embedded coding vector in the set of known watershed sample data embedded coding vectors to obtain a set of target small watershed semantic query feature deep implicit coding vectors and known watershed sample data semantic feature deep implicit coding vectors; S42, performing semantic response decision anchor coding on the target small watershed semantic query feature deep implicit coding vector and each known watershed sample data semantic feature deep implicit coding vector in the set of known watershed sample data semantic feature deep implicit coding vectors to obtain a set of target watershed-known watershed sample data semantic response anchor coding matrices; S43, based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrix, dynamically aggregate coding the set of target watershed-known watershed sample data semantic response anchor coding matrices to obtain the target small watershed semantic query dynamic response coding vector.
[0030] Specifically, the step S41 includes: using a deep implicit feature extraction module based on a fully connected coding network to process the query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors respectively to obtain the target data-deficient small watershed semantic query feature deep implicit coding vector and the set of the known watershed sample data semantic feature deep implicit coding vector, which is expressed by the formula: in, represents the set of embedding encoding vectors of known watershed sample data, , , and They represent the first, second, and third embedding vectors of the known watershed sample data. and The known watershed sample data is embedded in the coding vector, is the number of vectors in the set of embedding encoding vectors of the known watershed sample data, represents the query feature weight matrix, Represents the semantic feature weight matrix of known basin sample data, represents the query feature bias term, represents the semantic feature bias item of known basin sample data, represents the query vector, represents the sigmoid activation function, Indicates the deep implicit encoding vector of the semantic query feature of the target small watershed with missing data, express The corresponding semantic feature deep implicit encoding vector of known watershed sample data.
[0031] Here, in order to enhance the semantic query information of the target small watershed with insufficient information and the semantic expression ability of each known watershed sample data, the present application first uses a fully connected encoding network to perform deep implicit feature extraction on the query vector and the embedded encoding vector of each known watershed sample data, respectively, so as to utilize the powerful nonlinear mapping ability of the fully connected encoding network, learn the global nonlinear interaction relationship within each data, and extract a deeper level of semantic feature representation, thereby obtaining a set of deep implicit coding vectors of semantic query features of the target small watershed with insufficient information and deep implicit coding vectors of semantic features of known watershed sample data.
[0032] Specifically, the step S42 is expressed by the formula: in, represents the transpose of a vector, represents vector multiplication, represents the feature scale scaling factor, express and The semantic response anchor encoding matrix between the target basin and the known basin sample data.
[0033] That is, the semantic query feature deep implicit coding vector of the target small watershed with missing data and the semantic feature deep implicit coding vector of each known watershed sample data are further semantically response anchored coded respectively, and semantic interaction is performed through multiplication operation between vectors to capture the geographical characteristics of the target small watershed with missing data and the potential semantic association patterns between each known watershed sample data, thereby generating a set of target watershed-known watershed sample data semantic response anchor coding matrices.
[0034] Specifically, the step S43 includes: first, calculating the decision anchor adaptive splicing factor of each target basin-known basin sample data semantic response anchor coding matrix in the set of the target basin-known basin sample data semantic response anchor coding matrix to obtain a set of target basin-known basin sample data decision anchor adaptive splicing factors. Here, in order to more accurately measure the semantic similarity between each known basin sample data and the target small basin with insufficient data in terms of geographical features, the present application introduces a dynamic response aggregation coding method based on an attention mechanism to process the set of target basin-known basin sample data semantic response anchor coding matrices, and by calculating the decision anchor adaptive splicing factor of each target basin-known basin sample data semantic response anchor coding matrix, it is possible to adaptively focus on the known basin sample data that is most similar to the geographical features of the target small basin with insufficient data, so as to generate the most responsive feature representation by enhancing the aggregation of similar sample runoff data features.
[0035] In a specific example of the present application, based on the statistical eigenvalues of each target basin-known basin sample data semantic response anchor coding matrix in the set of target basin-known basin sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factors of each target basin-known basin sample data semantic response anchor coding matrix are calculated to obtain the set of target basin-known basin sample data decision anchor adaptive splicing factors, and the statistical eigenvalues include the maximum eigenvalue, the characteristic variance, the characteristic mean and the number of eigenvalues. More specifically, the sum of the characteristic variance and the drift coefficient of the target basin-known basin sample data semantic response anchor coding matrix is used as the numerator, and the square of the difference between the maximum eigenvalue and the characteristic mean of the target basin-known basin sample data semantic response anchor coding matrix is calculated multiplied by the number of its eigenvalues, and the drift coefficient and twice the characteristic variance are added as the denominator to obtain the target basin-known basin sample data decision anchor adaptive splicing factor, which is expressed as: in, Indicates the number of elements in the calculation matrix, represents the difference amplification factor, that is, the number of eigenvalues of the semantic response anchor coding matrix of the target basin-known basin sample data, represents the characteristic variance of the semantic response anchor coding matrix of the target basin-known basin sample data, represents the drift coefficient of the target basin-known basin sample data semantic response anchor coding matrix, represents the feature mean of the semantic response anchor coding matrix of the target basin-known basin sample data, To obtain the maximum value function, express The corresponding target basin-known basin sample data decision anchor adaptive splicing factor.
[0036] Here, the target basin-known basin sample data decision anchor adaptive splicing factor is used to reflect the semantic similarity between the geographical features of each known basin sample data and the target data-deficient small basin, as well as the dominance of the semantic interaction information between the two in the global context, which is the weight basis for the subsequent feature dynamic response aggregation coding.
[0037] In particular, in the process of calculating the adaptive splicing factor of the target basin-known basin sample data decision anchor, the drift coefficient is used to smooth the characteristic distribution fluctuation of the target basin-known basin sample data semantic response anchor coding matrix.
[0038] In a preferred example of the present application, for the target basin-known basin sample data decision anchor adaptive splicing factor, the drift coefficient , the distribution of the eigenvalue set of the target basin-known basin sample data semantic response anchor coding matrix transitions from the mean value being weakly interpretable overall to the maximum value being strongly interpretable locally. This application introduces the drift coefficient The weak to strong interpretable generalization is used to enhance the global dominance of the target basin-known basin sample data semantic response anchor coding matrix. Accordingly, the calculation process of the drift coefficient can be expressed as: in, is the intermediate transition representation value of the feature distribution equilibrium state of the target basin-known basin sample data semantic response anchor coding matrix, is a matrix No. eigenvalues, is a natural constant.
[0039] Here, As an intermediate state transition representation from weak interpretability to strong interpretability, each eigenvalue of the target basin-known basin sample data semantic response anchor encoding matrix , which is used as the importance score of the target basin-known basin sample data semantic response anchor encoding matrix for the global smooth state transition to analyze the intermediate state transition The importance score weights relative to the global state transition are globally controlled to achieve interpretable generalized inference of the weight-dependent adaptive splicing factors of the target basin-known basin sample data decision anchors.
[0040] Next, the set of adaptive splicing factors of the target basin-known basin sample data decision anchors is weighted based on the Softmax function to obtain a set of adaptive splicing weight factors of the target basin-known basin sample data decision anchors; based on the set of adaptive splicing weight factors of the target basin-known basin sample data decision anchors, the set of semantic response anchor coding matrices of the target basin-known basin sample data is weighted fused and feature reshaped to obtain the dynamic response coding vector of the semantic query of the target small basin with insufficient data, which is expressed as follows: in, represents the normalized exponential function, Representation Matrix The target watershed-known watershed sample data decision anchor adaptive splicing weight factor, Represents the target basin-known basin sample data semantic response anchor encoding fusion matrix, represents the characteristic reshape function, Represents the dynamic response encoding vector of the semantic query of the target small watershed with missing data.
[0041] That is, the set of adaptive splicing factors of the target basin-known basin sample data decision anchors is normalized into a probability distribution form with a value range in the interval [0,1] by using the Softmax function, and used as the weight factor of the subsequent feature dynamic response aggregation coding, by amplifying the significant differences between the semantic response anchor coding matrices of each target basin-known basin sample data to enhance the distinguishability and expression ability of the features. Finally, based on the generated weight factor, the set of semantic response anchor coding matrices of the target basin-known basin sample data is weighted fused to strengthen the runoff data features of the known basin sample data that are most similar to the geographical features of the target small basin with missing data, and generate the semantic query dynamic response coding vector of the target small basin with missing data.
[0042] In the above-mentioned method for inferring runoff in small watersheds with insufficient data based on deep learning, the step S5 determines the estimated runoff value of the target small watershed with insufficient data based on the semantic query dynamic response coding vector of the target small watershed with insufficient data. In a specific example of the present application, the step S5 includes: inputting the semantic query dynamic response coding vector of the target small watershed with insufficient data into the runoff estimator based on the RNN model to obtain the estimated runoff value of the target small watershed with insufficient data. It should be understood that the RNN model is a recursive neural network, which can effectively capture potential long-distance correlation dependency patterns in the input data through its internal recurrent connection structure. In the present application, the dynamic response coding vector of the semantic query of the target small watershed with missing data integrates the runoff data characteristics of the known watershed samples with similar geographical characteristics to the target small watershed. After the RNN model receives the dynamic response coding vector of the semantic query of the target small watershed with missing data as input, it can adaptively learn and mine the potential correlation rules between the geographical characteristics and precipitation data and runoff data in the known watershed sample data within a long range by performing step-by-step feature analysis and long-distance dependency mining, thereby realizing the intelligent inference of the runoff of the target small watershed with missing data and outputting the runoff inference value of the target small watershed with missing data.
[0043] In a preferred embodiment, the semantic query dynamic response encoding vector of the target small watershed with missing data is input into the runoff estimator based on the RNN model to obtain the runoff estimated value of the target small watershed with missing data, including: Based on the probability encoder of Sigmoid-softmax hybrid function, multi-scale probability distribution deduction is input to the dynamic response encoding vector of the semantic query of the target small watershed with insufficient data to obtain the multi-scale probability distribution encoding vector of the watershed semantic query, which is expressed as: in, represents the normalized exponential function, represents the S-type activation function, represents the dynamic response encoding vector of the semantic query of the target small watershed with missing data, represents a cascade function, Represents the multi-scale probability distribution encoding vector of the watershed semantic query; Calculate the Hadamard product between the multi-scale probability distribution encoding vector of the watershed semantic query and its transposed vector to obtain the global attribute association coupling matrix of the watershed semantic query; Based on the energy distribution reference value of the global attribute association coupling matrix of the watershed semantic query, the global attribute association coupling matrix of the watershed semantic query is normalized and modulated to obtain the normalized attribute response matrix of the watershed semantic query, which is expressed as: in, Indicates the domain attribute association coupling matrix of the watershed semantic query The value of the position, represents the energy distribution benchmark value of the global attribute association coupling matrix of the watershed semantic query, represents matrix multiplication, represents the transposed vector of the multi-scale probability distribution encoding vector of the watershed semantic query, represents the normalized attribute response matrix of the watershed semantic query; The normalized attribute response matrix of the watershed semantic query is input into the multi-scale gated mask generator to obtain the sparse attribute attention mapping matrix of the watershed semantic query, which is expressed as: in, represents the eigenvalues of each position of the normalized attribute response matrix of the watershed semantic query, represents the predetermined hyperparameters, represents a predetermined threshold, represents the mask function, Represents the sparse attribute attention mapping matrix of the watershed semantic query; The dynamic response encoding vector of the semantic query of the target small watershed with insufficient information is used as the benchmark feature vector, which is mapped to the mask space of the sparse attribute attention mapping matrix of the watershed semantic query to obtain the optimized dynamic response encoding vector of the semantic query of the target small watershed with insufficient information, which is expressed as: in, and represents different weight hyperparameters, The dynamic response encoding vector of semantic query of small watershed with insufficient data representing the target of optimization; The optimized target data-deficient small watershed semantic query dynamic response encoding vector is input into the runoff estimator based on the RNN model to obtain the runoff estimated value of the target data-deficient small watershed.
[0044] Accordingly, in this preferred embodiment, a dynamic noise explicit modeling factor of multi-scale distribution perception dimension is introduced into the dynamic response encoding vector of the semantic query of the target small watershed with insufficient data to strengthen the interactive coupling between different variables in the dynamic response encoding vector of the semantic query of the target small watershed with insufficient data based on attribute association, thereby realizing the adaptive scaling optimization of the source domain feature vector based on probability feedback. In this way, the key intrinsic mode features in the source domain features can be more effectively retained, and the noise components in the feature mapping can be selectively suppressed, and finally the decoding smoothness of the manifold expression of the dynamic response encoding vector of the semantic query of the target small watershed with insufficient data in the sequence decoding process is improved, so as to improve the decoding reasoning accuracy of the runoff inference value of the target small watershed with insufficient data.
[0045] In summary, the method for inferring runoff in small watersheds with insufficient data based on deep learning according to the embodiment of the present application is explained, which extracts the runoff data set of known watersheds from the prior database as reference samples, and uses a deep learning algorithm to perform semantic embedding encoding on the geographic data of the target small watershed with insufficient data and each reference data sample, so as to extract the semantic embedding representation of the geographic data of the target small watershed with insufficient data and each reference data sample, and then uses the geographic information of the target small watershed with insufficient data as a query, and performs high-dimensional semantic fuzzy matching encoding on the sample data of the known watershed, so as to use the runoff data in the known watershed samples with geographical similarity and the global precipitation data to infer the runoff of the target small watershed with insufficient data. In this way, global information can be fully utilized to improve the accuracy and applicability of runoff inference in small watersheds with insufficient data.
[0046] Furthermore, the present application also provides a system for estimating runoff in small watersheds with insufficient data based on deep learning.
[0047] Figure 4 FIG. 1 is a block diagram of a system for estimating runoff in a small watershed with insufficient data based on deep learning according to an embodiment of the present application. Figure 4As shown, according to the embodiment of the present application, the system 100 for inferring runoff in a small watershed with insufficient data based on deep learning includes: a module 110 for acquiring geographic data of a small watershed with insufficient data, for acquiring geographic data of a target small watershed with insufficient data; a module 120 for extracting runoff data of a known watershed, for extracting a runoff data set of a known watershed from a priori database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data and global precipitation data; a vectorization encoding module 130, for respectively performing vectorization encoding on each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with insufficient data. A set of known watershed sample data embedded coding vectors and a target small watershed with missing data embedded coding vector are obtained; a semantic fuzzy query coding module 140 is used to use the target small watershed with missing data embedded coding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedded coding vectors to obtain a target small watershed with missing data semantic query dynamic response coding vector; a runoff inference module 150 is used to determine the runoff inference value of the target small watershed with missing data based on the semantic query dynamic response coding vector of the target small watershed with missing data.
[0048] Here, those skilled in the art can understand that the specific operations of each module in the above-mentioned deep learning-based small watershed runoff estimation system with insufficient data have been referred to above. Figures 1 to 3 It has been introduced in detail in the description of the deep learning-based method for estimating runoff in small watersheds with insufficient data, and therefore, its repeated description will be omitted.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for estimating runoff in small watersheds with insufficient data based on deep learning, characterized in that: include: Obtain geographic data of target small watersheds with insufficient data; Extracting a runoff data set of a known watershed from a priori database, wherein each data sample in the runoff data set of the known watershed includes geographic data, runoff data, and global precipitation data; Vectorizing and encoding each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with missing data respectively to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data; Taking the target data-deficient small watershed geographic data embedded coding vector as the query vector, performing a watershed data high-dimensional semantic fuzzy query based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedded coding vectors to obtain a target data-deficient small watershed semantic query dynamic response coding vector; Based on the dynamic response coding vector of the semantic query of the target small watershed with missing data, the runoff estimated value of the target small watershed with missing data is determined.
2. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 1 is characterized in that: Each data sample in the runoff data set of the known watershed includes geographic data, runoff data and global precipitation data.
3. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 2 is characterized in that: Taking the target data-deficient small watershed geographic data embedded coding vector as the query vector, a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring is performed on the query vector and the set of known watershed sample data embedded coding vectors to obtain a target data-deficient small watershed semantic query dynamic response coding vector, including: Performing deep implicit feature extraction on the query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors to obtain a set of deep implicit coding vectors of semantic query features of target small watersheds with missing data and deep implicit coding vectors of semantic features of known watershed sample data; Performing semantic response decision anchor coding on the semantic query feature depth implicit coding vector of the target small watershed with missing data and each semantic feature depth implicit coding vector of the known watershed sample data in the set of semantic feature depth implicit coding vectors of the known watershed sample data to obtain a set of target watershed-known watershed sample data semantic response anchor coding matrices; Based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the set of target watershed-known watershed sample data semantic response anchor coding matrices is dynamically aggregated and encoded to obtain the dynamic response coding vector of the target small watershed semantic query with insufficient information.
4. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 3 is characterized in that: The query vector and each known watershed sample data embedding coding vector in the set of known watershed sample data embedding coding vectors are respectively subjected to deep implicit feature extraction to obtain a set of deep implicit coding vectors of semantic query features of target small watersheds with insufficient data and deep implicit coding vectors of semantic features of known watershed sample data, including: A deep implicit feature extraction module based on a fully connected coding network is used to process the query vector and each known watershed sample data embedded coding vector in the set of known watershed sample data embedded coding vectors respectively to obtain the deep implicit coding vector of the semantic query feature of the target small watershed with missing information and the set of deep implicit coding vectors of the semantic feature of the known watershed sample data.
5. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 4 is characterized in that: Based on the feature contribution of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the set of target watershed-known watershed sample data semantic response anchor coding matrices is dynamically aggregated and encoded to obtain the target data-deficient small watershed semantic query dynamic response coding vector, including: Calculating the decision anchor adaptive splicing factor of each target basin-known basin sample data semantic response anchor coding matrix in the set of target basin-known basin sample data semantic response anchor coding matrices to obtain a set of target basin-known basin sample data decision anchor adaptive splicing factors; The set of target watershed-known watershed sample data decision anchor adaptive splicing factors is weighted based on the Softmax function to obtain a set of target watershed-known watershed sample data decision anchor adaptive splicing weight factors; Based on the set of adaptive splicing weight factors of the target watershed-known watershed sample data decision anchors, the set of target watershed-known watershed sample data semantic response anchor coding matrices is weighted fused and feature-reshaped to obtain the dynamic response coding vector of the semantic query of the target small watershed with insufficient information.
6. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 5 is characterized in that: Calculating the decision anchor adaptive splicing factor of each target basin-known basin sample data semantic response anchor coding matrix in the set of target basin-known basin sample data semantic response anchor coding matrices to obtain a set of target basin-known basin sample data decision anchor adaptive splicing factors, including: Based on the statistical eigenvalues of each target watershed-known watershed sample data semantic response anchor coding matrix in the set of target watershed-known watershed sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factors of each target watershed-known watershed sample data semantic response anchor coding matrix are calculated to obtain the set of target watershed-known watershed sample data decision anchor adaptive splicing factors, and the statistical eigenvalues include the maximum eigenvalue, eigenvariance, eigenmean and eigenvalue number.
7. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 6 is characterized in that: Based on the statistical eigenvalues of each target basin-known basin sample data semantic response anchor coding matrix in the set of target basin-known basin sample data semantic response anchor coding matrices, the decision anchor adaptive splicing factor of each target basin-known basin sample data semantic response anchor coding matrix is calculated to obtain the set of target basin-known basin sample data decision anchor adaptive splicing factors, including: The sum of the characteristic variance and the drift coefficient of the target basin-known basin sample data semantic response anchor coding matrix is used as the numerator, and the square of the difference between the maximum eigenvalue and the characteristic mean of the target basin-known basin sample data semantic response anchor coding matrix is calculated multiplied by the number of its eigenvalues, and then the drift coefficient and twice the characteristic variance are added as the denominator to obtain the target basin-known basin sample data decision anchor adaptive splicing factor, wherein the drift coefficient is used to smooth the characteristic distribution fluctuation of the target basin-known basin sample data semantic response anchor coding matrix.
8. The method for estimating runoff in a small watershed with insufficient data based on deep learning according to claim 7 is characterized in that: Based on the semantic query dynamic response encoding vector of the target small watershed with missing data, the runoff estimated value of the target small watershed with missing data is determined, including: The dynamic response encoding vector of the semantic query of the target small watershed with missing data is input into the runoff estimator based on the RNN model to obtain the runoff estimated value of the target small watershed with missing data.
9. A deep learning-based small watershed runoff estimation system, characterized in that: include: A module for acquiring geographic data of small watersheds with missing data, used to acquire geographic data of target small watersheds with missing data; A known basin runoff data extraction module is used to extract a runoff data set of a known basin from a priori database, wherein each data sample in the runoff data set of the known basin includes geographic data, runoff data and global precipitation data; A vectorization encoding module is used to vectorize and encode each data sample in the runoff data set of the known watershed and the geographic data of the target small watershed with missing data to obtain a set of embedded coding vectors of the known watershed sample data and an embedded coding vector of the geographic data of the target small watershed with missing data; A semantic fuzzy query encoding module is used to use the target data-deficient small watershed geographic data embedded encoding vector as a query vector, and perform a high-dimensional semantic fuzzy query of watershed data based on adaptive decision anchoring on the query vector and the set of known watershed sample data embedded encoding vectors to obtain a dynamic response encoding vector for semantic query of the target data-deficient small watershed; The runoff inference module is used to determine the runoff inference value of the target small watershed with missing data based on the dynamic response coding vector of the semantic query of the target small watershed with missing data.
Citation Information
Patent Citations
Method for estimating runoff in non-data area based on ensemble kalman filter
CN106971034A
Intelligent voice processing method based on common sense understanding and Internet of Things system
CN113450784A
Small sample learning and LSTM (Long Short Term Memory)-based runoff prediction method for areas lacking data
CN114372631A
Cited By
Intelligent management and control system and method for bucket-wheel stacker-reclaimer based on digital twinning
CN120117427A
Material taking flow control system and method of bucket-wheel stacker-reclaimer
CN120122413A