Basin similarity classification method and device

GB2612682BActive Publication Date: 2026-08-26HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2022013018
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-08
Filing Date
2022-09-06
Publication Date
2026-08-26
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The selection of similar basins for hydrological forecasting in ungauged regions lacks a unified quantitative standard, leading to uncertainty and an inability to comprehensively capture hydrological rules, hindering effective hydrological similarity analysis and parameter regionalization.

Method used

A basin similarity classification method using a self-organizing map (SOM)-fuzzy c-mean (FCM)-based fusion algorithm, which involves collecting and analyzing hydrological, meteorological, and underlying surface data to construct characteristic indexes, calculate distance correlation coefficients, and perform clustering to determine similarity indexes, ultimately evaluating basin similarity through a fuzzy membership degree.

Benefits of technology

This method constructs a comprehensive basin similarity index system, enabling objective and efficient division of homogeneous hydrological-meteorological regions and sub-basin classification, providing a theoretical basis for selecting reference basins and transferring parameters in ungauged regions, thus advancing hydrological forecasting and warning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000002_0001
    Figure 00000002_0001
Patent Text Reader

Abstract

A method for classifying the similarity of basins comprises five steps. In step 1, hydrological, meteorological and underlying surface data of a basin is collected. Meteorological and underlying surfa
Need to check novelty before this filing date? Find Prior Art

Description

BASIN SIMILARITY CLASSIFICATION METHOD AND DEVICE TECHNICAL FIELD

[0001] The present disclosure belongs to the technical field of hydrology, and particularly relates to a basin similarity classification method and device, which are mainly applied to analogy basins judgment and hydrological warning and forecasting in ungauged regions. BACKGROUND ART

[0002] There has been a “staggering rise” in the number of extreme weather events over the past few years, driven largely by climatic changes and human activities. In the meanwhile, mountain torrents, debris flows and other serious flood disasters take place more and more frequently. In the regions where hydrological data are scarce, due to the lack of effective hydrological monitoring means, and poor implementation of disaster detection, disasters are relatively frequent, and losses are more serious, which severely constrains the sustainable development of social economy. Aiming at the problem of hydrological forecasting in ungauged regions, International Association of Hydrological Sciences (IAHS) launched PUB International Hydrological Programme in 2002. The purpose of the progamme is to make innovation in hydrological work of ungauged regions by shifting from a traditional method of observed data calibrating to a new method of mechanism analysis. At present, a common method for determining the model parameters of ungauged regions is the regionalization method, which refers to a process of selecting a gauged basin which is similar to a target basin, and using the model parameters calibrated by the gauged basins to deduce the model parameters of the ungauged basins, so as to realize the hydrological forecast of the target basin.

[0003] With the development of remote sensing and GIS technology, the data of climate and underlying surface can be easily obtained. By mining the effective information of the structural characteristics of ungauged basins, and making use of the similarity between climate and basin characteristics, the hydrological similarity is deduced in an approximate way, which makes it possible for runoff prediction in ungauged regions. In recent years, with the rapid development of computer technology, pattern recognition, clustering algorithm and fuzzy mathematics have made great progress from research theory to practical application, which provides technical support for the study of hydrological similarity. Therefore, establishing a comprehensive criterion for the basin meteorology and underlying surface similarity, and building an identification method for homogeneous hydrological regions in high-dimensional heterogeneous characteristic space have a certain theoretical basis and practical significance.

[0004] The selection of similar basins is often dependent on the experience of decision-makers, there is no unified quantitative standard, which leads to strong uncertainty in practical application. Besides, a general similarity index combination has not been formed yet, and a specific similarity index alone can hardly capture the comprehensive hydrological rules. How to select the similarity index and construct the evaluation system of similar basins has become a key point to be broken through in the study of hydrological similarity and parameter regionalization. Against this background, the present disclosure provides a general basin similarity classification method and device. SUMMARY

[0005] The technical problem to be solved by the present disclosure is to provide a basin similarity classification method and device, so as to provide a theoretical basis for similar basin selection and parameter transfer of an ungauged region.

[0006] To achieve the above objective, the present disclosure adopts the following technical solutions: a basin similarity classification method, including the following steps:

[0007] step 1, collecting hydrological, meteorological and underlying surface data of a predetermined basin, and extracting meteorological and underlying surface factors to construct characteristic indexes;

[0008] step 2, selecting a characteristic index by using a distance correlation coefficient between hydrological characteristics of the predetermined basin and the meteorological and underlying surface factors, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes;

[0009] step 3, according to a self-organizing map (SOM)-fuzzy c-mean (FCM)-based fusion algorithm, constructing a clustering ensemble on a grid scale, and conducting hydrological- meteorological regionalization on the predetermined basin;

[0010] step 4, based on a result of hydrological-meteorological regionalization, constructing a clustering ensemble on a sub-basin scale in a homogeneous meteorological region, and classifying sub-basins using the fusion model; and

[0011] step 5, according to a meteorological sub-region of the basin and categories of the sub- basins contained, constructing a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins, and evaluating a similarity between the basins using a maximum-minimum approach method to identify the similarity between the basins.

[0012] step 1 includes the following substeps:

[0013] step 11, collecting the hydrological, meteorological and underlying surface data of the predetermined basin, where the hydrological data include a runoff coefficient and a multi-year average daily runoff of the predetermined basin, the meteorological data include a multi-year average monthly precipitation, a multi-year average monthly potential evaporation and a multi- year average monthly temperature of the predetermined basin, and the underlying surface data include topographic characteristics, vegetation types, soil types and land use types of the predetermined basin;

[0014] step 12, according to the meteorological data of the predetermined basin, calculating following meteorological factors as characteristic indexes: annual average humidity index, annual maximum difference of monthly humidity index, snowfall proportion, annual average temperature, annual maximum difference of monthly temperature and snowfall time proportion: E,(t) [1-24 PO> E(t) P(t) M@)= 0 L,PO)=E,() 20) piy<E 0) E(t)

[0015] ml2 0) at dn=75

[0016] 1, , =max(MI(1,2,...12)) -min(MI(1, 2,...12))

[0017] ,,-ZPTOST) > P@) 1=1

[0018] z 1 emiz == YT

[0019] I, , =max(T(l,2,...12))-min(7(1, 2,...12))

[0020] p - ZDIOST) Ch 1=12 > D()

[0021] =

[0022] where MI(®) denotes a humidity index of a 7th month; In , mr , Ts , T., Tor , and D, denote the annual average humidity index, the annual maximum difference of monthly humidity index, the snowfall proportion, the annual average temperature, the annual maximum difference of monthly air temperature and the snowfall time proportion, respectively; PO ’ E® and T® denote a multi-year average precipitation in a sth month, a multi-year average potential evaporation in a 7th month and a multi-year average temperature value in a 7th month, respectively; D(®) denotes a number of days of a 1th month; ho denotes a temperature threshold, precipitation lower than the temperature is in a form of snow, and the temperature threshold is set to 0°C herein; and

[0023] step 13, according to underlying surface information of the predetermined basin, extracting following underlying surface factors as characteristic indexes, including topographic characteristic factors, soil vegetation characteristic factors and shape characteristic factors, ‘where the topographic characteristic factors include an average elevation, a maximum elevation difference, an area-elevation integral, an area-elevation curve slope, an average terrain index and an average slope of the predetermined basin. where the area-elevation integral represents earth surface mass in the predetermined basin, and the area-elevation curve slope reflects a topographic relief degree of the predetermined basin, both of which are calculated by an area- elevation curve; the average terrain index is an arithmetic mean of topography indexes of unit grids in the predetermined basin, and calculation formulas are as follows: HI={ f(x)dx

[0024] 2) = f(08) 0. £( ) A= [00251 0.8-0.2 Ln Tima Soin)

[0026]

[0027] where HI, 45 and TI denote the area-clevation integral value, area-elevation curve slope and average terrain index, respectively; the area-elevation curve J) js a fitted curve composed of *=a / 4 and ¥=h / H and a denotes an area above a certain contour line in the predetermined basin; # denotes a height difference between the contour line and a lowest point in the predetermined basin; 4 denotes a total area of the predetermined basin; # denotes a maximum relative height difference in the predetermined basin; £02) and S08) denote relative height differences on the area-clevation curve corresponding to area ratios are 0.2 and 0.8, respectively; 4 denotes a catchment area of an ith unit grid in the predetermined basin; A denotes a slope of the ith unit grid; and ” denotes a total number of grids in the predetermined basin: and

[0028] the soil vegetation characteristic factors include a sand content, a silt content, a clay content and a normalized vegetation index of a basin; and the shape characteristic factors include a basin area, a basin length, a basin shape factor, a basin-elongation ratio and a basin drainage density.

[0029] step 2 includes the following substeps:

[0030] said selecting a characteristic index from the meteorological and underlying surface factors by using a distance correlation coefficient, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes specifically includes:

[0031] calculating a distance correlation coefficient between different factors: dCov(X,Y) X,Y)= “GLuv \A,l) Cor) = Xaver)

[0032]

[0033] where X denotes a sequence of underlying surface or meteorological characteristic factors, and ¥ denotes a sequence of hydrological characteristic factors; 9CoT(X>Y) denotes a distance correlation coefficient between the sequence X and the sequence ¥ ; 4COV&X,Y) denotes a distance covariance of the sequence X and the sequence ¥ ; and 4¥37(X) and dVar(Y) denote distance standard deviations of the sequence X and the sequence ¥ respectively; and

[0034] setting a threshold @ , and clustering the underlying surface or meteorological characteristic factors with the distance correlation coefficient 9COT (XY) > ® og the similarity indexes. The threshold @ is set to be 0.5, and the following similarity indexes are determined according to the distance correlation coefficient:

[0035] three meteorological indexes, such as annual average precipitation, annual average humidity index and annual maximum difference of monthly humidity index; and seven underlying surface indexes, such as a sand content, a clay content and a normalized vegetation index of a basin, a basin area, a basin length, a basin shape factor, and a basin-elongation ratio.

[0036] step 3 includes the following substeps:

[0037] the hydrological-meteorological clustering ensemble on a grid scale is constructed, training regionalization is carried out by using the SOM-FCM-based fusion model, and the clustering process is divided into two stages: first, carrying out initial clustering by using an SOM algorithm, and obtaining a competitive output layer after training is finished; and second, using a weight vector corresponding to an SOM output layer node as a clustering sample of the FCM algorithm, and performing iterative calculation until a convergence condition is reached.

[0038] The clustering process according to an SOM-FCM-based fusion algorithm in step 3 includes:

[0039] step 31, initializing an SOM neural network: configuring a structure of a competitive layer, an initial neighborhood radius 5(0) , an initial learning rate 20) , a number of iterations k and a total number of iterations k, for the SOM neural network; normalizing an N- dimensional similarity index to obtain a training sample CG ; and setting weight vectors (0) = OBL ares EDU = 120M) oo to neurons in the ANTS NT JIVE LENT LN NTA Emenee 4 corresponding to neurons In the competitive layer, and initializing the weight vectors , where M denotes a number of neurons in the competitive layer, and an initial value of k is set to be 0;

[0040] step 32, inputting a training sample: randomly putting an jth sample _ T G,=(G,,Gy,..., Gy) to an input layer;

[0041] step 33, seeking a winner neuron: calculating a distance between G and W,k) , and selecting a neuron corresponding to a minimum distance as the winner neuron 7 ; 4, =1G,~W,()|= min, (6. -7,%))

[0042]

[0043] step 34, adjusting a connection weight of neurons: adjusting connection weight vectors ¥, of neuron in the neighborhood of the winner neuron 7 ; Wk +1) =W,(k) + 0(E)8; (kX ~W,(K)), J € 5; (k) Wk +1) =W,(k), j 2 67 (k)

[0044]

[0045] where 3k) denotes a learning rate; and sE denotes a neighborhood radius of the ‘winner neuron 7;

[0046] step 35, training an iteration counter & =£+1; and updating 9(k) ang 0k) 5

[0047] step 36, repeating steps 32-35 until the number of iterations ¥ reaches a predetermined total number of iterations k, through training to obtain an SOM output layer neural network;

[0048] step 37, taking a weight ; of each neuron of the SOM output layer neural network as an input vector of the FCM algorithm, setting a number of clusters ¢, a membership degree factor ™ , a limiting error € and a maximum number of iterations ky of the FCM algorithm, initializing the membership degree matrix U, and setting the number of iterations ¥ to 0;

[0049] step 38, calculating a cluster center according to the membership degree matrix U , where an ith cluster center C' is as follows: Surw, C= uy J=1

[0050] i

[0051] where denotes a membership degree of an input vector Ls with respect to G 3

[0052] step 39, according to the cluster center, updating the membership degree matrix U : Uy = - 2 cd mi te

[0053]

[0054] where 9 denotes an Euclidean distance between the input vector 7, and G ; and aq denotes an Euclidean distance between the input vector ¥, and G, g=12,..c ;

[0055] step 310, training the iteration counter ¥ = +1: and ® step 311, repeating steps 39-310 until [4 ue] = 0056] step 311, tin, s39-310 until I~ I= or the number of iterations ¥ P repeating step: reaches a maximum number of iterations k, .

[0057] In step 3, the structure of the SOM network is selected according to a principle of error minimization, evaluation is conducted according to a quantization error QE and a topology error TE, and both indexes can express clustering quality of SOM neural network: 0p-21G 7G) YG

[0058] TE=) v(G,)

[0059]

[0060] where CE denotes an average relative distance between an input sample and a corresponding winner neuron, and W.(G) denotes a weight vector of G corresponding to the winner neuron; and TE denotes a degree of proximity of samples adjacent in an input space in the competitive layer network, and if a sample adjacent to the sample G, remains adjacent to the latter in an output space, VG) isl.

[0061] The number of clusters ¢ of the SOM-FCM algorithm in step 3 is determined according to a Davies-Bouldin index ( DBI ), and € is optimized by evaluating a clustering result under different numbers of clusters: 1 5, +5, DBI ==) max 4 2 is l=)

[0062] v2 5-2 gpl)

[0063]

[0064] where € denotes a number of clusters; M, denotes a number of samples of category i; G denotes a cluster center of category ¢, and C denotes a cluster center of category 9, where i=1,2,,6 4=1.2,04¢ sand 5 denotes an average distance between samples of category 7 and G , and 5 denotes an average distance between samples of category 7 and G,

[0065] step 5, according to a meteorological sub-region of the basin and categories of the sub- basins contained, constructing a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins. Evaluate a similarity between the basins using a meximum-minimum approach method, where similarity between the basin 2 and the basin 22 is expressed as follows: 32060 nin) Simy p, =2— — 2 (i vis)

[0066] Np Np A =3 op = i Up, L A Ug, L AH,

[0067] J= Ju

[0068] where S55 denotes the similarity between the basin 3 and the basin 52; % and nt “5 denote area-weighted membership degrees of the basin 5 and the basin 5 with respect to A A a basin category i, respectively; 2” and *Y denote a ratio of an area of a jth sub-basin to the 3 A a basin category i, respectively; *” and * / denote a ratio of an area of a jth sub-basin to the total basin area in the basin B and the basin B respectively; Ny, and Ny, denote a number of u sub-basins in the basin B and the basin B, , respectively; “a and Yl i and Ys sub-basins in the basin ~! and the basin “2, respectively; "as and % / denote a fuzzy membership degree of the jth sub-basin with respect to the basin category i in the basin 5 and the basin ) , respectively; and A and V denote calculation of minimum and maximum values, respectively.

[0069] The present disclosure further provides a basin similarity classification device, including a processor and a memory, where the memory has a program or instruction stored therein, and the program or instruction is loaded and executed by the processor to implement the steps of the method according to any one of claims 1-9.

[0070] The present disclosure has the following beneficial effects: according to the basin similarity classification method and device, by calculating a distance correlation coefficient between basin factors, a basin similarity index system is constructed; division of homogeneous hydrological-meteorological regions is conducted using an SOM-FCM-based fusion model, and sub-basin classification is achieved using the fusion model; and a comprehensive basin similarity measurement is constructed by a fuzzy membership degree of the sub-basins in combination with a maximum-minimum approach method, thus comprehensively evaluating similarity between basins. Hydrological-meteorological regionalization and sub-basin clustering are conducted using the SOM-FCM-based model, which not only integrates features of self- organization and high nonlinear mapping in the SOM algorithm, but also includes the concept of fuzzy integration, making it possible to explain nonlinear and highly-heterogeneous hydrological data at boundaries that would have been hardly explained. In the meanwhile, an optimum number of clusters is selected according to internal clustering indexes, so as to guarantee efficiency and objective stability of a clustering result. The similarity between basins is expressed through the fuzzy membership degree, and an objective principle of quantification is given for the recognition of similar basins, which provides a theoretical basis for reference basin selection and parameter transfer of the ungauged basin, thus further advancing work in the ungauged region, such as model parameter regionalization and hydrological warning and forecasting is promoted. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] FIG. 1 is a schematic diagram for a calculation flowchart of a basin similarity classification method according to the present disclosure;

[0072] FIG. 2 is a diagram illustrating a distance correlation coefficient between hydrological characteristics and meteorological and underlying surface characteristics according to specific embodiments;

[0073] FIG. 3 is a network structure test diagram of the SOM neural network for hydrological- meteorological regionalization according to specific embodiments;

[0074] FIG. 4 is a diagram showing a relationship between the number of hydrological- meteorological sub-regions and Davies-Bouldin index according to specific embodiments;

[0075] FIG. 5 is a diagram showing a relationship between the number of clusters for sub- basins and Davies-Bouldin index according to specific embodiments; and

[0076] FIG. 6 is a schematic diagram showing the similarity between exemplary basins according to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] The present disclosure is described in further detail below in combination with the accompanying drawings and specific embodiments.

[0078] It should be understood that the specific embodiments described herein are merely intended to explain rather than limit the present disclosure.

[0079] As shown in FIG. 1, the present disclosure provides a basin similarity classification method, including the following steps:

[0080] step 1, collecting hydrological, meteorological and underlying surface data of a predetermined basin, and extracting meteorological and underlying surface factors to construct characteristic indexes;

[0081] step 2, selecting a characteristic index by using a distance correlation coefficient between hydrological characteristics of the predetermined basin and the meteorological and underlying surface factors, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes;

[0082] step 3, according to an SOM-FCM-based fusion algorithm, constructing a clustering ensemble on a grid scale, and conducting hydrological-meteorological regionalization on the predetermined basin;

[0083] step 4, based on a result of hydrological-meteorological regionalization, constructing a clustering ensemble on a sub-basin scale in a homogeneous meteorological region, and classifying sub-basins using the fusion model; and

[0084] step 5, according to a meteorological sub-region of the basin and categories of the sub- basins contained, constructing a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins, and evaluating a similarity between the basins using a maximum-minimum approach method to identify the similarity between the basins.

[0085] Taking Tunxi basin, Fenshuijiang basin, Maduhe basin, Banqgiao basin, Daheba basin, Chenhe basin, Daiying basin, Dage basin, Zhidan basin and Suide basin as 10 typical basins, the classification and similarity of small basins are studied. Specifically:

[0086] step 1, collect hydrological, meteorological and underlying surface data of a predetermined basin, and extract meteorological and underlying surface factors to construct characteristic indexes.

[0087] step 1 specifically includes:

[0088] step 11, collecting the hydrological, meteorological and underlying surface data of the predetermined basin, where the hydrological data include a runoff coefficient and a multi-year average daily runoff of the predetermined basin, the meteorological data include a multi-year average monthly precipitation, a multi-year average monthly potential evaporation and a multi- year average monthly temperature of the predetermined basin, and the underlying surface data include topographic characteristics, vegetation types, soil types and land use types of the predetermined basin;

[0089] step 12, according to the meteorological data of the predetermined basin, extracting following meteorological factors as characteristic indexes: annual average humidity index, annual maximum difference of monthly humidity index, snowfall proportion, annual average temperature, annual maximum difference of monthly temperature and snowfall time proportion: JPO>E (1) 1-£0 P(t) P(t) Mi®)=y 0 LPO=E) ) PQ) |z0 1,P()< E(t)

[0090] 1 mi2 1, =F MIG) by = MU) @

[0091] 4m 10092] 1,,, =max(MI(1,2,...12)) -min(MI(1,2,...12)) ©) , -ZPIOST) 4) > P@)

[0093] =1 1 mi T,=— 310) n=1 £0 3)

[0094] 12 Lm 10095] mr =m2x((1,2,...12))-min(7(1}2,...12)) 6) p,- ZPIOST) Nn SD)

[0096] EJ

[0097] where 7) denotes a humidity index of a 1th month; “nm, mr UR I, Tn , and D, denote the annual average humidity index, the annual maximum difference of monthly humidity index, the snowfall proportion, the annual average temperature, the annual maximum difference of monthly air temperature and the snowfall time proportion, respectively; PO) » E® and T® denote a multi-year average precipitation in a sth month, a multi-year average potential evaporation in a rth month and a multi-year average temperature value in a rth month, respectively; D(®) denotes a number of days of a tth month; % denotes a temperature threshold, precipitation lower than the temperature is in a form of snow, and the temperature threshold is set to 0°C herein; and

[0098] step 13, according to underlying surface information of the predetermined basin, extracting following underlying surface factors as characteristic indexes, including topographic characteristic factors, soil vegetation characteristic factors and shape characteristic factors, where the topographic characteristic factors include an average elevation (Hm), a maximum elevation difference (H;), an area-clevation integral (HI), an area-elevation curve slope (4s), an average terrain index (77) and an average slope ( A ) of the predetermined basin. where the area- elevation integral represents earth surface mass in the predetermined basin, and the area- elevation curve slope reflects a topographic relief degree of the predetermined basin, both of which are calculated by an area-elevation curve; the average terrain index is an arithmetic mean of topography indexes of unit grids in the predetermined basin, and calculation formulas are as follows: oy 2" firma ® 8) - £0. 2) Rn JQ: - A= Tos—02 ©) 8-02 = 08-0 Ag = 00] [01 1 q, TI==Y In(—%— a Cong == i —)

[0101] nig anf 1

[0102] where HI, 4s and TI denote the area-clevation integral value, area-clevation curve slope and average terrain index, respectively; the area-elevation curve S(®) is a fitted curve composed of ¥=a / 4 and y=hlH a4 a denotes an area above a certain contour line in the predetermined basin; h denotes a height difference between the contour line and a lowest point in the predetermined basin; 4 denotes a total area of the predetermined basin; 2 denotes a maximum relative height difference in the predetermined basin; £02) and S08) denote relative height differences on the area-elevation curve corresponding to area ratios are 0.2 and 0.8, respectively; % denotes a catchment area of an ith unit grid in the predetermined basin; A denotes a slope of the ith unit grid; and 7 denotes a total number of grids in the predetermined basin; and

[0103] the soil vegetation characteristic factors include a sand content (Sand), a silt content (Silt), a clay content (Clay) and a normalized vegetation index (NDVI) of a basin; and the shape characteristic factors include a basin area (4), a basin length (L), a basin shape factor (Rf), a basin-elongation ratio (Re) and a basin drainage density (Rd).

[0104] step 2, select a characteristic index by using a distance correlation coefficient between hydrological characteristics of the predetermined basin and the meteorological and underlying surface factors, draw a diagram illustrating a distance correlation coefficient between characteristic factors as shown in FIG. 2, and construct a distance correlation matrix of the basin runoff coefficient (¥ ) and basin characteristics to determine modeling and simulation training indexes.

[0105] Said selecting a characteristic index from the meteorological and underlying surface factors by using a distance correlation coefficient, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes specifically includes:

[0106] calculating a distance correlation coefficient between different factors: dCov(X,Y) A rer —y 11 avLoria, tl) Jadvar(X)avar(t) 0107] dVar(X)dVar(Y)

[0108] where X denotes a sequence of underlying surface or meteorological characteristic factors, and ¥ denotes a sequence of hydrological characteristic factors; 4CoT(X>Y) denotes a distance correlation coefficient between the sequence X and the sequence ¥ ; 4COV&X,Y) denotes a distance covariance of the sequence X and the sequence Y ; and V(X) and dVar(Y) denote distance standard deviations of the sequence X and the sequence Y respectively; and

[0109] setting a threshold @ , and clustering the underlying surface or meteorological characteristic factors with the distance correlation coefficient 4COT(X:Y)>@ os the similarity indexes. The threshold @ is set to be 0.5, and the following similarity indexes are determined according to the distance correlation coefficient:

[0110] three meteorological indexes, such as annual average precipitation, annual average humidity index and annual maximum difference of monthly humidity index; and seven underlying surface indexes, such as a sand content, a clay content and a normalized vegetation index of a basin, a basin area, a basin length, a basin shape factor, and a basin-elongation ratio.

[0111] step 3, according to an SOM-FCM-based fusion algorithm, construct a clustering ensemble on a grid scale, and conduct hydrological-meteorological regionalization on the predetermined basin.

[0112] The hydrological-meteorological clustering ensemble on a grid scale is constructed, training regionalization is carried out by using the SOM-FCM-based fusion model, and the clustering process is divided into two stages: first, carrying out initial clustering by using an SOM algorithm, and obtaining a competitive output layer after training is finished; and second, using a weight vector corresponding to an SOM output layer node as a clustering sample of the FCM algorithm, and performing iterative calculation until a convergence condition is reached.

[0113] The clustering process according to an SOM-FCM-based fusion algorithm includes:

[0114] step 31, initializing an SOM neural network: configuring a structure of a competitive layer, an initial neighborhood radius 50) , an initial learning rate 5(0) , a number of iterations k and a total number of iterations k for the SOM neural network; normalizing an N- dimensional similarity index to obtain a training sample G ; and setting weight vectors (6) = OF, (Ores ONG = 12M) (0 neurons in the SE REE RARE AAR ES corresponding to neurons in the competitive layer, and initializing the weight vectors , where M denotes a number of neurons in the competitive layer, and an initial value of k is set to be 0;

[0115] step 32, inputting a training sample: randomly putting an ith sample - T G,=(G,,G,,..,Gy) to an input layer;

[0116] step 33, seeking a winner neuron: calculating a distance between G, and Wk) , and selecting a neuron corresponding to a minimum distance as the winner neuron 7 ; 4, =|, - #0) = min, (IG, -,®))

[0117] Gir ZIT IARN= aa 7 TE) (12)

[0118] step 34, adjusting a connection weight of neurons: adjusting connection weight vectors ¥, of neuron in the neighborhood, winner neuron 7 ; Wyk +1) =W,(k)+3(k)8 (kX X —W,(k)), ] € 5] (k) W,(k+1)=W,(k), j & 57 (k)

[0119] 1 +) =W,(k),j ¢ 5] (k) “9

[0120] where 3K) denotes a learning rate; and 57 (k) denotes a neighborhood radius of the ‘winner neuron »;

[0121] step 35, training an iteration counter & =k +1; and updating 2%) ang (5);

[0122] step 36, repeating steps 32-35 until the number of iterations k reaches a predetermined total number of iterations k, through training to obtain an SOM output layer neural network;

[0123] step 37, taking a weight ¥; of each neuron of the SOM output layer neural network as an input vector of the FCM algorithm, setting a number of clusters ¢, a membership degree factor ™ , a limiting error € and a maximum number of iterations ky of the FCM algorithm, initializing the membership degree matrix U, and setting the number of iterations ¥ to 0;

[0124] step 38, calculating a cluster center according to the membership degree matrix U , where an ith cluster center ©! is as follows: SMuw, (14) = m Sur

[01258] - id

[0126] where denotes a membership degree of an input vector LJ with respect to G 3

[0127] step 39, according to the cluster center, updating the membership degree matrix U : Uy, = na) 0 d,

[0128] a=1\ 4g w

[0129] where 4 denotes an Euclidean distance between an input vector ", and G,

[0130] step 310, training the iteration counter ¥ =X +1; and i ) _ py ken) step 311, repeating steps 39-310 until [4 v E €

[0131] step 311, repeating steps 39-310 until ==" 17" or the number of iterations * reaches a maximum number of iterations ky .

[0132] The structure of the SOM network in the SOM-FCM algorithm is selected according to a principle of error minimization evaluation is conducted according to a quantization error QE and a topology error TE, and both indexes can express clustering quality of SOM neural network: G,-#,(G; op=2lG l (16) YlG|

[0133] 0134 E= 2vG) an

[0135] where QF Genotes an average relative distance between an input sample and a corresponding winner neuron, and W.(G) denotes a weight vector of G corresponding to the winner neuron; and TE denotes a degree of proximity of samples adjacent in an input space in the competitive layer network, and if a sample adjacent to the sample G remains adjacent to the latter in an output space, v@G) is 1.

[0136] FIG. 3 is a network structure test result of the SOM neural network for hydrological- meteorological regionalization, and finally, the number of model output nodes is 14 x 22.

[0137] The number of clusters ¢ of the SOM-FCM algorithm is determined according to a Davies-Bouldin index ( DBI ), and € is optimized by evaluating a clustering result under different numbers of clusters: WLS <I var feof) bl =—2 max A —————)]——

[0138] elt tn \le-cll; ) (18) v2 Se=l 3p 2 Gl 19

[0139] ag, & 4)

[0140] where © denotes a number of clusters; M, denotes a number of samples of category ; G denotes a cluster center of category , and G denotes a cluster center of category 7, where i=1,2,6 4=1.2,4¢ ; and 5 denotes an average distance between samples of category 7 and G ,and 5 denotes an average distance between samples of category 7 and G,

[0141] k is optimized, when & is 7, DBI reaches the minimum value, and the optimization process of the number of hydrological-meteorological sub-regions k is shown in FIG. 4.

[0142] step 4, based on a result of hydrological-meteorological regionalization, construct a clustering ensemble on a sub-basin scale in a homogeneous meteorological region, and classifying sub-basins using the fusion model.

[0143] Sub-basins are classified using the SOM-FCM algorithm, according to the principle of error minimization, the structure of the SOM network is selected based on a quantization error QE and a topology error TE, the clustering result under different numbers of clusters is evaluated according to DBI, and k is optimized. When k iso, DBI reaches the minimum value, and the optimization process of the number of classified sub-basins ¥ is shown in FIG. 5.

[0144] step 5, according to a meteorological sub-region of the basin and categories of the sub- basins contained, construct a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins, and evaluate a similarity between the basins using a maximum-minimum approach method to identify the similarity between the basins. With Dage basin as the target basin, the similarity among typical basins is shown in FIG. 6.

[0145] Evaluate a similarity between the basins using a maximum-minimum approach method, ‘where the similarity between the basin B and the basin B, is expressed as follows: (20) 3 (at nih) Simy g, =— — > (a vib) =

[0146] uw dy =) A uo @1) Ny Ny Ao oa i Up, “Sau Up, “Su

[0147] = 14

[0148] where S™5.5 denotes the similarity between the basin 2 and the basin 2; % and a “5 denote area-weighted membership degrees of the basin 5 and the basin 5 with respect to a basin category i, respectively; > and A denote a ratio of an area of a jth sub-basin to the a basin category i, respectively; & and as denote a ratio of an area of a jth sub-basin to the total basin area in the basin 5 and the basin 5, respectively; Nan and Na, denote a number of u ul sub-basins in the basin B and the basin B , respectively; 2v and % denote a fuzzy sub-basins in the basin “1 and the basin “2 , respectively; "as and %¢ denote a fuzzy membership degree of the jth sub-basin with respect to the basin category i in the basin 5 and the basin B , respectively; and A and V denote calculation of minimum and maximum values, respectively.

[0149] The embodiments of the present disclosure further provide a basin similarity classification device, including a processor and a memory, where the memory has a program or instruction stored therein, and the program or instruction is loaded and executed by the processor to implement the steps of the method according to the foregoing embodiments.

Claims

WHAT IS CLAIMED IS:

1. A basin similarity classification method, comprising: step 1, collecting hydrological, meteorological and underlying surface data of a predetermined basin, and extracting meteorological and underlying surface factors to construct characteristic indexes; step 2, selecting a characteristic index by using a distance correlation coefficient between hydrological characteristics of the predetermined basin and the meteorological and underlying surface factors, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes; step 3, according to a self-organizing map (SOM)-fuzzy c-mean (FCM)-based fusion algorithm, constructing a clustering ensemble on a grid scale, and conducting hydrological- meteorological regionalization on the predetermined basin; step 4, based on a result of hydrological-meteorological regionalization, constructing a clustering ensemble on a sub-basin scale in a homogeneous meteorological region, and classifying sub-basins using the fusion model; and step 5, according to a meteorological sub-region of the basin and categories of the sub- basins contained, constructing a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins, and evaluating a similarity between the basins using a maximum-minimum approach method to identify the similarity between the basins.

2. The method according to claim 1, wherein step 1 further comprises the following substeps: step 11, collecting the hydrological, meteorological and underlying surface data of the predetermined basin, wherein the hydrological data comprise a runoff coefficient and a multi- year average daily runoff of the predetermined basin, the meteorological data comprise a multi- year average monthly precipitation, a multi-year average monthly potential evaporation and a multi-year average monthly temperature of the predetermined basin, and the underlying surface data comprise topographic characteristics, vegetation types, soil types and land use types of the predetermined basin; step 12, according to the meteorological data of the predetermined basin, calculating following meteorological factors as characteristic indexes: annual average humidity index, annual maximum difference of monthly humidity index, snowfall proportion, annual average temperature, annual maximum difference of monthly temperature and snowfall time proportion; io £0 LUN CNC AY M(= ( 2s JP =Ey(D) Ep AO<ED 1 =z Ly=3 0 MI) 1, , =max(MI(1,2,...12)) -min(MI(1, 2,...12)) Jy > PIO ST) > Pe) fm] 1 ez T= T,, =max(T'(1,2,..12))-min(T(1,2,...12)) p,- ZPTOST) > Dr) =] wherein MI") denotes a humidity index of a fth month; In mr , Js, I. Tor , and D, denote the annual average humidity index, the annual maximum difference of monthly humidity index, the snowfall proportion, the annual average temperature, the annual maximum difference of monthly air temperature and the snowfall time proportion, respectively; PO) # E® and T® denote a multi-year average precipitation in a rth month, a multi-year average potential evaporation in a rth month and a multi-year average temperature value in a fth month, respectively; D(®) denotes a number of days of a th month; . denotes a temperature threshold, precipitation lower than the temperature is in a form of snow, and the temperature threshold is set to 0°C herein; and step 13, according to underlying surface information of the predetermined basin, extracting following underlying surface factors as characteristic indexes, comprising topographic characteristic factors, soil vegetation characteristic factors and shape characteristic factors, wherein the topographic characteristic factors comprise an average elevation, a maximum elevation difference, an area-elevation integral, an area-elevation curve slope, an average terrain index and an average slope of the predetermined basin, wherein the area-elevation integral represents earth surface mass in the predetermined basin, and the area-elevation curve slope reflects a topographic relief degree of the predetermined basin, both of which are calculated by an area-elevation curve; the average terrain index is an arithmetic mean of topography indexes of unit grids in the predetermined basin, and calculation formulas are as follows: HI =| f(x)dx 2)-f ©.8 102-708) A= 08-02 1$ a, TI=—) In(—— ne Cap wherein HI, 4s and TT denote the area-elevation integral, area-elevation curve slope and average terrain index; the area-elevation curve J) is a fitted curve composed of X= al Aand y=hlH ,and @ denotes an area above a certain contour line in the predetermined basin; h denotes a height difference between the contour line and a lowest point in the predetermined basin; 4 denotes a total area of the predetermined basin; H denotes a maximum relative height difference in the predetermined basin; / (®2) and S©8) denote relative height differences on the area-elevation curve corresponding to area ratios are 0.2 and 0.8, respectively; % denotes a catchment area of an ith unit grid in the predetermined basin; A denotes a slope of the ith unit grid; and 7 denotes a total number of grids in the predetermined basin; and the soil vegetation characteristic factors comprise a sand content, a silt content, a clay content and a normalized vegetation index of a basin; and the shape characteristic factors comprise a basin area, a basin length, a basin shape factor, a basin-elongation ratio and a basin drainage density.

3. The method according to claim 1, wherein said selecting a characteristic index from the meteorological and underlying surface factors by using a distance correlation coefficient, and setting a correlation threshold to determine meteorological and underlying surface similarity indexes in step 2 specifically comprises the following substeps: calculating a distance correlation coefficient between different factors: dCov(X,Y) X,7)= “uLuvia,i) ACer(E.T} ~ JaVar(X)dvar(Y) wherein X denotes a sequence of underlying surface or meteorological characteristic factors, and ¥ denotes a sequence of hydrological characteristic factors; 2CO7(XsY) denotes a distance correlation coefficient between the sequence X and the sequence ¥ ; dCoV(X.Y) denotes a distance covariance of the sequence X and the sequence ¥ ; and 974" (X) and aVar(Y) denote distance standard deviations of the sequence X and the sequence Y respectively; and setting a threshold ©? and clustering the underlying surface or meteorological characteristic factors with the distance correlation coefficient 9C07 (XY) > @ gs the similarity indexes.

4. The method according to claim 3, wherein the threshold @ is set to be 0.5, and the following similarity indexes are determined according to the distance correlation coefficient: three meteorological indexes, such as annual average precipitation, annual average humidity index and annual maximum difference of monthly humidity index; and seven underlying surface indexes, such as a sand content, a clay content and a normalized vegetation index of a basin, a basin area, a basin length, a basin shape factor, and a basin-elongation ratio.

5. The method according to claim 1, wherein in step 3, the hydrological-meteorological clustering ensemble on a grid scale is constructed, training regionalization is carried out by using the SOM-FCM-based fusion model, and the clustering process is divided into two stages: first, carrying out initial clustering by using an SOM algorithm, and obtaining a competitive output layer after training is finished; and second, using a weight vector corresponding to an SOM output layer node as a clustering sample of the FCM algorithm, and performing iterative calculation until a convergence condition is reached.

6. The method according to claim 1, wherein the clustering process according to an SOM- FCM-based fusion algorithm in step 3 comprises the following substeps: step 31, initializing an SOM neural network: configuring a structure of a competitive layer, an initial neighborhood radius 4(?), an initial learning rate 2(©), a number of iterations k and a total number of iterations ks for the SOM neural network; normalizing an N-dimensional similarity index to obtain a training sample G; and setting weight vectors (0) = OF 0 aes Wy (OG = L2H) to mewrons in the SINS NE RIVET LE VTA LN VTAAARS mmm mem a ding to neurons In the competitive layer, and initializing the weight vectors , wherein M denotes a number of neurons in the competitive layer, and an initial value of k is set to be 0; = 7 step 32, inputting a training sample: randomly putting an ith sample G,=(Gy,Gy,..., Gy to an input layer; step 33, seeking a winner neuron: calculating a distance between G, and W,(k) , and selecting a neuron corresponding to a minimum distance as the winner neuron r; d,,=|6,-,@®)]= min, (|G, -7,®)) step 34, mapping clustered weight categories of the neurons to a hydrological- meteorological clustering ensemble, and outputting the result of result of hydrological- meteorological regionalization: W,(k +1) =, (k) + 0(k)5; (k)(X ~W, (KD), j 7 (k) W,(k+1)=W,(k), j & 5 (k) wherein 2%) denotes a learning rate; and 5 ® wherein “\* / denotes a learning rate; and «+ ~” denotes a neighborhood radius of the ‘winner neuron 7; step 35, training an iteration counter =k +1; and updating 3k) ang SCF) 7 step 36, repeating steps 32-35 until the number of iterations % reaches a predetermined total number of iterations k through training to obtain an SOM output layer neural network; step 37, taking a weight % of each neuron of the SOM output layer neural network as an input vector of the FCM algorithm, setting a number of clusters €, a membership degree factor m | a limiting error £ and a maximum number of iterations k of the FCM algorithm, initializing the membership degree matrix U, and setting the number of iterations % to 0; step 38, calculating a cluster center according to the membership degree matrix U wheteli at 4h clustel ceiitar’S is as tollows: > uw, BE C =4— Sur J=1 wherein “denotes a membership degree of an input vector , with respect to G 3 step 39, according to the cluster center, updating the membership degree matrix U : u,= . 2 cfd Ym Py wherein 9, denotes an Euclidean distance between the input vector ¥, and G ; and dy denotes an Euclidean distance between the input vector LA and G = L2,..c 4 step 310, training the iteration counter ¥ =% +1; and *) step 311, repeating steps 39-310 until lo AE ‘ step 311, repeating steps 39-310 until © 7 17% or the number of iterations * reaches a maximum number of iterations k, N 7. The method according to claim 6, wherein in step 3, the structure of the SOM network is selected according to a principle of error minimization, evaluation is conducted according to a quantization error QE and a topology error TE, and both indexes can express clustering quality of SOM neural network: 216-76) E=&al trl EZ] TE=)Y v(G) wherein 2E denotes an average relative distance between an input sample and a corresponding winner neuron, and W.(G) denotes a weight vector of G, corresponding to the winner neuron; and TE denotes a degree of proximity of samples adjacent in an input space in the competitive layer network, and if a sample adjacent to the sample G remains adjacent to the latter in an output space, vG) is 1.

8. The method according to claim 7, wherein the number of clusters ¢ of the SOM-FCM algorithm in step 3 is determined according to a Davies-Bouldin index (DBI), and c is optimized by evaluating a clustering result under different numbers of clusters: 1& 5 +5, DBI == a -. = 2m =) wherein ¢ denotes a number of clusters; u, denotes a number of samples of category / ; G denotes a cluster center of category ?, and G denotes a cluster center of category 7, wherein {= 120 €=12,00 ; ang 5. denotes an average distance between samples of category and G , and 5 denotes an average distance between samples of category and C .

9. The method according to claim 1, wherein said according to a meteorological sub- region of the basin and categories of the sub-basins contained, constructing a comprehensive basin similarity measurement by a fuzzy membership degree of the sub-basins in step 5 specifically comprises: evaluating a similarity between the basins using a maximum-minimum approach method, wherein similarity between the basin 5 and the basin 5 is expressed as follows: : i Ad omy 2) 92 Er aah > (a vis) = N; Ng Al =3 TY =3 / Up ’ A, 0, Up, {a Au, = J=1 ; a wherein S™5.2 denotes the similarity between the basin 5 and the basin ©; #3 and a “5 denote area-weighted membership degrees of the basin 5 and the basin 5 with respect to A a basin category i, respectively; 2“ and #Y denote a ratio of an area of a jth sub-basin to the Aes and A total basin area in the basin 5 and the basin 5 , respectively; Nn and Np denote a number 4 4 of sub-basins in the basin B and the basin 5 , respectively; os and on of sub-basins in the basin “1 and the basin “2 , respectively; “as and = denote a fuzzy membership degree of the jth sub-basin with respect to the basin category i in the basin 5 and the basin 5 , respectively; and A and V denote calculation of minimum and maximum values, respectively.

10. A basin similarity classification device, comprising a processor and a memory, wherein the memory has a program or instruction stored therein, and the program or instruction is loaded and executed by the processor to implement the steps of the method according to any one of claims 1-9.