Mountain torrent debris flow monitoring and early warning method based on image sample intelligent generation

By placing cameras in the channels of flash floods and debris flows to collect video data and using AIGC to generate samples, and combining them with deep learning models to establish an early warning system, the problem of insufficient sample data in traditional methods was solved, and efficient and reliable early warning of flash floods and debris flows was achieved.

CN120726577AInactive Publication Date: 2025-09-30INST OF MOUNTAIN HAZARDS & ENVIRONMENT CHINESE ACADEMY OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511171058.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional flash flood and mudslide early warning methods lack sample data, resulting in insufficient accuracy and reliability of the early warning model. In addition, field monitoring is costly and susceptible to environmental interference, with low levels of automation and intelligence, making it difficult to respond quickly to emergencies.

Method used

By placing cameras at key locations in flash flood and debris flow channels to collect video data, using AIGC technology to generate diverse flash flood and debris flow scene samples, combining deep learning models to establish early warning models, performing data preprocessing and amplification, training and testing, and optimizing model parameters to improve early warning capabilities.

Benefits of technology

It effectively solved the problem of insufficient sample data sets, improved the generalization ability and accuracy of the early warning model, achieved rapid identification and reliable early warning of mountain torrents and mudslides, and improved the timeliness and reliability of the early warning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726577A_ABST
    Figure CN120726577A_ABST
Patent Text Reader

Abstract

The invention relates to a mountain torrent debris flow monitoring and early warning method based on image sample intelligent generation, and belongs to the technical field of geological disaster monitoring and early warning, and the method comprises the following steps: video collection and transmission: collecting video data of a mountain torrent debris flow channel in real time; data preprocessing and amplification: preprocessing the collected video data, intelligently generating a preprocessed original photo by using an AIGC technology, and amplifying a sample data set; training and testing an early warning model: training and testing a sample by using a deep learning model based on the amplified sample data set, and establishing an early warning model to realize monitoring and early warning of the mountain torrent debris flow; model evaluation and adjustment: evaluating the model obtained by training, and adjusting model parameters according to an evaluation result to optimize performance; the mountain torrent debris flow early warning method based on AIGC and deep learning has the beneficial effects that diversified mountain torrent debris flow scene samples are generated, and the problem of insufficient sample data sets is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of geological disaster monitoring and early warning, and in particular relates to a flash flood and debris flow monitoring and early warning method based on intelligent generation of image samples. Background Art

[0002] Flash floods and debris flows are extremely destructive natural disasters, characterized by suddenness, widespread occurrence, and difficulty in monitoring. They pose a serious threat to property safety. Due to the vast territory and complex geological environment, flash floods and debris flows are frequent, and their prevention and control have always been a key focus of geological disaster research.

[0003] Traditional flash flood and debris flow early warning methods have numerous limitations. First, they rely on historical event data and field monitoring data. However, flash flood and debris flow events are rare, and many regions lack sufficient sample data, resulting in insufficient accuracy and reliability of warning models. Second, traditional field monitoring requires the deployment of a large number of sensor devices, which is costly. Furthermore, in complex terrain and changing environments, equipment installation, maintenance, and data transmission are challenging, and are susceptible to environmental interference, leading to data loss or delays. Furthermore, traditional methods rely on manual analysis and empirical judgment, with low levels of automation and intelligence, making it difficult to respond quickly to emergencies.

[0004] In recent years, the application of artificial intelligence (AI) technology in flash flood and debris flow early warning has gradually attracted attention. Attempts have been made to combine satellite remote sensing and drone technology with machine learning algorithms for disaster risk assessment. High-precision sensor networks, geographic information systems (GIS), and AI have also been utilized. However, these technologies still generally face the problem of insufficient sample data, making it difficult to train reliable early warning models, especially in remote areas or small-scale flash flood and debris flow channels. Summary of the Invention

[0005] The present invention provides a flash flood and debris flow monitoring and early warning method based on intelligent generation of image samples, which is used to solve the technical problem of lack of sample data in traditional early warning methods. The flash flood and debris flow early warning method based on AIGC and deep learning effectively solves the problem of insufficient sample data set by generating diversified flash flood and debris flow scene samples.

[0006] In order to achieve the above object, the present invention is implemented by the following technical solutions:

[0007] A flash flood and debris flow monitoring and early warning method based on intelligent generation of image samples includes the following steps:

[0008] Step I): Video acquisition and transmission: Using cameras deployed at key locations along the path of the flash flood and debris flow, real-time video data of the flash flood and debris flow path is collected and transmitted to a data processing center via the network.

[0009] Step II): Data preprocessing and augmentation: The collected video data is preprocessed, including image denoising and enhancement to improve image quality. AIGC technology is then used to intelligently generate the preprocessed original photos to augment the sample dataset.

[0010] Step III): Early warning model training and testing: Based on the amplified sample dataset, a deep learning model is used to train and test the samples, establish an early warning model, and implement early warning monitoring of flash floods and debris flows.

[0011] Step IV): Model evaluation and adjustment: Evaluate the trained model and adjust the model parameters based on the evaluation results to optimize performance.

[0012] Optionally, in the video capture and transmission step, the specific steps for deploying cameras are as follows:

[0013] Step a): The cameras are arranged to cover the entire cross section of the ditch, including the ditch bed, ditch walls, and hillsides on both sides of the ditch;

[0014] Step b): increasing the density of cameras in potentially dangerous areas of the trench;

[0015] Step c): The camera is installed on a stable foundation and has lightning protection, waterproof and dustproof measures;

[0016] Step d): The camera collects video and transmits data via wireless or wired signals, and the data transmission delay time of the camera is controlled within 1 second.

[0017] Optionally, in the data preprocessing and amplification step, the data preprocessing includes image denoising and enhancement operations; wherein, image denoising adopts a Gaussian filtering method, and image enhancement adopts a histogram equalization method.

[0018] Optionally, image enhancement uses histogram equalization to improve the contrast and brightness of the image and make the details in the image clearer:

[0019] The steps of the histogram equalization method are as follows:

[0020] Step 1): Calculate the percentage of pixels for each grayscale value and obtain the PDF of the histogram;

[0021] Step 2): Accumulate the PDF of each gray level to get the CDF of the histogram:

[0022] Step 3): Quantize the CDF and map it to the output image.

[0023] Optionally, in the data preprocessing and amplification steps, the specific process of using AIGC technology to amplify the sample data set is: input the preprocessed original photos into a generative AI model based on a generative adversarial network, and through adversarial training of the generator and the discriminator, generate pictures of channel sections flooded by mountain torrents and mudslides in different proportions. The generated pictures and the original photos constitute an amplified data set.

[0024] Optionally, in the early warning model training and testing steps, the deep learning model used is a convolutional neural network to establish an early warning model. The early warning model is constructed as follows:

[0025] The input layer is used to receive pre-processed images with a uniform size of 224×224 pixels;

[0026] The convolution layer includes the first convolution layer and the second convolution layer. The first convolution layer contains 32 3×3 convolution kernels with a step size of 1 and uses the ReLU activation function; the second convolution layer contains 64 3×3 convolution kernels with a step size of 1 and uses the ReLU activation function.

[0027] The pooling layer is a 2×2 maximum pooling layer set after each convolutional layer with a stride of 2;

[0028] The fully connected layer includes the first fully connected layer and the second fully connected layer. The first fully connected layer contains 128 neurons and uses the ReLU activation function; the second fully connected layer contains 64 neurons and uses the ReLU activation function;

[0029] The output layer contains 2 neurons, and the neurons use the softmax activation function to output probability distribution.

[0030] Optionally, the model training of the early warning model uses the cross entropy loss function as the optimization target, selects the Adam optimization algorithm, and sets the parameters as follows: the initial value of the learning rate is 0.001, β1=0.9, β2=0.999, the number of training rounds is 50, and the performance indicators of the model on the validation set are recorded every 5 rounds.

[0031] Optionally, in the model evaluation and adjustment step, the generated sample dataset is evaluated, and the evaluation indicators include fidelity, diversity, and consistency; among them, fidelity is measured by manual annotation or automatic evaluation algorithm to evaluate the visual similarity between the generated image and the real image; diversity is evaluated by IS, and when IS is greater than 6, the generated image meets the diversity requirements; consistency is measured by mean square error, structural similarity, and peak signal-to-noise ratio.

[0032] Optionally, in the model evaluation and adjustment step, accuracy, recall, F1 value, and area under the ROC curve are used as performance evaluation indicators.

[0033] Optionally, in the model evaluation and adjustment step, the model is adjusted according to the performance indicators of the model on the validation set; when the model is overfitting, the strategy of increasing data augmentation diversity, adjusting the model structure or reducing the learning rate is adopted; when the model is underfitting, the strategy of increasing model complexity, adjusting optimization algorithm parameters or further preprocessing the data is adopted.

[0034] Beneficial effects of the present invention:

[0035] 1. This invention uses AIGC (artificial intelligence generated content) and deep learning to develop a flash flood and debris flow early warning method. By generating diverse flash flood and debris flow scene samples, it effectively solves the problem of insufficient sample data sets. This not only improves the generalization ability and accuracy of the early warning model, but also provides a basis for early warning of other similar natural disasters.

[0036] 2. Through AIGC technology, the present invention can automatically generate a large number of flash flood and mudslide images, video sample data and rich model training data sets. Combined with deep learning algorithms, the model can automatically extract and analyze key features in these generated data, thereby improving the ability to identify flash floods and mudslides. The automation and intelligent characteristics of AIGC technology enable the early warning system to quickly respond to sudden flash floods and mudslides, further improving the timeliness and reliability of the early warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 Schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION

[0039] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0040] Example 1;

[0041] like Figure 1 As shown, this embodiment provides a flash flood and debris flow monitoring and early warning method based on intelligent generation of image samples, which is characterized by comprising the following steps:

[0042] Step I): Video acquisition and transmission: Cameras deployed at key locations along the path of flash floods and debris flows collect real-time video data from the path and transmit it to a data processing center via the network. This ensures stable acquisition and efficient transmission of video data, meeting the system's real-time requirements.

[0043] Step II): Data preprocessing and amplification: The collected video data is preprocessed, including image denoising and enhancement to improve image quality. AIGC technology is then used to intelligently generate the preprocessed original photos and expand the sample data set. Among them, the use of AIGC technology to intelligently generate the original photos and expand the sample data set provides rich sample resources for model training. The generative AI model generates high-quality new sample data, effectively solving the problem of sample scarcity in traditional methods.

[0044] Step III): Early warning model training and testing: Based on the expanded sample dataset, a deep learning model is used to train and test the samples, establish an early warning model, and realize the monitoring and early warning of flash floods and mudslides. Specifically, based on the expanded sample dataset, a deep learning model is used to train and test the image samples of each ditch to establish an early warning model. The appropriate deep learning algorithm and model architecture are selected. Through training on a large number of sample datasets, it can accurately identify the signs of flash floods and mudslides and output reliable early warning signals.

[0045] Step IV): Model evaluation and adjustment: Evaluate the trained model and adjust the model parameters based on the evaluation results to optimize performance.

[0046] Example 2;

[0047] Based on Example 1, in the video acquisition and transmission steps, the camera layout and original photo acquisition are:

[0048] On the basis of completing the field survey, the reasonable deployment of cameras is a key step in obtaining high-quality original photos. The deployment of cameras needs to comprehensively consider factors such as the topography, geological conditions and hydrological characteristics of the ditch to ensure that the key information of the mountain torrent and debris flow ditch can be fully and clearly captured.

[0049] The camera deployment operation is to select key locations in the flash flood and debris flow channel based on the field survey results. Specifically, the camera deployment operation steps include:

[0050] Step a): The cameras are deployed to cover the entire cross-section of the ditch, including the ditch bed, ditch walls, and hillsides on both sides of the ditch, ensuring that the entire scene of the flash flood and debris flow is captured;

[0051] Step b): Increase the density of cameras for key monitoring in potentially dangerous areas of the channel, such as potential landslides, narrow channel areas, and water inlets.

[0052] Step c): The camera should be installed on a stable foundation to avoid damage or displacement due to geological disasters or water erosion. At the same time, the camera should have lightning protection, waterproof and dustproof measures to ensure long-term stable operation;

[0053] Step d): The camera collects video and transmits data via wireless or wired signals. The collected video data can be transmitted to the data processing center in real time. The delay time of the camera's data transmission should be controlled within 1 second to meet the system's real-time requirements.

[0054] After the cameras are deployed, take photos of each flash flood and debris flow gully from the front to obtain the original monitoring photos of the gully. The requirements for collecting the original photos are as follows:

[0055] Shooting angle: The camera should shoot perpendicular to the channel direction to ensure that the captured channel cross-section image is clear and complete. The error of the shooting angle should be controlled within ±5° to ensure the accuracy and consistency of the image.

[0056] Shooting frequency: The frequency of original photo collection is determined based on the hydrological characteristics and rainfall conditions of the flash flood and debris flow gully. During the non-rainy season, original photos are collected 1-2 times a day. During the rainy season, the frequency can be increased appropriately, based on rainfall intensity and gully water level fluctuations, to 1-2 times per hour to capture signs of flash flood and debris flow.

[0057] Image quality: The original photos collected should have high resolution and clarity, and be able to clearly show the topography, landforms, water flow conditions and distribution of ditch bed materials. The image resolution should be no less than 1920×1080 pixels, and the signal-to-noise ratio should be greater than 30dB to meet the needs of subsequent data processing and model training.

[0058] Example 3;

[0059] Based on Example 1, data preprocessing and amplification include the following methods:

[0060] Raw data preprocessing involves preprocessing the original photos before they are fed into the generative AI model. This preprocessing step includes image denoising and image enhancement operations to improve image quality.

[0061] The image denoising operation uses the Gaussian filtering method to remove the noise interference in the image:

[0062] (1);

[0063] In formula (1), The coordinates of any point in the calculation area, that is, the position of the pixel, the origin (0,0) is the coordinates of the center point of the calculation area, and each position Corresponding two-dimensional Gaussian function value , find the center position of the filter kernel in the image processing process, and the positions of other pixels in the filter kernel relative to the center position of the filter kernel are , the pixel position The weight of The decision reflects the rule that the farther away from the center, the lower the weight; is the variance, is the standard deviation coefficient, The larger it is, the flatter the curve is, and the stronger the smoothing effect of the corresponding filter is. The smaller it is, the sharper the curve is, the stronger the ability to retain details is, and the standard deviation of the Gaussian function is adjusted. , which can flexibly control the degree of smoothing according to the noise characteristics and image detail requirements.

[0064] For weighted fair normalization, when filtering the image, the surrounding pixels should be fairly aggregated so that the entire Gaussian function covers all pixels. For example, a 3×3 filter kernel with 9 weights summing to 1 will not brighten or darken the image as a whole when using weighted averaging pixels, thus avoiding destroying the original brightness information.

[0065] is a natural constant, the exponential part The faster the growth, the faster the weight decays. During filtering, only pixels within a very small range (e.g., a 3×3 kernel) near the center pixel have a higher weight, while the weight of pixels farther from the center pixel is almost negligible. The filtered image reduces noise, but retains more pixels of image details (e.g., edges and textures), making it appear less blurry.

[0066] The image is convolved using the Gaussian function. The expression of Gaussian filtering is:

[0067] (2);

[0068] In formula (2), The coefficients of the Gaussian kernel, whose size is given by , Decide; Indicates that the filtered image is at position The pixel value of Indicates that the original image is at position The pixel value of the original image is Relative offset The input pixel value at position; a and b represent the radius range of the neighborhood window.

[0069] Indicates that the convolution operation adopts the weighted summation of the discrete domain, traversing 、 For all offsets within the range, the Gaussian kernel coefficients The image pixel at the corresponding position Multiply and add to get .

[0070] Image enhancement uses histogram equalization to improve the contrast and brightness of the image, making the details in the image clearer. The preprocessed image can better meet the input requirements of the generative model and improve the quality of the generated samples.

[0071] The steps of the histogram equalization algorithm are as follows:

[0072] Step 1): Calculate the percentage of pixels for each grayscale value and get the PDF of the histogram:

[0073] (3);

[0074] In formula (3), It is The probability density function value of a value (or interval), is the grayscale of the input image, is the total number of pixels in the input image, Is grayscale The total number of pixels.

[0075] Step 2): Accumulate the PDF of each gray level to get the CDF of the histogram:

[0076] (4);

[0077] in, Indicates the data The cumulative distribution function value corresponding to a value (or interval), It is The probability density function value of the value (or interval) is summed from arrive , for PDF from the beginning to the The cumulative probability result is obtained by adding up the intervals.

[0078] Step 3): Quantize the CDF and map it to the output image middle:

[0079] (5);

[0080] In formula (5), start and end represent the minimum grayscale and maximum grayscale of the mapping interval, respectively.

[0081] The original photos were fed into a generative AI model. Based on a generative adversarial network (GAN), a generator and discriminator were trained adversarially to generate images of channel sections submerged at varying proportions by flash floods and debris flows. The generator's goal was to produce images that were as realistic as possible, while the discriminator's goal was to distinguish between the generated images and real ones. Through continuous iterative training, the generator was able to generate high-quality new sample data. The generated images, combined with the original photos, formed an augmented dataset, providing a rich sample resource for subsequent model training.

[0082] The generator is represented by G(z;θg), where z represents the input random noise and θg represents the generator parameters. The discriminator is represented by D(X;θd), where X represents the input image of a real channel with a flash flood and mudslide or a fake image of a channel with a flash flood and mudslide created by the generator, and θd represents the generator parameters. GAN continuously improves the generator's generation ability and the discriminator's detection ability by alternately optimizing G and D. The optimization process can be described as follows:

[0083] (6);

[0084] Where Pr represents the topography, flow state, and distribution of ditch bed materials in the real flash flood debris flow channel image; Pz represents the topography, flow state, and distribution of ditch bed materials in the generated flash flood debris flow channel image; It reflects the minimum-maximum relationship between the generator G and the discriminator D; is the probability distribution that the real data X obeys, is the potential noise Obey the prior probability distribution; when optimizing the discriminator, fix the generator and train the discriminator to classify the real data X~Pr as 1 (giving a high score) and the generated data X=G(z)~Pg as 0 (giving a low score). When optimizing the generator, fix the discriminator and train the generator to improve the quality of the generated samples so that the discriminator gives the highest possible score.

[0085] In formula (6), the optimization objective of the discriminator is about Taking the derivative and setting it to 0, we can get the theoretical optimal solution of the discriminator:

[0086] (7);

[0087] in, is the output of the discriminator, Represents the probability distribution of real data (for example, real samples in the training set), describing the probability of the occurrence of real data sample X. represents the generative model, the probability distribution of the generated data, that is, the probability law followed by the generated sample X; and the theoretical optimal solution of the generator should make the topography, water flow state, and distribution of ditch bed materials in the generated flash flood and debris flow channel image exactly the same as the topography, water flow state, and distribution of ditch bed materials in the real flash flood and debris flow channel image, that is, Pg = Pr. At this time, the discriminator always has D*(X) = 0.5, that is, it is completely unable to distinguish between true and false data samples, and both parties reach a Nash equilibrium.

[0088] However, GAN training is very difficult and prone to gradient vanishing and mode collapse problems. Substituting the optimal discriminator form of formula (7) into formula (6), it can be derived that:

[0089] (8);

[0090] in, It is the value function / loss function of GAN, which is used to characterize the adversarial game process between the generator and the discriminator. The optimal discriminator is in the GAN (Generation Adversarial Network) framework, where the discriminator's goal is to distinguish real samples (from ) and generate samples (from ), express and The JS (Jensen-Shannon) divergence between these two distributions. Formula (8) shows that the essence of the optimization of the discriminator is to measure the JS distance between the real data distribution and the generated data distribution; correspondingly, the optimization process of the generator is to minimize the JS distance, thereby achieving the alignment of Pg to Pr.

[0091] After the generative model is trained, the trained generator is used to augment samples from real flash flood and debris flow channel photos. The method uses real flash flood and debris flow channel photos as conditional inputs, combined with a random noise vector, to generate images showing different proportions of the channel cross-section being inundated by the flash flood and debris flow. For example, images are generated for various random conditions, such as a channel cross-section being inundated by the flash flood and debris flow at one-quarter, one-half, and 100% of the channel cross-section. These generated images, combined with the original photos, form an augmented dataset, providing a rich sample resource for subsequent deep learning model training.

[0092] Evaluation metrics include image realism, diversity, and consistency;

[0093] Realism can be measured by manual annotation or automatic evaluation algorithms to assess how visually similar the generated images are to the real images;

[0094] Diversity is measured by calculating the topography, flow conditions, and gully bed material distribution of the generated flash flood and debris flow channel images, ensuring that the generated samples cover different flash flood and debris flow inundation conditions. The IS is used to assess the quality and diversity of the generated images. A higher IS indicates better quality and higher diversity of the generated samples. In this invention, when the IS is greater than 6, the generated image is considered to meet the diversity requirements. Diversity is evaluated as follows:

[0095] Feature extraction: First, the Inception model is used to extract topographic features, flow conditions, and gully bed material features from the generated and real flash flood debris flow channel images.

[0096] Calculate probability distribution: For each image, calculate its probability distribution for each category;

[0097] Calculate the within-class variance and between-class variance:

[0098] Intra-class variance: For each generated image, the entropy (i.e. uncertainty) of its probability distribution is calculated. This reflects whether the distribution of generated images in each category is concentrated;

[0099] Inter-class variance: Calculates the KL divergence (Kullback-Leibler divergence) between the average probability distribution of all generated images and the average probability distribution of real images. This reflects the difference between the category distribution of generated images and the category distribution of real images;

[0100] Calculate Inception Score: IS is the harmonic mean of intra-class variance and inter-class variance, and the formula is as follows:

[0101] (9);

[0102] in, is an exponential function used to convert the logarithmic result back to the original dimension; For real data distribution , i.e., traversing the generated model input (or the conditions associated with the real data samples) and calculating the overall average; KL divergence (Kullback-LeiblerDivergence), which measures the Under the KL-valued distribution, the difference between the generated sample category distribution p(Y|X) and the marginal category distribution p(Y) (which can be understood as the overall category distribution of the generated samples) is calculated. A larger KL-valued divergence indicates a more significant difference between the conditional distribution and the marginal distribution, reflecting clear and diverse categories of the generated samples. X represents a generated sample, Y represents the text label corresponding to the sample, p(Y) is the marginal distribution, or the marginal category distribution, and p(Y|X) is the conditional distribution, or the sample category distribution. A larger KL-valued divergence between the marginal distribution p(Y) and the conditional distribution p(Y|X) indicates higher quality of the generated image.

[0103] Consistency is measured by comparing the generated images of flash flood and debris flow channels with photographs of real flash flood and debris flow channels in terms of channel topography, flow patterns, and channel bed material distribution. This ensures that the generated images have been appropriately augmented while preserving the original information. By evaluating the quality of the dataset, we can further optimize the parameters of the generated model and improve the quality and effectiveness of the augmented samples.

[0104] Mean-Square Error (MSE), Structural Similarity (SSIM) and Peak Signal to Noise Ratio (PSNR) are used to measure the consistency of the generated channel image.

[0105] Mean Square Error (MSE) is a commonly used indicator to measure the difference between two images. It is achieved by calculating the square of the difference between the corresponding pixel values ​​in the two images and then finding the average value. The smaller the MSE value, the more similar the two images are and the smaller the image quality loss is. The calculation formula is:

[0106] (10);

[0107] Among them, M and N represent the width and height of the image respectively. and are the images to be compared, I1(x,y) and I2(x,y) represent the pixel values ​​at pixel position (x,y) of the first and second images respectively. For the double summation operation, traverse the image and All pixel positions, calculate each pixel position and The squares of the pixel value differences are then added together to summarize the overall pixel difference between the two images.

[0108] Structural Similarity (SSIM) is based on the characteristics of the human visual system and simulates the human eye's perception of image differences. The SSIM value is between -1 and 1. The closer the value is to 1, the more similar the two images are. The calculation formula is:

[0109] (11);

[0110] in, and are the images to be compared, μ1 is The average brightness of the image, μ2 is The average brightness of the image; yes The average squared brightness of the image, yes The average square brightness of the image; σ1 2 yes Image brightness fluctuation, contrast variance, σ2 2 yes Image brightness fluctuation and contrast variance; yes and The covariance of two images reflects the correlation between the brightness changes of the two images. c1 and c2 are two small constants used to maintain the stability of the calculation and prevent the denominator from being zero.

[0111] molecular and the denominator The combination of adjustment to achieve image brightness matching; molecules and the denominator The combined adjustment achieves image contrast matching.

[0112] Through covariance Correlate the structural change trends of the two images. If the structures are similar (for example, the edges and textures have the same direction), the covariance will be larger and the overall structure of the images will be similar.

[0113] final The value of is in the range of [−1,1]. The closer it is to 1, the more similar the two images are in brightness, contrast, and structure. If it is negative, it means that the image structures are very different (or even opposite). It is used to evaluate the effects of image denoising, compression, and super-resolution.

[0114] Peak Signal to Noise Ratio (PSNR) compares the difference between the original image and the reconstructed image and converts it into a logarithmic value. The larger the PSNR value, the more similar the two images are. The calculation method of Peak Signal to Noise Ratio is:

[0115] (12);

[0116] in, and are the images to be compared, is the peak signal-to-noise ratio of the image, Is the mean square error of the image, which reflects the average of the squares of the differences in the corresponding pixel values ​​of the two images. Represents the maximum possible pixel value range of the image. For example, for an 8-bit image, L=255 represents the maximum possible pixel value, that is, the peak value.

[0117] Calculate first Get the error degree between images, and then pass The error is converted into a ratio relative to the peak value and finally the error is converted into a ratio relative to the peak value. 10 The larger the PSNR value, the better. The larger the value, the better the image quality. Generally, in the field of image compression and restoration, it can be used to evaluate the effectiveness of the algorithm. For example, the higher the PSNR between the compressed image and the original image, the smaller the compression loss.

[0118] Example 4;

[0119] Based on Example 1, for early warning model training and testing, a convolutional neural network (CNN) was selected as the deep learning model;

[0120] The specific early warning model is constructed as follows:

[0121] The input layer is used to receive the preprocessed flash flood debris flow channel images, and the image size is uniformly adjusted to 224×224 pixels to meet the model input requirements.

[0122] The convolutional layer consists of the first and second convolutional layers. The first convolutional layer contains 32 convolution kernels with a size of 3×3 and a stride of 1. It uses the ReLU activation function to extract low-level image features such as edges and textures. The second convolutional layer contains 64 convolution kernels with a size of 3×3 and a stride of 1. It also uses the ReLU activation function to further extract higher-level features.

[0123] The pooling layer is a maximum pooling layer set after each convolutional layer. The pooling window size is 2×2 and the step size is 2. It is used to reduce the dimension of the feature map and reduce the amount of calculation while retaining important features.

[0124] The fully connected layer consists of the first and second fully connected layers. After convolution and pooling, the fully connected layer flattens the feature map into a one-dimensional vector and inputs it to the fully connected layer. The first fully connected layer contains 128 neurons and uses the ReLU activation function; the second fully connected layer contains 64 neurons and also uses the ReLU activation function.

[0125] The output layer contains two neurons, which correspond to the classification results of whether flash floods and debris flows occur or not, and use the softmax activation function to output the probability distribution.

[0126] The cross entropy loss function is used as the optimization objective of the model. The cross entropy loss function can effectively measure the difference between the model's predicted value and the true value, guiding the model to learn the correct classification boundary.

[0127] The Adam optimization algorithm was selected for model training. It combines the advantages of momentum optimization and adaptive learning rate adjustment, enabling rapid convergence and finding optimal model parameters during training. The parameters of the Adam optimization algorithm were set as follows: an initial learning rate of 0.001, β1 = 0.9, and β2 = 0.999.

[0128] Based on the dataset size and model complexity, the number of training rounds was set to 50. During training, performance metrics on the validation set, such as accuracy, recall, and F1 value, were recorded every five rounds to promptly detect overfitting or underfitting of the model and adjust the training strategy.

[0129] Example 5;

[0130] Based on Example 1, for model evaluation and adjustment:

[0131] Performance evaluation involves comprehensively assessing model performance during training using metrics such as accuracy, recall, F1 value, and area under the receiver operating characteristic (ROC) curve (AUC). Accuracy measures the proportion of samples correctly classified by the model; recall measures the proportion of positive samples correctly identified by the model; the F1 value is the harmonic mean of accuracy and recall, comprehensively reflecting the model's classification performance; and the AUC value indicates the model's classification ability at different thresholds; values ​​closer to 1 indicate better model performance.

[0132] The model adjustment strategy involves adjusting the model based on its performance on the validation set. If the model overfits, such as high accuracy on the training set but low accuracy on the validation set, the following strategies can be used for adjustment: increasing the diversity of data augmentation and introducing more data perturbations; adjusting the model structure, such as adding dropout layers or reducing the number of neurons in the convolutional layer; and reducing the learning rate to make the model more stable during training. If the model underfits, such as low accuracy on both the training and validation sets, you can try increasing the model complexity, such as increasing the number of neurons in the convolutional or fully connected layers; adjusting the optimization algorithm parameters, such as increasing the learning rate or adjusting the momentum parameter; and further preprocessing and feature engineering the data to improve data quality and separability.

[0133] The algorithm in this paper needs to be run on a high-performance computing server equipped with an Intel Xeon processor, 64GB of memory, and an NVIDIA RTX 3090 GPU, running the Linux operating system. The deep learning framework used is PyTorch 1.12.0, and Python version 3.9.

[0134] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A flash flood and debris flow monitoring and early warning method based on intelligent generation of image samples, characterized in that: The steps include: Step I): Video acquisition and transmission: Using cameras deployed at key locations along the path of the flash flood and debris flow, real-time video data of the flash flood and debris flow path is collected and transmitted to a data processing center via the network. Step II): Data preprocessing and augmentation: The collected video data is preprocessed, including image denoising and enhancement to improve image quality. AIGC technology is then used to intelligently generate the preprocessed original photos to augment the sample dataset. Step III): Early warning model training and testing: Based on the amplified sample dataset, a deep learning model is used to train and test the samples, establish an early warning model, and implement early warning monitoring of flash floods and debris flows. Step IV): Model evaluation and adjustment: Evaluate the trained model and adjust the model parameters based on the evaluation results to optimize performance.

2. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the video acquisition and transmission steps, the specific steps for deploying cameras are as follows: Step a): The cameras are arranged to cover the entire cross section of the ditch, including the ditch bed, ditch walls, and hillsides on both sides of the ditch; Step b): increasing the density of cameras in potentially dangerous areas of the trench; Step c): The camera is installed on a stable foundation and has lightning protection, waterproof and dustproof measures; Step d): The camera collects video and transmits data via wireless or wired signals, and the data transmission delay time of the camera is controlled within 1 second.

3. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the data preprocessing and amplification steps, the data preprocessing includes image denoising and enhancement operations; wherein, the image denoising adopts a Gaussian filtering method, and the image enhancement adopts a histogram equalization method.

4. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 3 is characterized in that: The image enhancement adopts the histogram equalization method to improve the contrast and brightness of the image and make the details in the image clearer: The steps of the histogram equalization method are as follows: Step 1): Calculate the percentage of pixels for each grayscale value and obtain the PDF of the histogram; Step 2): Accumulate the PDF of each gray level to get the CDF of the histogram: Step 3): Quantize the CDF and map it to the output image.

5. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the data preprocessing and amplification steps, the specific process of amplifying the sample data set using AIGC technology is as follows: the preprocessed original photos are input into a generative AI model based on a generative adversarial network, and through adversarial training of the generator and the discriminator, pictures of channel sections being flooded by mountain torrents and mudslides at different proportions are generated. The generated pictures and the original photos constitute an amplified data set.

6. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the early warning model training and testing steps, the deep learning model used is a convolutional neural network to establish an early warning model. The early warning model is constructed as follows: The input layer is used to receive pre-processed images with a uniform size of 224×224 pixels; The convolution layer includes the first convolution layer and the second convolution layer. The first convolution layer contains 32 3×3 convolution kernels with a step size of 1 and uses the ReLU activation function; the second convolution layer contains 64 3×3 convolution kernels with a step size of 1 and uses the ReLU activation function. The pooling layer is a 2×2 maximum pooling layer set after each convolutional layer with a stride of 2; The fully connected layer includes the first fully connected layer and the second fully connected layer. The first fully connected layer contains 128 neurons and uses the ReLU activation function; the second fully connected layer contains 64 neurons and uses the ReLU activation function; The output layer contains 2 neurons, and the neurons use the softmax activation function to output probability distribution.

7. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 6 is characterized in that: The model training of the early warning model adopts the cross entropy loss function as the optimization target, selects the Adam optimization algorithm, and sets the parameters as follows: the initial value of the learning rate is 0.001, β1=0.9, β2=0.999, the number of training rounds is 50, and the performance indicators of the model on the validation set are recorded every 5 rounds.

8. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the model evaluation and adjustment step, the generated sample data set is evaluated, and the evaluation indicators include fidelity, diversity and consistency; wherein, the fidelity is measured by manual annotation or automatic evaluation algorithm to evaluate the visual similarity between the generated image and the real image; the diversity is evaluated by IS, and when IS is greater than 6, the generated image meets the diversity requirement; the consistency is measured by mean square error, structural similarity and peak signal-to-noise ratio.

9. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1 is characterized in that: In the model evaluation and adjustment step, accuracy, recall, F1 value and area under the ROC curve are used as performance evaluation indicators.

10. The method for monitoring and early warning of mountain torrents and debris flows based on intelligent generation of image samples according to claim 1, characterized in that: In the model evaluation and adjustment step, the model is adjusted according to the performance indicators of the model on the validation set; when the model is overfitting, the strategy of increasing data amplification diversity, adjusting the model structure or reducing the learning rate is adopted; when the model is underfitting, the strategy of increasing model complexity, adjusting optimization algorithm parameters or further preprocessing the data is adopted.

Citation Information

Patent Citations

  • Risk prediction model training method and device, risk prediction model using method and device, equipment and medium

    CN116957059A

  • Debris flow geological disaster monitoring and early warning method

    CN119832717A

  • Debris flow intelligent monitoring method and system based on visual computing video image acquisition and analysis

    CN120032479A