A flood discharge detection model training method, device, equipment and medium
By performing various sample enhancement processes on video image data of reservoir spillways and improving the feature extraction and fusion of the YOLOv8 model, the problems of insufficient samples and poor recognition effect of reservoir flood discharge detection models were solved, and efficient and accurate flood discharge detection was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing reservoir discharge detection models suffer from problems such as high sample quantity and quality requirements, large data noise, large domain differences, class imbalance, and lack of representativeness due to the limited number of samples used in reservoir discharge events. This leads to overfitting of target recognition models and poor cross-domain transfer performance, making them difficult to use effectively in production applications.
Multiple sample augmentation strategies were employed to process video image data of the reservoir spillway, including noise injection, color transformation, geometric transformation, and weather simulation. These strategies were combined with an improved YOLOv8 model for feature extraction and fusion to construct a lightweight flood discharge detection model.
This improved the model's computational efficiency and target detection accuracy in scenarios with limited sample data, enhanced the model's generalization ability and robustness, and ensured the real-time performance and high accuracy of flood discharge detection.
Smart Images

Figure CN121095696B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence image processing, and in particular to a flood discharge detection model training method and device, equipment and medium. BACKGROUND
[0002] Reservoirs are important water conservancy facilities for pest control and benefit, and spillways, as an important part of reservoirs, play a key role in ensuring the flood control function of reservoirs. The traditional reservoir flood discharge safety hazards are mainly detected through manual inspection or video monitoring combined with manual analysis, which has the disadvantages of large management blind area, poor monitoring timeliness, lagging hidden danger analysis, and low intelligence, etc., resulting in increased flood discharge safety risk of spillways during flood season. Based on the deep learning image recognition theory, a reservoir flood discharge intelligent detection model based on video images is constructed to further improve the perception ability and safety management level of reservoirs, ensure the safe and stable operation of reservoirs during flood season, accelerate the realization of the development goal of modern management of reservoirs and smart water conservancy in the new stage, and has great significance.
[0003] However, the mainstream target detection model has relatively high requirements for the number and quality of samples, while reservoir flood discharge is not normal, so the reservoir flood discharge detection model is a typical few-shot training model. As known, few-shot data has many problems such as large data noise, large domain difference, class imbalance, lack of representativeness, etc., which leads to overfitting of the target recognition model trained by the few-shot data and poor cross-domain migration effect, and has great limitations in production application.
[0004] Therefore, in view of these problems, it is of great application value to construct a lightweight reservoir flood discharge detection model that is suitable for few-shot data, complex weather environment and high-precision real-time detection. SUMMARY
[0005] The present application provides a flood discharge detection model training method, device, equipment and medium, which can improve the calculation efficiency and target detection accuracy of the flood discharge detection model in the few-shot data scenario.
[0006] In a first aspect, the present application provides a flood discharge detection model training method, comprising:
[0007] The pre-acquired video image data of the reservoir spillway is used as a first training sample set, and a sample enhancement is performed on the first training sample set according to a preset sample enhancement strategy to obtain a second training sample set; wherein the sample enhancement strategy includes one or a combination of the following: an independent enhancement strategy, a combined enhancement strategy or an inlay enhancement strategy;
[0008] According to the second training sample set, an initial flood discharge detection model is trained to obtain a final flood discharge detection model, wherein the initial flood discharge detection model is an improved YOLOv8 model obtained by modifying a flood discharge convolution feature extraction module and updating a feature fusion module.
[0009] The embodiment of the present application balances sample distribution and increases sample quantity by providing multiple sample enhancement strategies, solves the problem of low sample quantity and narrow coverage caused by low occurrence rate of flood discharge events, thereby improving the generalization ability and robustness of the model; the calculation amount of the model is reduced by lightening the YOLOv8 model, thereby ensuring the real-time performance of the flood discharge detection, and the feature extraction ability of the model is strengthened by modifying the feature fusion module of the YOLOv8 model, thereby improving the detection accuracy, and finally the second training sample set obtained by sample enhancement is used to train the initial flood discharge detection model obtained by structural modification, thereby obtaining a final flood discharge detection model with high precision and strong robustness. Compared with the prior art, the present application can improve the calculation efficiency and target detection accuracy of the flood discharge detection model in the few sample data scene.
[0010] Further, the independent enhancement strategy includes one or a combination of the following: a noise injection strategy, a color transformation strategy, a geometric transformation strategy, or a weather simulation strategy.
[0011] The sample enhancement of the first training sample set according to the preset sample enhancement strategy to obtain the second training sample set comprises:
[0012] When the sample enhancement strategy is a noise injection strategy, a noise image of any noise type is superimposed on each original sample image in the first training sample set to obtain the second training sample set.
[0013] When the sample enhancement strategy is a color transformation strategy, each original sample image in the first training sample set is respectively color transformed according to any color transformation mode to obtain the second training sample set; wherein the color transformation mode includes an HSV transformation mode, a mean blur mode, a contrast adjustment mode, a brightness adjustment mode, or a histogram equalization mode.
[0014] When the sample enhancement strategy is a geometric transformation strategy, each original sample image in the first training sample set is respectively geometrically transformed according to any geometric transformation type to obtain the second training sample set; wherein the geometric transformation type includes a translation transformation, a rotation transformation, a scaling transformation, a mirror flip, or an image distortion.
[0015] When the sample enhancement strategy is a weather simulation strategy, each original sample image in the first training sample set is subjected to weather simulation processing through simulation simulation technology to obtain a second training sample set.
[0016] The embodiment of the application generates an enhanced sample image with injected noise through a noise injection strategy to improve the noise resistance of the flood discharge recognition model; simulates the effects of images under different lighting and different imaging conditions through a color transformation strategy to enhance the adaptability of the model to lighting changes; simulates camera perspective changes through a geometric transformation strategy to enrich the diversity of the sample space and enhance the detection accuracy of the model for images under different perspectives; covers rare weather scenarios through a weather simulation strategy to solve the problem of single environment of sample images; and finally improves the detection stability and accuracy of the model under complex natural environments through multi-dimensional sample enhancement strategies.
[0017] Further, the weather simulation strategy includes a sunlight simulation strategy, a rain simulation strategy, a cloud simulation strategy, or a snowflake simulation strategy.
[0018] The weather simulation processing of each original sample image in the first training sample set through simulation simulation technology to obtain a second training sample set includes:
[0019] When the weather simulation strategy is a sunlight simulation strategy, a pre-generated halo image is linearly superimposed on each original sample image in the first training sample set through image fusion technology to obtain a second training sample set; wherein the halo image is generated by performing light source component attenuation simulation calculation on a randomly generated light source point in the corresponding original sample image through a mirror attenuation function.
[0020] When the weather simulation strategy is a rain simulation strategy, a pre-generated rain pattern image is linearly superimposed on each original sample image in the first training sample set through image fusion technology to obtain a second training sample set; wherein the rain pattern image is generated by color modulation and transparency control on a pre-generated rain line image; and the rain line image is generated by drawing the length and inclination angle of the rain line from each raindrop in a pre-created basic raindrop noise image.
[0021] When the weather simulation strategy is a cloud simulation strategy, each original sample image in the first training sample set is subjected to fogging processing through a pre-set transmittance image to obtain a second training sample set.
[0022] When the weather simulation strategy is a snowflake simulation strategy, a second training sample set is obtained by superimposing each original sample image in the first training sample set with a pre-generated final snowflake image one by one through an image fusion technology, wherein the final snowflake image is generated by adjusting the transparency and brightness of a plurality of snowflakes in a pre-generated initial snowflake image, and the initial snowflake image is generated by randomly generating a plurality of snowflakes on a corresponding original sample image through a random point generation and a mask operation.
[0023] The embodiment of the present application simulates strong light reflection interference through the sunlight simulation strategy to avoid feature loss caused by overexposure, simulates rain streak shielding through the rain simulation strategy to improve the target recognition ability in rainy days, simulates fog concentration differences through the cloud and fog simulation strategy, adapts to snow scenes through the snowflake simulation strategy, and accurately simulates complex weather through various weather simulation strategies to make up for the lack of real data and ensure the reliability of the model in extreme weather.
[0024] Further, the first training sample set is subjected to sample enhancement according to a preset sample enhancement strategy to obtain a second training sample set, including:
[0025] When the sample enhancement strategy is a mosaic enhancement strategy, a plurality of original sample images are randomly selected from the first training sample set, and the original sample images are subjected to size scaling processing to obtain scaled sample images, so as to splice the scaled sample images into a new image with the same size as each original sample image.
[0026] The same number of original sample images are randomly selected from the remaining original sample images in the first training sample set to splice into new images with the same size, until all original sample images are spliced.
[0027] Each new image spliced is determined as the second training sample set.
[0028] The embodiment of the present application splices a plurality of samples into an image through the mosaic enhancement strategy, which on the one hand reduces local storage resource consumption while expanding samples through independent enhancement strategies and combined enhancement strategies, thereby reducing the performance requirements of the model on hardware, and on the other hand increases the number of targets in a single image to enrich positive samples and alleviate the problem of insufficient model detection performance caused by positive and negative sample imbalance.
[0029] Further, the pre-constructed initial flood discharge detection model is trained according to the second training sample set to obtain a final flood discharge detection model, including:
[0030] inputting the second training sample set into a backbone network layer of the initial flood discharge detection model to perform feature extraction through a stacked flood discharge convolution feature extraction module to generate an initial feature map set; wherein the flood discharge convolution feature extraction module comprises a CBS module, a C2f module after lightweight modification, an SPPF module and a CSAM hybrid attention mechanism;
[0031] fusing the initial feature map set through a preset multi-scale feature fusion module to obtain a multi-scale feature map; wherein the multi-scale feature fusion module is obtained by parallel connection of a hollow convolution kernel and an attention mechanism;
[0032] performing output prediction according to the multi-scale feature map to obtain an output prediction result, and performing loss evaluation on the target prediction result through a preset comprehensive loss function to obtain a loss evaluation result; wherein the comprehensive loss function comprises a classification loss function, a boundary regression loss function and a focal loss function;
[0033] updating the parameters of the initial flood discharge detection model according to the loss evaluation result to obtain a final flood discharge detection model.
[0034] The embodiment of the application extracts the initial feature map by stacking the flood discharge convolution feature extraction module composed of the CBS module, the C2f module after lightweight modification, the SPPF module and the CSAM hybrid attention mechanism, so as to use the depth separable convolution instead of the ordinary convolution, and then realize the lightweight feature extraction of the model; Through the multi-scale feature fusion module, different levels of features are fused to improve the multi-scale target detection capability; Through the classification loss, the flood discharge / non-flood discharge state discrimination is optimized, through the boundary regression loss, the gate, water flow and other targets are accurately positioned, and through the focal loss, the discrete coordinate prediction is realized to adapt to the target with fuzzy boundary and difficult to define, and the robustness is improved; Finally, through the model structure modification and loss function update, the model detection precision and efficiency are improved.
[0035] Further, the feature extraction through the stacked flood discharge convolution feature extraction module generates an initial feature map set, specifically:
[0036] a plurality of layers of combined convolution modules are used to perform multiple ordinary convolution and depth separable convolution calculations on the second training sample set in turn to obtain a first feature map set containing different levels of feature maps; wherein the combined convolution module comprises the CBS module and the C2f module after lightweight modification;
[0037] the SPPF module is used to perform multi-scale spatial pooling on the highest level feature map in the first feature map set to obtain a second feature map
[0038] The CSAM mixed attention mechanism is used for enhancing the second feature map to obtain a third feature map.
[0039] The first feature map set, the second feature map and the third feature map are merged to obtain an initial feature map set.
[0040] The improved flood discharge convolution feature extraction module is used to extract the convolution feature map, so that the deep separable convolution is used instead of the ordinary convolution, the model calculation amount is reduced, the mixed attention mechanism is used to enhance the key features and suppress the background interference, so that the balance between the model lightening and the feature enhancement is realized, the small target detection capability is improved while the resource consumption is reduced.
[0041] Further, the initial feature map set is fused by a preset multi-scale feature fusion module to obtain a multi-scale feature map, specifically:
[0042] The initial feature map set is subjected to convolution operation by a plurality of convolution blocks with different hole rates to obtain a plurality of first intermediate feature maps;
[0043] The first intermediate feature maps are subjected to feature enhancement by a preset mixed attention mechanism to obtain a plurality of second intermediate feature maps;
[0044] The second intermediate feature maps are subjected to splicing operation to obtain the multi-scale feature map.
[0045] The multi-hole rate convolution block is used to expand the receptive field and capture context information of different scales, the mixed attention mechanism is used to highlight the key features and suppress the background interference, and the splicing operation is used to integrate feature information from different sources and improve the feature richness, so that the detection robustness of the complex scene is finally improved.
[0046] In a second aspect, an embodiment of the present application provides a flood discharge detection model training device, which comprises a sample acquisition module and a model training module, wherein,
[0047] The sample acquisition module is used to select part of the pre-acquired video image data of the reservoir spillway as a first training sample set, and to obtain a second training sample set by performing sample enhancement on the first training sample set according to a preset sample enhancement strategy; wherein the sample enhancement strategy comprises an independent enhancement strategy, a combined enhancement strategy and an inlay enhancement strategy.
[0048] The model training module is used to train a pre-constructed initial flood discharge detection model according to the second training sample set to obtain a final flood discharge detection model; wherein the initial flood discharge detection model is an improved YOLOv8 model obtained by modifying the flood discharge convolution feature extraction module and updating the feature fusion module.
[0049] The embodiment of the present application provides a plurality of sample enhancement strategies through the sample acquisition module, balances sample distribution, increases sample quantity, solves the problem of small sample quantity and narrow coverage caused by low flood discharge event occurrence rate, and thus improves generalization capability and robustness of a model; the model training module is used for performing lightweight modification on a YOLOv8 model, reduces calculation amount of the model, and thus guarantees real-time performance of flood discharge detection, the feature fusion module modification is performed on the YOLOv8 model, the feature extraction capability of the model is strengthened, and thus detection precision is improved, finally, the second training sample set obtained through sample enhancement is used to train the initial flood discharge detection model obtained through structure modification, and thus a final flood discharge detection model with high precision and strong robustness is obtained.
[0050] In a third aspect, an embodiment of the present application provides a terminal device, including a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete communication with each other through the communication bus.
[0051] The memory is used for storing at least one executable instruction, and the executable instruction makes the processor execute the operations of the flood discharge detection model training method in any one of the above aspects.
[0052] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, including a stored computer program, wherein when the computer program runs, the computer readable storage medium controls a device or apparatus where the computer readable storage medium is located to execute the flood discharge detection model training method in any one of the above aspects.
[0053] The above description is only a summary of the technical scheme of the embodiment of the present application, in order to more clearly understand the technical means of the embodiment of the present application, the embodiment can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the embodiment of the present application more obvious and easy to understand, the specific embodiment of the present application is described as follows. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 A flood discharge detection model training method provided by the embodiment of the present application is shown in the figure;
[0055] Figure 2 An RSD-YOLO network structure diagram provided by the embodiment of the present application is shown in the figure;
[0056] Figure 3 A DWBlock and CBS module structure diagram provided by the embodiment of the present application is shown in the figure;
[0057] Figure 4 A LW_C2f module structure diagram provided by the embodiment of the present application is shown in the figure;
[0058] Figure 5 This is a schematic diagram of the SPPF module structure provided in an embodiment of the present invention;
[0059] Figure 6 This is a schematic diagram of the CSAM module structure provided in an embodiment of the present invention;
[0060] Figure 7 This is a schematic diagram of the multi-scale feature fusion module MSFF results provided in an embodiment of the present invention;
[0061] Figure 8 This is a technical roadmap for a flood discharge detection model provided in an embodiment of the present invention;
[0062] Figure 9 This is a structural diagram of a flood discharge detection model training device provided in an embodiment of the present invention. Detailed Implementation
[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] Example 1:
[0065] like Figure 1 As shown, an embodiment of the present invention provides a method for training a flood discharge detection model, which includes the following steps:
[0066] S11, the pre-collected video image data of the reservoir spillway is used as the first training sample set, and the first training sample set is augmented according to the preset sample augmentation strategy to obtain the second training sample set; wherein, the sample augmentation strategy includes one or more of the following: independent augmentation strategy, combined augmentation strategy or mosaic augmentation strategy.
[0067] In a specific embodiment, the collection process of the video image data of the reservoir spillway includes: obtaining real-time video images of the reservoir spillway through a video monitoring system, if the spillway is currently in a flood discharge state (generally in the case of forecast flood), adjusting the camera, focusing on collecting the gate, inlet section, discharge channel section, and side wall of the reservoir spillway, avoiding chaotic parts such as the stilling basin and the end of the efficiency facility, and collecting images of each part and saving them to the local, then fixing a certain view angle unchanged, constantly adjusting the focal length to change the image magnification (the image magnification starts from 1, until the maximum, and each time it is enlarged by 2 times), to obtain different magnification images under a certain fixed view angle, and the flood discharge samples of the same reservoir cover different weather conditions (if it is really impossible to meet, it can be compensated by sample enhancement), and the number of reservoir flood discharge image samples is increased. The above method is used to collect a round of image data of the spillway without flood discharge, covering the gate, inlet section, discharge channel section, and side wall.
[0068] In a specific embodiment, before sample enhancement, sample annotation will also be performed, wherein the sample annotation process specifically includes: using labelme annotation software to annotate the video images of the spillway, and the annotation categories are subdivided into two categories of flood discharge and non-flood discharge. When annotating, a rectangular frame is drawn to identify the target type. At the same time, to reduce the influence of precision evaluation error caused by class imbalance, the number of sample images in the flood discharge and non-flood discharge states should be as close as possible, and the ratio of the two should be controlled within the range of [0.2, 5].
[0069] It should be noted that the reservoir flood discharge is a phenomenon that the reservoir management unit discharges water to the downstream river channel through the spillway (flood discharge tunnel) in a planned and controlled manner in order to control the water level of the reservoir and ensure the safety of downstream flood control. Therefore, the reservoir flood discharge is a human-initiated, controlled, and dynamic process, and there is a spatial dependence relationship between the flood discharge state and the water flow, the spillway, and the gate. Only when the spillway and the gate are the background of the water flow can the reservoir be accurately identified and judged whether it is in flood discharge. This is different from the conventional detection target. If the spillway surface is filled with water due to the storage of the stilling basin or natural rainfall, it does not belong to flood discharge. Therefore, the difference between the spillway water accumulation and the spillway flood discharge should be distinguished in the annotation process.
[0070] In the embodiment, the independent enhancement strategy includes one or more combinations of the following: a noise injection strategy, a color transformation strategy, a geometric transformation strategy, or a weather simulation strategy.
[0071] The sample enhancement is performed on the first training sample set according to a preset sample enhancement strategy to obtain a second training sample set.
[0072] When the sample enhancement strategy is the noise injection strategy, a noise image of any noise type is superimposed on each original sample image in the first training sample set to obtain the second training sample set.
[0073] When the sample enhancement strategy is the color transformation strategy, color transformation is performed on each original sample image in the first training sample set according to any color transformation mode to obtain the second training sample set, wherein the color transformation mode includes an HSV transformation mode, a mean blur mode, a contrast adjustment mode, a brightness adjustment mode, or a histogram equalization mode.
[0074] When the sample enhancement strategy is the geometric transformation strategy, geometric transformation is performed on each original sample image in the first training sample set according to any geometric transformation type to obtain the second training sample set, wherein the geometric transformation type includes a translation transformation, a rotation transformation, a scaling transformation, a mirror flip, or an image twist.
[0075] When the sample enhancement strategy is the weather simulation strategy, weather simulation processing is performed on each original sample image in the first training sample set by simulation technology to obtain the second training sample set.
[0076] In a specific embodiment, the noise injection strategy is specifically: one of Poisson noise, uniform noise, and Gaussian noise is randomly selected to generate a noise image with the same size as the input image, and then superimposed into the input image to generate an enhanced sample image with injected noise, which helps to improve the noise resistance of the flood discharge recognition model. The mathematical principle is as follows:
[0077] I out = I in + I noise
[0078] wherein I in represents an input image, I noise represents a noise image, and I out represents an output image.
[0079] In a specific embodiment, the color transformation strategy is specifically: randomly selecting one mode from HSV transformation, mean blur, contrast adjustment, brightness adjustment, and histogram equalization to perform color transformation on the input image.
[0080] HSV transformation: is the process of converting the image from RGB color space to HSV color space, which can be realized by the built-in function of Opencv.
[0081] Mean smoothing: replacing the original value of a pixel with the average value of all pixels in its neighborhood (such as 3*3 neighborhood), which can be realized by the built-in function of Opencv.
[0082] Contrast adjustment: first performing HSV transformation on the image, then adjusting the value range of the brightness component V (such as limiting it to the range of 5% to 95%), and then returning to the RGB space through inverse operation, adjusting the difference between light and dark of the image to achieve contrast enhancement / weakness, which can be realized by the built-in function of Opencv.
[0083] Brightness adjustment: first performing HSV transformation on the image, then increasing or decreasing the same value on the brightness component V of all pixels, and then returning to the RGB space through inverse operation, adjusting the brightness value of the image as a whole or in part to improve the brightness of the image and thus improve the visual perception, which can be realized by the built-in function of Opencv.
[0084] Histogram equalization: using a transformation function to map the input image's gray level to the output image, making the output image's gray levels relatively evenly distributed, thus enhancing the contrast of the image, which can be realized by the built-in function of Opencv.
[0085] In a specific embodiment, the geometric transformation strategy is specifically: randomly selecting one mode from translation transformation, rotation transformation, scaling transformation, mirror flip, and image distortion to perform geometric transformation on the input image, simulating the scene effects of different imaging conditions such as camera position, focal length, and viewing angle changes, to enrich the number and types of image samples.
[0086] In this embodiment, the weather simulation strategy includes sunlight simulation strategy, rain simulation strategy, cloud and fog simulation strategy, or snowflake simulation strategy.
[0087] The weather simulation strategy includes sunlight simulation strategy, rain simulation strategy, cloud and fog simulation strategy, or snowflake simulation strategy.
[0088] When the weather simulation strategy is a sunlight simulation strategy, a pre-generated halo image is linearly superimposed on each of the original sample images in the first training sample set by an image fusion technique to obtain a second training sample set; wherein the halo image is generated by performing attenuation simulation calculation on a light source component of a randomly generated light source point in a corresponding original sample image by a mirror attenuation function;
[0089] When the weather simulation strategy is a rain simulation strategy, a pre-generated rain streak image is linearly superimposed on each of the original sample images in the first training sample set by an image fusion technique to obtain a second training sample set; wherein the rain streak image is generated by color modulation and transparency control on a pre-generated rain line image; and the rain line image is generated by drawing the length and inclination angle of a rain line from each raindrop in a pre-created basic raindrop noise image;
[0090] When the weather simulation strategy is a cloud and fog simulation strategy, each of the original sample images in the first training sample set is foggy processed by a pre-set transmittance image to obtain a second training sample set;
[0091] When the weather simulation strategy is a snowflake simulation strategy, a pre-generated final snowflake image is superimposed on each of the original sample images in the first training sample set by an image fusion technique to obtain a second training sample set; wherein the final snowflake image is generated by adjusting the transparency and brightness of a plurality of snowflakes in a pre-generated initial snowflake image; and the initial snowflake image is generated by random point generation and mask operation to generate a plurality of snowflakes on a corresponding original sample image.
[0092] In a specific embodiment, the sunlight simulation strategy specifically includes: first, randomly selecting a rectangular region in an image as a simulated light region, and randomly selecting a point in the region as the center position of the sunlight; then, assuming that the color value of the light source S is (255, 235, 205), performing attenuation simulation calculation on the brightness component of the light source using a radial attenuation function at the center of the light source to generate a halo image with the light source as the center and the brightness of the light spot gradually decreasing; then, applying a Gaussian blur function to the halo image for gradual smoothing; and finally, linearly superimposing the generated halo image on the background image using an image fusion technique to realize the scene effect of simulating sunlight irradiation. Wherein, the calculation method of the radial attenuation function is as follows:
[0093]
[0094] I out = (1 - a) x I in + a x I sun
[0095] where λ is the attenuation coefficient, R represents the sun spot radius, I sun represents the simulated halo map, S represents the light source color value, S c represents the light spot center position, I in represents the input image, I out represents the output image, and α represents the weight, generally in the range of [0, 1], and the default value is 0.2.
[0096] In a specific embodiment, the rain simulation strategy is specifically: first, a basic raindrop noise map is created using a random number generator, then a rain line map is drawn from each raindrop as a starting point according to the length and inclination angle (i.e. the angle with the vertical direction) of the rain line, the appearance characteristics of the rain pattern are controlled by adjusting the color and transparency properties of the rain, then the rain pattern is blurred using Gaussian smoothing technology, then the generated rain pattern image is linearly superimposed with the background image using image fusion technology, and finally the fusion image is processed using histogram equalization, thereby simulating the scene effect of a rainy day. The core mathematical expression is as follows:
[0097] I out = (1-α) × I in + α × I rain
[0098] where I out represents the rain map, I in represents the background image, and I rain is the rain pattern image, and α represents the weight, generally in the range of [0, 1], and the default value is 0.2.
[0099] In a specific embodiment, the cloud and fog simulation strategy is specifically: first, a point is randomly selected in the image as the fogging center point, and its atmospheric light value A can be set to (255, 255, 255) or a color value close to white, and a transmittance map with the same size as the input image is set, and the default initial value of the transmittance map is 0; then a plurality of rectangular / circular regions are randomly set as fogging regions, and the distance of each pixel in the fogging region to the fogging center point is calculated, and the transmittance value of each pixel point is calculated according to the distance according to the transmittance calculation formula (6), and the pixel value of the corresponding position in the transmittance map is updated with the transmittance value, to obtain a complete transmittance map corresponding to the input image, and then the fogging model is substituted to obtain a fogging image, and finally the fogging image is subjected to brightness enhancement and color transformation using a contrast enhancement algorithm, thereby achieving simulation of a foggy scene. The process of simulating the fog map and calculating the transmittance is as follows:
[0100] I haze = I in t(x) + A[1-t(x)];
[0101] t(x) = e -βd(x)
[0102] wherein, I haze represents the simulation of the fog map, I in represents the background image, t(x) represents the transmittance, A represents the atmospheric light value, d(x) represents the distance from the pixel point to the center point of the fog, and β represents the scattering coefficient, which is set to -0.3502 by default. By adjusting the scattering coefficient β in the formula, different concentrations of cloud and fog effect simulation can be realized, making the simulation effect of the cloud and fog scene more rich.
[0103] In a specific embodiment, the snowflake simulation strategy is as follows: first, a random point is generated in the image range using a random number generator as the position of the snowflake center; then the shape of the snowflake is set to a circle, and the size of the snowflake is defined by adjusting the radius of the pixel point; then the generation density and distribution range of the snowflake are controlled by adjusting the generation frequency of the random point, and the generation of the snowflake is increased or decreased in a specific area of the image through mask operation; the appearance characteristics of the snowflake are controlled by adjusting the transparency and brightness properties of the snowflake to simulate the snow scene effect under different lighting conditions; then the generated snowflake image is superimposed with the original image using image fusion technology, and finally the color and contrast of the fused image are adjusted to ensure that the visual effect of the snow-added image is natural, thereby realizing the simulation and simulation of the snow scene. The core mathematical expression is as follows:
[0104] I out = (1-α) × I in + α × I snow
[0105] wherein, I out represents the rain map, I in represents the background image, I snow is the rain streak image, and α represents the weight, which is in the range of [0, 1] and is set to 0.2 by default.
[0106] In a specific embodiment, in order to simulate the complex conditions of the natural environment, a combined enhancement strategy is designed to simulate complex scenes, so as to expand the scene types and the number of samples, and to supplement and improve the single enhancement strategy. There are various forms of combined enhancement, including binary combination, ternary combination and the like. According to the characteristics of the number of samples and the target scene, a reasonable choice can be made to balance the calculation efficiency and the number of samples. The binary combination can be combined by using any two of the following strategies: ① injecting noise, ② color transformation, ③ geometric transformation, and ④ weather simulation, to generate new scene synthesis images. There are multiple mode choices in each single strategy, so the binary combination can be the combination of different strategies, or the combination of different modes in a single strategy. In this way, the number of combinations is greatly increased, thereby improving the richness of the scene. The ternary combination can be implemented by analogy with the binary combination strategy, except that the number of combined strategies is increased to 3. The rest can be implemented by analogy.
[0107] In the embodiment, the sample enhancement is performed on the first training sample set according to the preset sample enhancement strategy to obtain a second training sample set, including:
[0108] When the sample enhancement strategy is the mosaic enhancement strategy, a plurality of original sample images are randomly selected from the first training sample set, and size scaling processing is performed on the original sample images to obtain scaled sample images, so as to splice the scaled sample images into a new image with the same size as each original sample image.
[0109] Continue to randomly select a same number of original sample images from the remaining original sample images in the first training sample set to splice into new images with the same size, until all original sample images are spliced;
[0110] Each new image spliced is determined as the second training sample set.
[0111] In a specific embodiment, the mosaic enhancement strategy is specifically: four sample pictures are spliced into a large picture as a new training sample image through random scaling, random cropping and random arrangement, which can increase the number of targets in a single image, enrich positive samples, alleviate the problem of insufficient model detection performance caused by positive and negative sample imbalance, and reduce local storage resource consumption through dynamic generation in memory, thereby reducing the performance requirements of the model on hardware. The specific implementation process is as follows: first, all image samples in the current directory are constructed into a set, and the loop starts from the first sample, and three images are randomly selected from the remaining N-1 sample images to form four sample images with the current sample. Then a new image with all black background color is generated, which has the same size as a single image sample. A point is randomly selected in the image, and the horizontal and vertical coordinates of the point are used as boundaries to divide the new image into four sub-regions. Then each sub-region is replaced with the four images. In order to maintain the integrity of the image during the replacement process, necessary size scaling or cropping of the image to be replaced is performed. At the same time of performing size scaling or cropping of the image, the processing process is also mapped to the labeled target box. When filling the area, the coordinates of the target box are also adjusted synchronously according to the difference of the original point, so as to realize synchronous transformation of the sample image and the detection target, thereby completing the mosaic enhancement of the sample.
[0112] S12, training an initial flood discharge detection model pre-constructed according to the second training sample set to obtain a final flood discharge detection model; wherein the initial flood discharge detection model is an improved YOLOv8 model obtained through modification of a flood discharge convolution feature extraction module and update of a feature fusion module.
[0113] In a specific embodiment, the flood discharge detection model trained by the present application (hereinafter referred to as RSD-YOLO model) is obtained by modification and optimization based on the YOLOv8 model, as shown in Figure 2 The RSD-YOLO model is composed of three parts: backbone network, neck network and target detection head. The main function of the backbone network is to extract shallow spatial and deep semantic features. The neck network fuses network features from different levels through the aggregation of FPN (bottom-up) and PAN (top-down) multi-level network structures. The function of the target detection head is to predict target types and boundary regression. As shown in Figure 3As shown in the figure, the CBS module is a basic module that includes three basic operations of two-dimensional convolution (Conv2d), batch normalization (BatchNormalize), and activation function (SiLU), and the feature map is down-sampled (the width and height of the feature map become half of the original) once for each execution of the CBS module. The DWBlock is designed based on the design concept of the ResNet residual network, which sums the input feature map and the feature map after two times of deep separable convolution, can overcome the problem of gradient disappearance or gradient explosion that easily occurs in the deepening process of the deep learning network, and realize deeper feature extraction and expression;
[0114] As shown in the figure, Figure 4 The LW_C2f module is also a feature extraction module, but this module does not change the size of the feature map. By adjusting the number of DWBlock, the transmission and reuse of shallow features in the module can be realized, and the use of deep separable convolution instead of ordinary convolution in DWBlock can further reduce the parameter amount and make the model more lightweight;
[0115] As shown in the figure, Figure 5 The SPPF module is a spatial multi-scale feature fusioner, which uses a serial cascade way to perform maximum pooling operation on the same feature map using multiple pooling kernels of different sizes (such as 5x5, 9x9, 13x13), simulates different receptive fields to capture different scale context information in the feature map, to enhance the model's ability to understand different size objects;
[0116] As shown in the figure, Figure 6 The CSAM attention module is a lightweight hybrid attention mechanism composed of channel attention and spatial attention in series, which can not only enhance channel information, but also enhance spatial information, highlight key features, suppress background interference, and improve small target detection capability;
[0117] Upsample is an up-sampling operation, which is a key component of the feature pyramid network (FPN), and its core function is to increase the spatial resolution of the feature map through the nearest neighbor / bilinear interpolation algorithm to adapt to the need of feature fusion at higher resolution scale;
[0118] Concat is a splicing operation that can realize the stacking of feature maps of the same size in the channel dimension, integrate feature information from different sources, and improve the feature richness;
[0119] As shown in the figure, Figure 7As shown, MSFF is a cross-domain feature fusion module. The MSFF module first simulates convolution operations of different receptive fields through three convolution blocks with different hole rates, then respectively performs feature enhancement through CSAM modules, finally splices three groups of features and performs gradient flow optimization through a residual module, and finally restores the channel number through a 1*1 convolution to efficiently fuse low-resolution but semantically rich deep features with high-resolution and detailed shallow feature information.
[0120] In the embodiment, the training of the pre-constructed initial flood discharge detection model according to the second training sample set to obtain the final flood discharge detection model comprises: inputting the second training sample set into the backbone network layer of the initial flood discharge detection model to perform feature extraction through the stacked flood discharge convolution feature extraction module to generate an initial feature map set; wherein the flood discharge convolution feature extraction module comprises a CBS module, a C2f module after lightweight modification, an SPPF module and a CSAM hybrid attention mechanism; the initial feature map set is fused through a pre-set multi-scale feature fusion module to obtain a multi-scale feature map; wherein the multi-scale feature fusion module is obtained by parallel connection of a hole convolution kernel and an attention mechanism; output prediction is performed according to the multi-scale feature map to obtain an output prediction result, and a loss evaluation result is obtained by performing loss evaluation on the target prediction result through a pre-set comprehensive loss function; wherein the comprehensive loss function comprises a classification loss function, a boundary regression loss function and a focal loss function; the parameters of the initial flood discharge detection model are updated according to the loss evaluation result to obtain the final flood discharge detection model.
[0121] In the embodiment, the feature extraction through the stacked flood discharge convolution feature extraction module to generate an initial feature map set is specifically: a plurality of layers of continuous combined convolution modules are used to sequentially perform multiple ordinary convolution and depth separable convolution calculations on the second training sample set to obtain a first feature map set containing feature maps of different levels; wherein the combined convolution module comprises the CBS module and the C2f module after lightweight modification; the SPPF module is used to perform multi-scale spatial pooling on the highest level feature map in the first feature map set to obtain a second feature map; the CSAM hybrid attention mechanism is used to enhance the second feature map to obtain a third feature map; and the first feature map set, the second feature map and the third feature map are combined to obtain the initial feature map set.
[0122] In a specific embodiment, the generation process of the initial feature map is realized by a backbone network module of the RSD-YOLO model; the process of generating the initial feature map by the backbone network module is specifically as follows: the image sample is first subjected to feature extraction by 2 times of CBS module, the width and height of the extracted feature map are reduced to 1 / 4 of the input, and the feature channel number is increased to 128 dimensions; then, the shallow layer feature information is continuously transmitted to the deep layer feature through multiple DWBlock of the LW_C2f module, while the size of the feature map is kept unchanged, to obtain the P1 level feature map; then, the CBS+LW_C2f module is executed three times in sequence to obtain the P2, P3 and P4 level feature maps; then, the P4 level feature map is subjected to multi-scale spatial pooling by the SPPF module to obtain the P5 feature map, to capture the context information of different scales and enhance the feature expression capability for different size targets, and then the important channel and position information of the P4 layer feature map are further enhanced by the CSAM module to obtain the P6 feature map, to suppress redundant and noise information, and finally the P2, P3 and P6 initial feature maps are sent to the neck network layer for multi-scale feature fusion.
[0123] In the embodiment, the initial feature map set is fused by a preset multi-scale feature fusion module to obtain a multi-scale feature map, specifically: a plurality of first intermediate feature maps are obtained by performing convolution operation on the initial feature map set through a plurality of convolution blocks with different hole rates; a plurality of second intermediate feature maps are obtained by performing feature enhancement on the first intermediate feature maps through a preset hybrid attention mechanism; and the multi-scale feature map is obtained by performing splicing operation on the second intermediate feature maps.
[0124] In a specific embodiment, the generation process of the multi-scale feature map is realized by a neck network module of the RSD-YOLO model; the process of generating the multi-scale feature map by the neck network module is specifically as follows: the P6 feature map output by the terminal CSAM module of the backbone layer network is subjected to 2x upsampling, spliced with the P3 layer feature map through Concat, and then subjected to adaptive fusion by the MSFF module to obtain the P7 feature map; the P7 is subjected to 2x upsampling again, spliced with the P2 feature map through Concat, and then subjected to adaptive fusion by the MSFF module to obtain the P8 feature map; the P8 is subjected to two-dimensional convolution downsampling, spliced with the P7 feature map through Concat to obtain the P9 feature map; the P9 is subjected to two-dimensional convolution downsampling, spliced with the P6 feature map through Concat to obtain the P10 feature map; and finally the P8, P9 and P10 feature maps of the three levels are combined to obtain the multi-scale feature map.
[0125] In a specific embodiment, the multi-scale feature maps composed of P8, P9 and P10 are respectively subjected to class prediction and regression prediction by the target detection head, and the decoupling head is trained for precision evaluation and reverse gradient update by using a comprehensive loss function including classification loss, regression loss and object loss. The comprehensive loss function is generally divided into three parts: classification loss L cls , boundary regression loss L bbox and focal loss L dfl .
[0126] L = λ1L cls + λ2L bbox + λ3L dfl
[0127] Wherein, L cls represents the classification loss between the predicted type of the target and the real type, L bbox represents the regression loss between the predicted boundary of the target and the real boundary, and L dfl represents the focal loss between the predicted boundary of the target and the real boundary; λ1, λ2 and λ3 represent the weights of the three types of losses, respectively. Considering the importance of the positioning accuracy of the boundary box to the identification of the flood discharge scene, the class is the second, therefore, in the application, λ1 = 2, λ2 = 7 and λ3 = 1.
[0128] L cls The classification loss adopts the cross-entropy loss of binary form, and its mathematical expression is as follows:
[0129]
[0130] Wherein, c represents the class number, y c represents the probability that the model accurately predicts the target belonging to class c, and C represents the total number of classes.
[0131] L bbox The boundary regression loss is further optimized on the basis of CIoU, and the predicted box is more accurate and the convergence speed is faster by comprehensively considering the area, center point and aspect ratio. The boundary regression loss function calculated in this paper is shown in the following formula:
[0132]
[0133] Wherein, IoU is the ratio of the intersection area to the union area of the predicted box and the real box (intersection and union ratio), d represents the diagonal length of the minimum bounding rectangle of the predicted boundary box and the real boundary box, α is a penalty term that can be adjusted, the greater α is, the more obvious the penalty effect is, and the default value is 1, and ρ(b, b gt ) represents the Euclidean distance between the center of the predicted box and the real box.
[0134] DFL models continuous coordinate values as probability distributions over discrete intervals, and obtains the final predicted coordinates through weighted summation (integration). Assuming the true coordinates of the target boundary are y, they are discretized into n coordinate points {y0, y1, y2, ..., y...}. n-1 The model predicts the coordinates y. i The probability is P(y) i If the boundary coordinates of the target are then predicted, the final predicted boundary coordinates are:
[0135]
[0136] L dfl The focus loss uses a binary form of cross-entropy loss metric, selecting the interval [y] closest to the predicted value y. i ,y i+1 To simplify the calculation, the focus loss is calculated as follows:
[0137] L dfl =-[(y i+1 -y)log(P(y i ))+(yy i )log(P(y i+1 ))]
[0138] Among them, y i y represents the discrete coordinates of the ground truth bounding box, and y represents the coordinates of the bounding box predicted by the model.
[0139] In one specific embodiment, the training process of the RSD-YOLO model is as follows:
[0140] 1) First, convert the annotation files of the spillway images to the standard YOLO format, and then place them in the same directory as the corresponding image data as training samples. Divide the training samples into training set and validation set at a ratio of 2:1, and use few-shot augmentation technology to expand the training set data to increase the number of training set samples by at least 10 times.
[0141] 2) The weights and bias parameters of the RSD-YOLO model were initialized using random numbers (ranging from [0,1]). The total number of iteration epochs was set to 100, the initial learning rate was 0.001, the gradient descent was solved using the AdamW algorithm, the number of samples in a single batch was 8, the momentum of the weight update was 0.937, and the decay rate of the weights was 0.00005.
[0142] 3) From the first round of training, read the first batch of training set data in a loop, initialize the RSD-YOLO model with the current w and b to calculate the loss and gradient value of the batch, and update the weight w and bias parameter b by the back propagation method, judge whether all batches of samples are read, if yes, proceed to the next round of training, otherwise continue the next batch of loop.
[0143] 4) According to the average loss value of all batches of training set data as the current round of loss value, and calculate the current precision index, judge whether the current round reaches the set maximum iteration round or the training precision no longer improves, if yes, stop iteration calculation, output the final model weight w and bias parameter b, otherwise continue the next round of training.
[0144] 5) The trained RSD-YOLO model is evaluated by using the validation set data. The mean average precision (mAP) is commonly used in the field of target detection to evaluate the detection performance of the model. This application uses mAP0.5 and mAP0.5:0.9 two evaluation indexes, which represent the mAP value when the IoU threshold is 0.5 and the average mAP value when the IoU threshold is in the range of [0.5, 0.55, 0.6,..., 0.95]. The AP value represents the area under the precision-recall curve, and the larger the value, the higher the model detection accuracy. There is an inverse constraint relationship between them. The calculation formulas of precision, recall and AP are as follows:
[0145]
[0146] Where P represents precision, R represents recall, TP represents the number of correctly classified positive samples, FP represents the number of false negative samples, and FN represents the number of missed positive samples.
[0147] In a specific embodiment, when the loss function value of the model on the training set tends to be stable and the validation set accuracy is higher than 90%, the construction of the RSD-YOLO model is completed. In the inference stage, the real-time video image data of the reservoir is obtained through the Guangdong Water Resources Video Monitoring System, then each frame of image is read and the RSD-YOLO model is called for real-time detection and analysis, and the current flood discharge state information of the reservoir is output, and abnormal flood discharge phenomenon is warned, helping the reservoir management personnel to find the safety hidden danger of the spillway in time, reducing the risk of poor reservoir flood discharge.
[0148] In order to better illustrate the working principle and step flow of the present application, see Figure 8 An example of a technical roadmap of a flood discharge detection model provided by an embodiment of the present application.
[0149] The embodiment of the present application balances the sample distribution and increases the sample quantity by providing multiple sample enhancement strategies, solves the problem of small sample quantity and narrow coverage caused by low flood discharge event occurrence rate, thereby improving the generalization ability and robustness of the model; the calculation amount of the model is reduced by lightening the YOLOv8 model, thereby ensuring the real-time performance of the flood discharge detection, and the feature extraction ability of the model is enhanced by reforming the feature fusion module of the YOLOv8 model, thereby improving the detection accuracy, and finally using the second training sample set obtained through sample enhancement to train the initial flood discharge detection model obtained through structural reform, thereby obtaining a final flood discharge detection model with high precision and strong robustness. Compared with the prior art, the present application can improve the calculation efficiency and target detection accuracy of the flood discharge detection model in the few-sample data scene.
[0150] Embodiment two:
[0151] As shown in Figure 9 , the present embodiment provides a flood discharge detection model training device, which comprises a sample acquisition module 001 and a model training module 002, wherein,
[0152] The sample acquisition module 001 is used for selecting part of the pre-acquired video image data of the reservoir spillway as a first training sample set, and performing sample enhancement on the first training sample set according to a preset sample enhancement strategy to obtain a second training sample set; wherein the sample enhancement strategy includes an independent enhancement strategy, a combined enhancement strategy and an inlay enhancement strategy;
[0153] In the present embodiment, the independent enhancement strategy includes one or more combinations of the following: noise injection strategy, color transformation strategy, geometric transformation strategy or weather simulation strategy; the sample acquisition module 001 performs sample enhancement on the first training sample set according to the preset sample enhancement strategy to obtain a second training sample set, including:
[0154] When the sample enhancement strategy is a noise injection strategy, a noise image of any noise type is superimposed on each original sample image in the first training sample set to obtain a second training sample set;
[0155] When the sample enhancement strategy is a color transformation strategy, each original sample image in the first training sample set is respectively color transformed according to any color transformation mode to obtain a second training sample set; wherein the color transformation mode includes an HSV transformation mode, a mean blur mode, a contrast adjustment mode, a brightness adjustment mode or a histogram equalization mode;
[0156] When the sample enhancement strategy is a geometric transformation strategy, for each original sample image in the first training sample set, a second training sample set is obtained by performing geometric transformation according to any geometric transformation type, wherein the geometric transformation type includes translation transformation, rotation transformation, scaling transformation, mirror flip or image distortion.
[0157] When the sample enhancement strategy is a weather simulation strategy, a second training sample set is obtained by performing weather simulation processing on each original sample image in the first training sample set through simulation technology.
[0158] The model training module 002 is configured to train a pre-constructed initial flood discharge detection model according to the second training sample set to obtain a final flood discharge detection model, wherein the initial flood discharge detection model is an improved YOLOv8 model obtained by updating a flood discharge convolution feature extraction and a feature fusion module.
[0159] In this embodiment, the model training module 002 trains a pre-constructed initial flood discharge detection model according to the second training sample set to obtain a final flood discharge detection model, including: the model training module 002 inputs the second training sample set into the backbone network layer of the initial flood discharge detection model to perform feature extraction through the stacked flood discharge convolution feature extraction module to generate an initial feature map set; wherein the flood discharge convolution feature extraction module includes a CBS module, a C2f module after light-weight modification, an SPPF module and a CSAM hybrid attention mechanism; the initial feature map set is fused through a pre-set multi-scale feature fusion module to obtain a multi-scale feature map; wherein the multi-scale feature fusion module is obtained by parallel connection of a hollow convolution kernel and an attention mechanism; output prediction is performed according to the multi-scale feature map to obtain an output prediction result, and the target prediction result is loss evaluated through a pre-set comprehensive loss function to obtain a loss evaluation result; wherein the comprehensive loss function includes a classification loss function, a boundary regression loss function and a focal loss function; the parameters of the initial flood discharge detection model are updated according to the loss evaluation result to obtain a final flood discharge detection model.
[0160] The more detailed working principle and step flow of this embodiment can be but not limited to referring to the related description of embodiment one.
[0161] The embodiment of the present application provides a plurality of sample enhancement strategies through a sample acquisition module 001, balances sample distribution, increases sample quantity, solves the problem of small sample quantity and narrow coverage caused by low flood discharge event occurrence rate, and thus improves the generalization ability and robustness of a model; a model training module 002 is used for lightening the YOLOv8 model, reducing the calculation amount of the model, thus ensuring the real-time performance of flood discharge detection, and the YOLOv8 model is modified through a feature fusion module, the feature extraction ability of the model is strengthened, thus improving the detection precision, and finally, a second training sample set obtained through sample enhancement is used to train an initial flood discharge detection model obtained through structural modification, and thus a final flood discharge detection model with high precision and strong robustness is obtained.
[0162] Embodiment three:
[0163] The embodiment provides a terminal device, comprising a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete mutual communication through the communication bus.
[0164] The memory is used for storing at least one executable instruction, and the executable instruction makes the processor execute the operations of the flood discharge detection model training method in any one of the above embodiments.
[0165] Embodiment four:
[0166] The embodiment of the present application provides a computer readable storage medium, the computer readable storage medium comprises a stored computer program, wherein the computer program controls the device or apparatus where the computer readable storage medium is located to execute the flood discharge detection model training method in any one of the above embodiments when the computer program is running.
[0167] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0168] The above-mentioned specific embodiments further illustrate the purpose, technical solutions and advantages of the present application. It should be understood that the above-mentioned specific embodiments are only for the specific embodiments of the present application and do not limit the protection scope of the present application. It is particularly pointed out that any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for training a flood discharge detection model, characterized in that, include: The pre-collected video image data of the reservoir spillway is used as the first training sample set, and the first training sample set is augmented according to the preset sample augmentation strategy to obtain the second training sample set; wherein, the sample augmentation strategy includes one or more of the following: independent augmentation strategy, combined augmentation strategy or mosaic augmentation strategy. Based on the second training sample set, the pre-constructed initial flood discharge detection model is trained to obtain the final flood discharge detection model; wherein, the initial flood discharge detection model is an improved YOLOv8 model obtained by modifying the flood discharge convolution feature extraction module and updating the feature fusion module; The step of training a pre-constructed initial flood discharge detection model based on the second training sample set to obtain a final flood discharge detection model includes: inputting the second training sample set into the backbone network layer of the initial flood discharge detection model to extract features through a stacked flood discharge convolutional feature extraction module to generate an initial feature map set; wherein the flood discharge convolutional feature extraction module includes a CBS module, a lightweight modified C2f module, an SPPF module, and a CSAM hybrid attention mechanism; fusing the initial feature map set through a preset multi-scale feature fusion module to obtain a multi-scale feature map; wherein the multi-scale feature fusion module is obtained by parallel connection of dilated convolutional kernels and attention mechanisms; performing output prediction based on the multi-scale feature map to obtain an output prediction result, and evaluating the loss of the output prediction result through a preset comprehensive loss function to obtain a loss evaluation result; wherein the comprehensive loss function includes a classification loss function, a boundary regression loss function, and a focus loss function; updating the parameters of the initial flood discharge detection model based on the loss evaluation result to obtain the final flood discharge detection model; Specifically, the feature extraction through stacked flood discharge convolutional feature extraction modules to generate an initial feature map set involves: performing multiple ordinary convolutions and depthwise separable convolutions on the second training sample set sequentially using several layers of consecutive combined convolutional modules to obtain a first feature map set containing feature maps of different levels; wherein the combined convolutional modules include the CBS module and the C2f module after lightweight modification; performing multi-scale spatial pooling on the highest-level feature map in the first feature map set using the SPPF module to obtain a second feature map; enhancing the second feature map using the CSAM hybrid attention mechanism to obtain a third feature map; and merging the first feature map set, the second feature map, and the third feature map to obtain the initial feature map set.
2. The flood discharge detection model training method as described in claim 1, characterized in that, The independent enhancement strategies include one or more combinations of the following: noise injection strategy, color transformation strategy, geometric transformation strategy, or weather simulation strategy; The step of performing sample augmentation on the first training sample set according to a preset sample augmentation strategy to obtain a second training sample set includes: When the sample augmentation strategy is a noise injection strategy, a noise map of any noise type is superimposed on each original sample image in the first training sample set to obtain a second training sample set. When the sample enhancement strategy is a color transformation strategy, each original sample image in the first training sample set is color transformed according to any color transformation mode to obtain the second training sample set; wherein, the color transformation mode includes HSV transformation mode, mean blur mode, contrast adjustment mode, brightness adjustment mode or histogram equalization mode. When the sample augmentation strategy is a geometric transformation strategy, the original sample images in the first training sample set are subjected to geometric transformation according to any geometric transformation type to obtain the second training sample set; wherein, the geometric transformation type includes translation transformation, rotation transformation, scaling transformation, mirror flipping or image distortion. When the sample enhancement strategy is a weather simulation strategy, the second training sample set is obtained by performing weather simulation processing on each original sample image in the first training sample set through simulation technology.
3. The flood discharge detection model training method as described in claim 2, characterized in that, The weather simulation strategy includes a sunshine simulation strategy, a rain simulation strategy, a cloud and fog simulation strategy, or a snowflake simulation strategy. The second training sample set is obtained by performing weather simulation processing on each original sample image in the first training sample set using simulation technology, including: When the weather simulation strategy is a sunlight simulation strategy, the pre-generated halo map is linearly superimposed with each original sample image in the first training sample set through image fusion technology to obtain the second training sample set; wherein, the halo map is generated by performing attenuation simulation calculation of the light source component of randomly generated light source points in the corresponding original sample image through a mirror attenuation function. When the weather simulation strategy is a rain simulation strategy, the pre-generated rain pattern image is linearly superimposed with each original sample image in the first training sample set using image fusion technology to obtain the second training sample set; wherein, the rain pattern image is generated by color modulation and transparency control of the pre-generated rain line image; the rain line image is generated by drawing the length and tilt angle of the rain line with each raindrop as the starting point in the pre-created basic raindrop noise image; When the weather simulation strategy is a cloud and fog simulation strategy, fogging is applied to each original sample image in the first training sample set using a preset transmittance map to obtain a second training sample set. When the weather simulation strategy is a snowflake simulation strategy, the pre-generated final snowflake image is superimposed on each of the original sample images in the first training sample set using image fusion technology to obtain the second training sample set. The final snowflake image is generated by adjusting the transparency and controlling the brightness of several snowflakes in the pre-generated initial snowflake image. The initial snowflake image is obtained by generating several snowflakes on the corresponding original sample image through random point generation and masking operations.
4. The flood discharge detection model training method as described in claim 1, characterized in that, The step of performing sample augmentation on the first training sample set according to a preset sample augmentation strategy to obtain a second training sample set includes: When the sample augmentation strategy is a mosaic augmentation strategy, several original sample images are randomly selected from the first training sample set, and the original sample images are scaled to obtain scaled sample images, so as to stitch the scaled sample images into a new image with the same size as each original sample image. Continue to randomly select the same number of original sample images from the remaining original sample images in the first training sample set to stitch them together into a new image of the same size, until all original sample images have been stitched together; The newly stitched images are used as the second training sample set.
5. The flood discharge detection model training method as described in claim 1, characterized in that, The initial feature map set is fused using a preset multi-scale feature fusion module to obtain a multi-scale feature map, specifically as follows: By performing convolution operations on the initial feature map set using several convolutional blocks with different dilation rates, several first intermediate feature maps are obtained. By using a pre-defined hybrid attention mechanism, the first intermediate feature map is enhanced to obtain several second intermediate feature maps; The second intermediate feature map is then stitched together to obtain a multi-scale feature map.
6. A flood discharge detection model training device, characterized in that, It includes a sample acquisition module and a model training module, among which, The sample acquisition module is used to select a portion of pre-collected video image data of the reservoir spillway as a first training sample set, and perform sample enhancement on the first training sample set according to a preset sample enhancement strategy to obtain a second training sample set; wherein, the sample enhancement strategy includes independent enhancement strategy, combined enhancement strategy and mosaic enhancement strategy. The model training module is used to train the pre-constructed initial flood discharge detection model based on the second training sample set to obtain the final flood discharge detection model; wherein, the initial flood discharge detection model is an improved YOLOv8 model obtained by modifying the flood discharge convolution feature extraction module and updating the feature fusion module. The step of training a pre-constructed initial flood discharge detection model based on the second training sample set to obtain a final flood discharge detection model includes: inputting the second training sample set into the backbone network layer of the initial flood discharge detection model to extract features through a stacked flood discharge convolutional feature extraction module to generate an initial feature map set; wherein the flood discharge convolutional feature extraction module includes a CBS module, a lightweight modified C2f module, an SPPF module, and a CSAM hybrid attention mechanism; fusing the initial feature map set through a preset multi-scale feature fusion module to obtain a multi-scale feature map; wherein the multi-scale feature fusion module is obtained by parallel connection of dilated convolutional kernels and attention mechanisms; performing output prediction based on the multi-scale feature map to obtain an output prediction result, and evaluating the loss of the output prediction result through a preset comprehensive loss function to obtain a loss evaluation result; wherein the comprehensive loss function includes a classification loss function, a boundary regression loss function, and a focus loss function; updating the parameters of the initial flood discharge detection model based on the loss evaluation result to obtain the final flood discharge detection model; Specifically, the feature extraction through stacked flood discharge convolutional feature extraction modules to generate an initial feature map set involves: performing multiple ordinary convolutions and depthwise separable convolutions on the second training sample set sequentially using several layers of consecutive combined convolutional modules to obtain a first feature map set containing feature maps of different levels; wherein the combined convolutional modules include the CBS module and the C2f module after lightweight modification; performing multi-scale spatial pooling on the highest-level feature map in the first feature map set using the SPPF module to obtain a second feature map; enhancing the second feature map using the CSAM hybrid attention mechanism to obtain a third feature map; and merging the first feature map set, the second feature map, and the third feature map to obtain the initial feature map set.
7. A terminal device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the flood discharge detection model training method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device or apparatus containing the computer-readable storage medium to perform the flood discharge detection model training method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent dike crack dangerous case identification method based on edge enhancement
CN118865184A
Remote sensing target detection large model construction method based on context feature deep learning
CN120356096A