A method and system for monitoring the operation of a fish driving fence or fish collecting hopper based on image generation

Through the image generation method, combined with the YOLOv5s object detection model and style transfer model, the problem of fish catcher or fish trap recognition and motion state judgment in water conservancy projects is solved, automatic recognition and motion state judgment is realized, work flow is simplified and cost is reduced.

CN119339288BActive Publication Date: 2025-06-20HYDROPOWER WATER CONSERVANCY GUIHUA DESIGN ZONGYUAN +5
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411382919.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-06-20
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

In water conservancy projects, traditional image processing methods are difficult to adapt to the appearance changes of fish catches or fish collecting buckets built by different power stations, and deep learning networks are difficult to collect data in this application scenario, resulting in insufficient detection performance and high cost.

Method used

Using an image generation method, by obtaining video streams, drawing frames, and inputting pre-trained YOLOv5s object detection model, combining the style transfer model to generate a diversified data set, and training the object detection model to achieve automatic identification and motion state judgment of fish catcher or fish collecting bucket.

Benefits of technology

It realizes efficient and automatic identification of the operating status of fish catches or fish collecting buckets in hydropower station scenarios, simplifies the recording process of power station staff, reduces data collection costs, and avoids false operation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339288B_ABST
    Figure CN119339288B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for monitoring the operation of a fish driving fence or fish collecting hopper based on image generation. The method includes: obtaining a video stream containing the fish driving fence or fish collecting hopper; sequentially extracting images from the video stream according to a preset frame extraction frequency, and inputting the extracted images into a corresponding pre-trained object detection model to obtain corresponding object detection frames, wherein each of the object detection models is constructed based on a network including YOLOv5s; sequentially accumulating a preset number of object detection frame width sequences, and fitting the target motion trend according to the preset number of object detection frame width sequences to determine the current target motion state; determining the start time and end time of video recording according to the change of the target motion state to save the target motion video. In the embodiments of the present disclosure, the object detection frame detected in the monitoring video of the fish driving fence or fish collecting hopper is used to judge the target motion state, so as to determine to record and save the target motion video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of water conservancy and hydropower engineering, and particularly to a method, a system, an electronic device, and a computer-readable storage medium for monitoring the operation of a fish driving grid or a fish collecting hopper based on image generation. Background Art

[0002] With the continuous increase in the number and scale of domestic water conservancy projects, higher requirements have been put forward for the protection of fish ecology. The construction of a fish lift system can remedy the impact of human activities on fish migration channels, enabling fish to swim over the dam smoothly to specific waters for spawning, maintaining the normal reproduction of river fish, and protecting the continuity of the ecosystem. Using a fish driving grid to collect fish and a fish collecting hopper to lift fish are important steps in the fish passing process of the fish lift. In water conservancy projects, it is necessary to monitor the operating status of the fish driving grid and the fish collecting hopper.

[0003] Considering that the actual scenario is complex and changeable, and the appearances of fish driving grids or fish collecting hoppers built in different power stations are not exactly the same, etc., traditional image processing methods are not applicable. Although deep learning networks are widely used in image processing, for deep learning problems, data collection is a very important and costly issue. As a rare application scenario, it is difficult to obtain image data of fish driving grids or fish collecting hoppers in hydropower stations, and the sample data collected on site has the defect of insufficient diversity. Therefore, using a deep learning network for detection in the hydropower station application scenario to improve the detection performance while controlling the cost has become a new research trend. Summary of the Invention

[0004] The purpose of the embodiments of the present disclosure is to provide a method, a system, an electronic device, and a computer-readable storage medium for monitoring the operation of a fish driving grid or a fish collecting hopper based on image generation, so as to solve the foregoing problems existing in the prior art.

[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present disclosure are as follows:

[0006] On the one hand, the embodiments of the present disclosure propose a method for monitoring the operation of a fish driving grid or a fish collecting hopper based on image generation, and the method includes:

[0007] Obtain a video stream containing a fish driving grid or a fish collecting hopper;

[0008] Extract images from the video stream in sequence according to a preset frame extraction frequency, and input the extracted images into a corresponding pre-trained target detection model to obtain corresponding target detection frames; wherein, each of the target detection models is constructed based on the YOLOv5s network, and the training samples of each of the target detection models include: an initial image data set containing the target and different style image data sets generated by passing the initial image data set through a style transfer model, and the target includes a fish driving grid or a fish collecting hopper;

[0009] A preset number of target detection frame width sequences obtained by successive accumulation, and fitting the target motion trend according to the width sequences of the preset number to determine the current target motion state;

[0010] According to the change of the target motion state, determine the start time and end time of video recording, and save the recorded target motion video.

[0011] Optionally, the target detection model is obtained by the following training steps:

[0012] Obtain a number of images containing fish driving fences or fish collecting hoppers, and mark the fish driving fences or fish collecting hoppers in each image to construct a corresponding initial data set;

[0013] According to a variety of pre-selected images with different styles, use a style transfer model to perform style transformation on the corresponding initial data set respectively to obtain data sets corresponding to different styles;

[0014] Integrate the initial data set and the corresponding data sets with different styles, and perform data transformation and amplification on the integrated data set to obtain a corresponding amplified data set;

[0015] Train the corresponding YOLOv5s network with the corresponding amplified data set to obtain a corresponding trained target detection model.

[0016] Optionally, the loss function used in the process of training the target detection model includes: a localization loss function, a confidence loss function, and a classification loss function. Among them, both the confidence loss function and the classification loss function use BCEloss, and the localization loss function uses CIOU loss;

[0017] According to the corresponding weights of different loss functions, obtain the overall loss function of the target detection model until the overall loss function converges to obtain the target detection model.

[0018] Optionally, after using the style transfer model to perform style transformation on the corresponding initial data set, the method further includes:

[0019] Filter out the style images with a blocky feeling, and retain the style images with continuous backgrounds, prominent target of interest, and no deformation.

[0020] Optionally, before the preset number of target detection frame width sequences obtained by successive accumulation, the method further includes:

[0021] Judge whether the image is the first frame image of the video stream. If so, calculate and save the aspect ratio of the target detection frame in the first frame image;

[0022] Otherwise, calculate the aspect ratio of the target detection box in the current frame image, and calculate the ratio of the absolute value of the difference between the aspect ratio of the target detection box in the current frame image and the aspect ratio of the target detection box in the first frame image to the aspect ratio of the target detection box in the first frame image, and determine whether the ratio exceeds the first threshold. If so, the current frame target is a false detection and is skipped.

[0023] Optionally, the fitting the target motion trend according to a preset number of width sequences to determine the target motion state includes:

[0024] Use the least squares method to fit a preset number of target detection box width sequences to obtain the slope k of the fitting line;

[0025] If the slope k is greater than or equal to the first preset slope threshold θ, the width of the target detection box gradually increases, and the target moves closer to the target position;

[0026] If the slope k is less than or equal to the second preset slope threshold -θ, the width of the target detection box gradually decreases, and the target moves away from the target position;

[0027] If the slope k is less than the first preset threshold θ and greater than the second preset threshold -θ, the width of the target detection box remains unchanged, and the target is stationary.

[0028] Optionally, the determining the start time and end time of video recording according to the change of the target motion state includes:

[0029] When the target changes from a stationary state to moving closer to the target position, record the current time point as the target start motion time and start video recording;

[0030] When the target changes from moving closer to the target position to a stationary state and the ratio of the width of the target detection box in the current image frame to the width of the target detection box in the first frame is greater than the second preset threshold, or there is no target detection box in several consecutive images, record the current time point as the target end motion time point and end video recording.

[0031] Another aspect of the embodiments of the present disclosure provides a running monitoring system for a fish driving fence or fish collecting hopper based on image generation, and the system includes:

[0032] An acquisition module, configured to acquire a video stream including a fish driving fence or a fish collecting hopper;

[0033] A detection module, configured to sequentially extract images from a video stream according to a preset frame extraction frequency, input the extracted images into a corresponding pre-trained object detection model, and obtain corresponding object detection frames; wherein each of the object detection models is constructed based on the YOLOv5s network, and the training samples of each of the object detection models include: an initial image data set containing an object and different style image data sets generated by passing the initial image data set through a style transfer model, and the object includes a fish driving fence or a fish collecting hopper;

[0034] A judgment module, configured to sequentially accumulate a preset number of object detection frame width sequences, and fit the object motion trend according to the preset number of width sequences to judge the current object motion state;

[0035] A recording module, configured to determine the start time and end time of video recording according to the change of the object motion state, and save the recorded object motion video.

[0036] Another aspect of the embodiments of the present disclosure provides an electronic device, including: one or more processors;

[0037] A storage unit, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the method as described above.

[0038] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above can be implemented.

[0039] The beneficial effects of the embodiments of the present disclosure are:

[0040] A method and system for monitoring the operation of a fish driving fence or a fish collecting hopper based on image generation proposed by the embodiments of the present disclosure can automatically and efficiently identify the operation state of the fish driving fence or the fish collecting hopper based on vision, simplify the manual recording work process of power station staff, and save the corresponding video to avoid the problem of false operations. Description of the Drawings

[0041] Figure 1 It is a schematic flow chart of a method for monitoring the operation of a fish driving fence or a fish collecting hopper based on image generation according to an embodiment of the present disclosure;

[0042] Figure 2 It is a schematic flow chart of a method for monitoring the operation of a fish driving fence or a fish collecting hopper based on image generation including the training process of an object detection model according to an embodiment of the present disclosure;

[0043] Figure 3 It is a schematic flow chart of the training process of a style transfer model according to an embodiment of the present disclosure;

[0044] Figure 4 It is a schematic structural diagram of a fish driving fence or fish collecting hopper operation monitoring system based on image generation according to an embodiment of the present disclosure;

[0045] Figure 5 It is a schematic diagram of the target detection result of the fish driving fence according to an embodiment of the present disclosure. Among them, the fish driving fence is located at the farthest distance from the camera, and it will gradually approach the camera when it starts to move. Specific implementation manners

[0046] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to explain the embodiments of the present disclosure, and are not used to limit the embodiments of the present disclosure.

[0047] In order to automatically monitor the operation status of the fish driving fence or fish collecting hopper, standardize the operations of power station staff, save the corresponding video segments as supervision evidence, or count the operation times of the fish lift system, the embodiments of the present disclosure propose a fish driving fence or fish collecting hopper operation monitoring method and system based on image generation.

[0048] As Figure 1 shown, on the one hand, the embodiments of the present disclosure propose a fish driving fence or fish collecting hopper operation monitoring method based on image generation, and the method includes:

[0049] Step S100, obtaining a video stream including a fish driving fence or a fish collecting hopper.

[0050] In the embodiments of the present disclosure, the fish driving fence and the fish collecting hopper are both part of the fish transportation system. First, the fish driving fence moves horizontally to transport the fish school to the target position, and then, the fish collecting hopper vertically lifts the fish school. However, the fish transportation system is relatively large, and a single camera can only clearly capture the movement process of one of the fish driving fence or the fish collecting hopper. Therefore, cameras are installed near the fish driving fence or the fish collecting hopper respectively, and the shooting range of the corresponding camera covers the movement range of the corresponding fish driving fence or fish collecting hopper. The subsequent monitoring method steps for the operation of the two targets are the same. Among them, the target detection model for the fish driving fence or the fish collecting hopper in step S200 is also trained using the same method, and the difference lies only in the different parameter weights of the target detection model. Therefore, the method of the embodiments of the present disclosure can be used to monitor the fish driving fence or the fish collecting hopper. The fish driving fence and the fish collecting hopper need to install the cameras at appropriate positions and fix them respectively to capture the target movement process, ensuring that the target of the fish driving fence or the fish collecting hopper is directly facing the camera, so that the change in the size of the target detection frame can reflect the movement change of the target. The camera of the fish driving fence is set directly in front of it, and the camera of the fish collecting hopper is set directly above it. For the fish driving fence, the camera is installed directly in front of the target position and can capture the entire operation process of the fish driving fence. Among them, the front refers to the direction of transporting the fish school to the target position, and the rear refers to the direction away from the target position; for the fish collecting hopper, the camera is installed directly above it and can capture the entire operation process of the fish collecting hopper. When the fish driving fence or the fish collecting hopper starts to operate, it will approach the camera. The camera for the fish driving fence cannot capture the fish collecting hopper, and the camera for the fish collecting hopper cannot capture the fish driving fence either, and there is only one target detection frame in each frame of the image. The appearance of multiple ones belongs to false detection.

[0051] Step S200: Sequentially extract the images in the video stream according to a preset frame extraction frequency, and input the extracted images into the corresponding pre-trained target detection model to obtain the corresponding target detection frames. Among them, each of the target detection models is constructed based on the YOLOv5s network, and each training sample of the target detection model includes: an initial image data set containing the target and different style image data sets generated by passing the initial image data set through a style transfer model, and the target includes the fish driving fence or the fish collecting hopper.

[0052] The embodiments of the present disclosure process the images in the video stream in real time, use the YOLOv5s network to detect the target, obtain the target detection frame and coordinates. The coordinates of the target detection frame are used to calculate the height, width, etc. of the target detection frame, save the aspect ratio of the width and height of the target detection frame in the first frame image of the video, and calculate the width w of the target detection frame at the current moment according to the target detection frame. t Perform frame-by-frame processing on the video stream, including performing target detection, screening false detections, recording the width of the target detection frame, etc. on each frame of the image, for judging the target movement state.

[0053] In the embodiments of the present disclosure, the target boxes of the fish driving fence or fish collecting hopper in each frame of the video are obtained in real time. The specific process includes inputting the real-time video into the target detection model to obtain the target boxes of interest, such as Figure 5 as shown; at the same time, some misdetected targets are filtered according to the aspect ratio of the target detection box in the first frame image.

[0054] Exemplarily, before inputting the extracted image into the corresponding pre-trained target detection model, the method further includes: preprocessing the image to obtain an image with a long side of 640 pixels.

[0055] It should be noted that, in order to ensure real-time detection and implementation at the edge, the embodiments of the present disclosure adopt a one-stage target detector based on YOLOv5s for skip-frame detection. For example, detection is performed every 5 frames. The image frames to be detected are scaled proportionally to a long side of 640 pixels and input into the backbone feature extraction network Backbone. Then, three feature maps of different sizes are output through the Neck layer network, and then the 1×1 convolution is used to predict the results in the Head layer.

[0056] Exemplarily, the target detection model is obtained by the following training steps:

[0057] Step S210: Obtain a number of images containing the fish driving fence or fish collecting hopper, and mark the fish driving fence or fish collecting hopper in each image to construct the corresponding initial dataset.

[0058] In the embodiments of the present disclosure, in step S210, images containing the fish driving fence or fish collecting hopper in the actual scenes of different hydropower stations are collected, and the fish driving fence or fish collecting hopper in the images are marked respectively to form the corresponding initial dataset. The initial dataset contains images of the fish driving fence or fish collecting hopper from multiple angles and multiple distances. The size of the fish driving fence or fish collecting hopper will be different at different distances from the camera.

[0059] Step S220: According to a variety of pre-selected images with different styles, use the style transfer model to perform style transformation on the corresponding initial dataset respectively to obtain datasets corresponding to different styles.

[0060] In the embodiments of the present disclosure, in step S220, a variety of target style images are selected, and the style transfer model is used to convert the real images containing the fish driving fence or fish collecting hopper obtained in step S210 into corresponding style images respectively. In step S220, 12 target style images containing different textures, shapes, colors and distribution characteristics are selected, and 12 different style datasets are generated on the basis of the initial dataset with the help of the style transfer model and screened manually. The screening process includes: retaining the style images with continuous background, prominent target of interest and no large deformation, and filtering out the style images with obvious blockiness.

[0061] In the embodiments of the present disclosure, a style transfer model is used to convert real image data into style data, and its corresponding network architecture is as Figure 3 shown. The style transfer model consists of an image generation network and a loss network. The image generation network is essentially an autoencoder network, which consists of a pair of downsampling and upsampling modules and five residual modules. Through the image generation network, the style of the input image can be changed to the expected style on the basis of keeping the high-level semantic information of the input image unchanged. The loss network is essentially a feature extraction network, which is composed of a pre-trained VGG-16 network and is used to measure the differences in content and style between the generated style image and the input content image and style image. In order to change the style of the input image on the premise of ensuring that the image content remains unchanged, the style transfer model uses a feature reconstruction loss and a style reconstruction loss to constrain the generated style image. The original style transfer model includes three loss functions, namely, a feature reconstruction loss, a style reconstruction loss, and a total variation regularization loss. Among them, the purpose of the total variation regularization loss is to make the generated style image smoother. The style transfer model of the embodiments of the present disclosure only uses the first two to constrain the generated image, making the foreground and background more distinguishable, and training until the weighted sum of the two converges.

[0062] Feature reconstruction loss: The shallow feature maps of the loss network often contain pixel point information similar to the input, including color, texture, etc., and the deep feature maps often contain abstract semantic information, such as spatial structure information. The feature reconstruction loss constrains the deep features of the loss network, so that the generated image changes the color and texture information on the premise of retaining the original shape and contour information. Let x represent the input content image, y represent the input style image, represent the generated image, C j 、H j 、W j respectively represent the number of channels and the height and width of the feature variable, φ j (·) represents the j-th layer feature of the loss network, then the feature reconstruction loss function can be expressed as:

[0063]

[0064] Style reconstruction loss: The Gram matrix is the inner product operation of matrices and can be used to measure the correlation between matrices. In the process of style transfer, the embodiments of the present disclosure use the Gram matrix to measure the correlation between the generated image and the input style image at different feature layers. If the distances between them at different feature layers are close, it means that they have similar color and texture information, that is, similar styles. The Gram matrix can be expressed as:

[0065]

[0066] Among them, c represents all channels of the features of the j-th layer of the loss network, and c' represents the transpose of c. By adjusting φ j (x) to the form of C j ×H j W j , the Gram matrix can be calculated efficiently. At this time, G j φ (x) = ψψ T / C j H j W j . The style reconstruction loss reflects the difference in the Gram matrices between the generated image and the input style image. Minimizing the style reconstruction loss can make the style of the generated image transform towards the expected style. The style reconstruction loss function can be expressed as:

[0067]

[0068] Among them, ψ represents a matrix of shape C j ×H j W j . F represents the Frobenius norm of the matrix, that is, the square root of the sum of the squares of each element of the matrix.

[0069] Combining the feature reconstruction loss and the style reconstruction loss, the overall loss function of the style transfer model can be expressed as:

[0070]

[0071] Among them, λ C and λ S represent the weights of the feature reconstruction loss and the style reconstruction loss respectively. J represents the style feature layer, and j ∈ J. In this embodiment, λ C is set to 1, and λ S is set to 100. j only refers to a single layer, and J represents a set of multi-layer networks.

[0072] Step S230: Integrate the initial data set and the data sets with corresponding different styles, and perform data transformation and augmentation on the integrated data set to obtain a corresponding augmented data set.

[0073] In step S230 of the present disclosure embodiment, the style data set screened in step S220 and the corresponding initial data set obtained in step S210 are integrated to form a training set for the corresponding object detection network. Step S230 can also perform further data transformation and augmentation on the integrated data set to obtain a corresponding augmented data set, so as to train a fishing fence or fish collecting bucket object detection model based on the YOLOv5s network framework. Train a fishing fence or fish collecting bucket object detection network based on the YOLOv5s framework. The data transformation methods include: any one or more of mosaic, scale scaling, random flipping, and translation, etc.

[0074] Step S240: Train the corresponding YOLOv5s network with the corresponding amplified dataset to obtain the corresponding trained object detection model.

[0075] In the embodiment of the present disclosure, the diversity of the dataset is improved in steps S210 to S230, preventing the object detection model from learning irrelevant features, so that the trained object detection model has stronger generalization ability and improves the overall performance. After comprehensively considering the detection accuracy and detection speed, the YOLOv5s network is selected as the object detection network.

[0076] Exemplarily, the loss function used in the process of training the object detection model includes: a localization loss function, a confidence loss function, and a classification loss function. Among them, both the confidence loss function and the classification loss function use BCEloss, and the localization loss function uses CIOU loss;

[0077] According to the corresponding weights of different loss functions, the overall loss function of the object detection model is obtained until the overall loss function converges, and the object detection model is obtained.

[0078] The overall loss function of the object detection model in the embodiment of the present disclosure comprehensively considers three factors: the overlapping area of the rectangular frame, the distance between the center points, and the aspect ratio, improving the stability and convergence speed of the training of the object detection model.

[0079] Step S300: Cumulatively obtain the width sequences of a preset number of object detection frames in sequence, and fit the target motion trend according to the width sequences of the preset number of object detection frames to determine the current target motion state.

[0080] In the embodiment of the present disclosure, the width of the object detection frame is calculated based on the coordinates of the object detection frame. In the embodiment of the present disclosure, the width sequence of the object detection frame for a preset period of time can be obtained to fit the target motion trend and determine the target motion state. It can also be detected once every 5 frames, and the detection results of 24 detections are fitted to determine the target motion trend. For a camera with a frame rate of 30, the preset time is about 4 seconds. If the change in the width of the object detection frame is not obvious, it is caused by the normal jitter of the detection results of different frames.

[0081] For the detection results of every 24 frames, the slope method is used to determine the current motion state of the target. The specific process is as follows: The least squares method is used to fit the width data of 24 target detection frames, and the data trend of the sequence is judged according to the slope k of the fitted straight line. In the embodiments of the present disclosure, the motion state of the target can be obtained according to the change in the width of the target detection frame. The change in width and the change in motion state are related to the installation position of the camera. Therefore, in the embodiments of the present disclosure, the camera is installed directly in front of the fish driving fence or directly above the fish collecting hopper, and a clearer working state video of the fish driving fence or the fish collecting hopper can be obtained. The embodiments of the present disclosure detect the target in the video stream in real time and judge its motion state. For example, the video stream is framed according to the preset frame extraction frequency, and each frame is detected. The motion state of the target is judged once for the accumulated detection results of every 24 frames in turn. The first motion state judgment is for frames 1-24, the second motion state judgment is for frames 2-25, and so on. Whether to start recording the video is determined according to the judgment result of the change in the motion state of the target.

[0082] As Figure 2 shown, exemplarily, before successively accumulating to obtain a preset number of target detection frame sequences, the method further includes:

[0083] Judging whether the image is the first frame image of the video stream. If so, calculating and saving the aspect ratio of the target detection frame in the first frame image;

[0084] If not, calculating the aspect ratio of the target detection frame of the current frame image, calculating the ratio of the absolute value of the difference between the aspect ratio of the target detection frame of the current frame image and the aspect ratio of the target detection in the first frame image to the aspect ratio of the target detection in the first frame image, and judging whether the ratio exceeds the first threshold. If so, the target of the current frame image is a false detection and is skipped; if not, the target detection frames of the current frame image are accumulated.

[0085] Generally, it is defaulted that the target has not started to move at this time in the first frame of the video stream and is located at the farthest distance from the camera. Because the camera can capture the complete motion process of the target, theoretically there is a target in each frame image. When the ratio of the absolute value of the difference between the aspect ratio r1 of the target detection frame in the current frame and the aspect ratio r0 of the target detection in the first frame to the aspect ratio of the target detection in the first frame image exceeds the first preset threshold, the aspect ratio of the current target and the first frame target is quite different, and the target detection frame in the current frame is determined to be a false detection target. In the embodiments of the present disclosure, the first preset threshold is about 0.3, that is, when |r1 - r0| / r0 ≥ 0.3, it is determined as a false detection target, and the images of the false detection targets are filtered until the target detection frames of the preset number of images are detected.

[0086] Exemplarily, the fitting of the target motion trend according to the width sequences of the preset number of target detection frames to judge the target motion state includes:

[0087] Using the least squares method to fit the width sequences of a preset number of target detection frames to obtain the slope k of the fitted straight line;

[0088] If the slope k is greater than or equal to the first preset slope threshold θ, the width of the target detection frame gradually increases, and the target moves closer to the target position;

[0089] If the slope k is less than or equal to the second preset slope threshold -θ, the width of the target detection frame gradually decreases, and the target moves away from the target position;

[0090] If the slope k is less than the first preset slope threshold θ and greater than the second preset slope threshold -θ, the width of the target detection frame remains basically unchanged, and the target is stationary.

[0091] Specifically, in the embodiments of the present disclosure, the slope method is used to fit the width sequences of a preset number m of frames of target detection frames. m can be 24 frames. The best linear function matching of m data is found by the sum of the squares of the minimum errors, that is, according to the given discrete data, the best slope and intercept are calculated. This embodiment mainly considers the slope calculation, and the specific formula is as follows:

[0092]

[0093] Among them, p represents the data serial number, q represents the width of each frame of the target box, represents the average value of the target widths of m frames, represents the average value of the data serial numbers, p i represents the i-th data serial number. In the embodiments of the present disclosure, m takes the value of 24, and p i takes values from 1 to 24.

[0094] When the slope k ≥ the first preset slope threshold θ and θ > 0, it is determined that the target moves towards the target position, that is, the target moves closer to the camera; when the slope k ≤ the second preset slope threshold -θ, it is determined that the target moves away from the target position, that is, the target moves in the direction away from the camera; when the second preset slope threshold -θ < the slope k < the first preset slope threshold θ, it is determined that the target is in a stationary state.

[0095] Step S400: Determine the start time for saving the video and the end time for recording the video according to the change of the adjacent target motion state, so as to save the recorded target motion video.

[0096] That is to say, the closer the fish driving fence or fish collecting hopper transports the fish school to the target position, the closer it will be to the camera. At this time, the video is recorded and saved. When the fish driving fence or fish collecting hopper is reset, the video recording is not performed. Specifically, when the target of interest, that is, the fish driving fence or fish collecting hopper, changes from a stationary state to a moving state approaching, it is marked as the start time of the target running, and the video writing starts at the same time; when the target of interest, that is, the fish driving fence or fish collecting hopper, changes from the starting approaching state to a stationary state, it is marked as the end time of the target movement, and at this time the video writing ends.

[0097] The present disclosure proposes a method for monitoring the operation of a fish driving fence or a fish collecting hopper based on image generation. Based on intelligent recognition technology and information means, it simplifies the process of power station operators manually recording the operation time of the fish lift system, avoids the possibility of operation false alarms, and at the same time saves the video of the target operation time period, effectively improving the efficiency of subsequent review and analysis; based on image generation technology, it expands the original data set, reduces the data collection cost, enhances the generalization ability of the target detection model, and improves the overall performance.

[0098] Exemplarily, determining the start time and end time of video recording according to the change of the adjacent target motion state includes:

[0099] When the target changes from a stationary state to a moving state approaching the target position, record the current time point as the start time of the target movement and start recording the video;

[0100] When the target changes from a moving state approaching the target position to a stationary state and the ratio of the width of the target detection frame in the current image frame to the width of the first-frame target detection frame is greater than the second preset threshold, or the target cannot be detected in several consecutive image frames, that is, the target partially or completely exceeds the image field of view, record the current time point as the end time point of the target movement and end recording the video.

[0101] Considering that the actual installation position of the camera may be difficult to capture the complete motion trajectory of the target of interest, the embodiment of the present disclosure also marks the time when the target of interest partially or completely exceeds the visible range of the camera as the end time of the movement. In order to increase the robustness of the algorithm, the embodiment of the present disclosure records the aspect ratio of the fish driving fence or fish collecting hopper target in the first frame of the video as a constraint condition to filter out some misdetected targets; in order to avoid misjudgment of the motion state caused by the jitter of the detection frame in the video, the embodiment of the present disclosure also sets the change of the target box width as another judgment condition for the end time of the movement.

[0102] It should be noted that, in order to avoid the influence of the detection frame jitter on the target motion state, the embodiment of the present disclosure also sets the growth ratio of the target box width as the judgment condition for the end node of the video recording, that is, when the target motion state changes and the ratio of the target width in the current image frame to the first-frame target width is greater than the set threshold, the video recording will end.

[0103] In the embodiments of the present disclosure, a more general deep learning object detection method is adopted to realize the video saving of the start and end operation times of the fish driving fence and the fish collecting hopper during the operation of the fish lift system, which is beneficial for subsequent reference and analysis.

[0104] As Figure 4 shown, on the other hand, the embodiments of the present disclosure provide a fish driving fence or fish collecting hopper operation monitoring system based on image generation, and the system includes:

[0105] An acquisition module 100, configured to acquire a video stream including a fish driving fence or a fish collecting hopper;

[0106] A detection module 200, configured to sequentially extract images from the video stream according to a preset frame extraction frequency, and input the extracted images into a corresponding pre-trained object detection model to obtain corresponding object detection frames; wherein, each of the object detection models is constructed based on the YOLOv5s network, and the training samples of each of the object detection models include: an initial image data set including the object and different style image data sets generated by passing the initial image data set through a style transfer model, and the object includes a fish driving fence or a fish collecting hopper;

[0107] A judgment module 300, configured to sequentially accumulate the width sequences of a preset number of obtained object detection frames, and fit the target motion trend according to the width sequences of the preset number of object detection frames to judge the current target motion state;

[0108] A recording module 400, configured to determine the start time and end time of video recording according to the change of the target motion state, so as to save the recorded target motion video.

[0109] In the embodiments of the present disclosure, data augmentation is mainly proposed based on a style transfer model, the position information of the fish driving fence or fish collecting hopper target is monitored in real time by using a deep learning object detection network, and the operation state is judged according to the result, and whether to record is determined according to the target operation state. The object detection model (the weight of the trained detection model) is exported in rknn format and deployed to the RK3588 platform, and the embedded neural network processor NPU is called for acceleration to realize real-time processing at the edge. The RK3588 platform is an edge computing platform, and the above inference process is carried out on this platform. The edge side refers to installing an AI industrial host at a suitable position in the hydropower station, which is equipped with an RK3588 processor, and the camera video stream is directly transmitted to this edge computing platform to complete processes such as detection, inference, and video saving.

[0110] The system further includes: a misdetection judgment module, configured to judge whether the image is the first frame image of the video stream, and if so, calculate and save the aspect ratio of the object detection frame in the first frame image;

[0111] Otherwise, calculate the aspect ratio of the target detection bounding box in the current frame image, and calculate the ratio of the absolute value of the difference between the aspect ratio of the target detection bounding box in the current frame image and the aspect ratio of the target detection bounding box in the first frame image to the aspect ratio of the target detection bounding box in the first frame image. Determine whether the ratio exceeds the first threshold. If so, the target in the current frame image is a false detection and is skipped.

[0112] Another aspect of the embodiments of the present disclosure provides an electronic device, including: one or more processors;

[0113] A storage unit for storing one or more programs, which when executed by the one or more processors, can enable the one or more processors to implement the method as described above.

[0114] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the method as described above.

[0115] The method of the embodiments of the present disclosure first collects image data of fish driving fences or fish collecting hoppers in the actual scenarios of hydropower stations. Considering problems such as few training samples, many actual application scenarios, and differences in fish driving fences or fish collecting hoppers in different hydropower stations, the present invention uses a style transfer model to generate diverse data, screens and labels them to expand the training set; trains the target detection network YOLOv5s according to the synthetic data set to achieve the target detection of fish driving fences or fish collecting hoppers; performs frame-by-frame detection on the video stream, uses the slope method to fit the change trend of the width of the target detection bounding box in the image frame, and sets a threshold to judge the target motion state; records the start and end times of the target running according to the change of the motion state of the fish driving fence or fish collecting hopper, and saves the corresponding video segment for subsequent reference.

[0116] The above are only the preferred embodiments of the embodiments of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present disclosure, several improvements and refinements can be made, and these improvements and refinements should also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A method for monitoring the operation of a fish-driving fence or a fish-collecting bucket based on image generation, characterized in that: The method comprises: Get a video stream containing a fish trap or fish collecting bucket; Extracting images from the video stream in sequence according to a preset frame extraction frequency, inputting the extracted images into a corresponding pre-trained target detection model to obtain a corresponding target detection frame, wherein each of the target detection models is constructed based on a YOLOv5s network, and the training samples of each of the target detection models include: an initial image data set containing a target and a different style image data set generated by the initial image data set through a style transfer model, wherein the target includes a fish trap or a fish collecting bucket; A preset number of target detection frame width sequences are accumulated in sequence, and a target motion trend is fitted according to the preset number of width sequences to determine the current target motion state; The step of fitting the target motion trend according to a preset number of width sequences to determine the target motion state includes: The least square method is used to fit the preset number of target detection box width sequences to obtain the slope k of the fitting line; If the slope k is greater than or equal to a first preset slope threshold θ, the width of the target detection frame gradually increases, and the target moves closer to the target position; If the slope k is less than or equal to the second preset slope threshold -θ, the width of the target detection frame gradually decreases, and the target moves away from the target position; If the slope k is less than the first preset slope threshold θ and greater than the second preset slope threshold -θ, the width of the target detection frame remains unchanged and the target is stationary; According to the change of the target motion state, determine the time to start recording the video and the time to end recording the video, and save the recorded target motion video; The step of determining the time to start recording the video and the time to end recording the video according to the change of the target motion state includes: When the target changes from a stationary state to moving towards the target position, the current time point is recorded as the target start movement time, and video recording begins; When the target changes from moving towards the target position to a stationary state and the ratio of the target detection frame width in the current image frame to the target detection frame width in the first frame is greater than a second preset threshold, or there is no target detection frame in several consecutive images, the current time point is recorded as the time point when the target ends its movement, and the video recording ends.

2. The method according to claim 1, characterized in that The target detection model is obtained using the following training steps: Obtain several images containing fish-driving fences or fish-collecting buckets, and mark the fish-driving fences or fish-collecting buckets in each image to construct a corresponding initial data set; According to the pre-selected images of different styles, the style transfer model is used to transform the styles of the corresponding initial data sets to obtain data sets corresponding to different styles; Integrate the initial data set and corresponding data sets of different styles, and perform data transformation and amplification on the integrated data set to obtain a corresponding amplified data set; The corresponding YOLOv5s network to be trained is trained with the corresponding amplified data set to obtain the corresponding trained target detection model.

3. The method according to claim 2, characterized in that The loss functions used in the process of training the target detection model include: positioning loss function, confidence loss function and classification loss function, wherein the confidence loss function and the classification loss function both use BCEloss, and the positioning loss function uses CIOUloss; According to the corresponding weights of different loss functions, the overall loss function of the target detection model is obtained until the overall loss function converges to obtain the target detection model.

4. The method according to claim 2, characterized in that: After the style transfer model is used to perform style transformation on the corresponding initial data set, the method further includes: Filter out blocky style images and retain style images with continuous background, prominent objects of interest and no deformation.

5. The method according to any one of claims 1 to 4, characterized in that: Before the preset number of target detection frame width sequences accumulated in sequence, the method further includes: Determine whether the image is the first frame of the video stream. If so, calculate the aspect ratio of the target detection frame of the first frame and save it. If not, calculate the aspect ratio of the target detection frame of the current frame image to calculate the ratio of the absolute value of the difference between the aspect ratio of the target detection frame of the current frame image and the aspect ratio of the target detection in the first frame image to the aspect ratio of the target detection in the first frame image, and determine whether the ratio exceeds the first threshold. If so, the target of the current frame image is a false detection and is skipped.

6. A fish-driving fence or fish-collecting bucket operation monitoring system based on image generation, characterized in that: The system comprises: An acquisition module, used for acquiring a video stream containing a fish-driving fence or a fish-collecting bucket; A detection module is used to extract images from a video stream in sequence according to a preset frame extraction frequency, and input the extracted images into a corresponding pre-trained target detection model to obtain a corresponding target detection frame; wherein each of the target detection models is constructed based on a YOLOv5s network, and the training samples of each of the target detection models include: an initial image data set containing a target and a different style image data set generated by the initial image data set through a style transfer model, wherein the target includes a fish trap or a fish collecting bucket; A judgment module is used to sequentially accumulate a preset number of target detection frame width sequences, and fit the target motion trend according to the preset number of width sequences to judge the current target motion state; The step of fitting the target motion trend according to a preset number of width sequences to determine the target motion state includes: The least square method is used to fit the preset number of target detection box width sequences to obtain the slope k of the fitting line; If the slope k is greater than or equal to a first preset slope threshold θ, the width of the target detection frame gradually increases, and the target moves closer to the target position; If the slope k is less than or equal to the second preset slope threshold -θ, the width of the target detection frame gradually decreases, and the target moves away from the target position; If the slope k is less than the first preset slope threshold θ and greater than the second preset slope threshold -θ, the width of the target detection frame remains unchanged and the target is stationary; The recording module is used to determine the time to start and end video recording according to the change of the target motion state, and save the recorded target motion video; The step of determining the time to start recording the video and the time to end recording the video according to the change of the target motion state includes: When the target changes from a stationary state to moving towards the target position, the current time point is recorded as the target start movement time, and video recording begins; When the target changes from moving towards the target position to a stationary state and the ratio of the target detection frame width in the current image frame to the target detection frame width in the first frame is greater than a second preset threshold, or there is no target detection frame in several consecutive images, the current time point is recorded as the time point when the target ends its movement, and the video recording ends.

7. An electronic device, characterized in that: include: one or more processors; A storage unit, used to store one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the method described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Behavior recognition method and system based on monitoring video

    CN115620212A

  • Water depth adjustable type fish gathering platform and control system

    CN117385801A