A method, system, medium and device for calculating the proportion of river ice cover area

Through the segmentation model, the first frame of the ice surface video is segmented and passed to the video object segmentation model for tracking mask detection of subsequent frames, which solves the problem of manual labeling in the prior art, improves the efficiency of calculating the proportion of river ice area, and realizes an accurate estimation of the river ice area.

CN119228871BActive Publication Date: 2025-05-09齐鲁空天信息研究院 +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411280511.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-05-09
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to automatically calculate the proportion of river ice-sealed area, especially when the river is long and the ice surface morphology is different, it is necessary to manually mark the first frame of the image, resulting in low computing efficiency.

Method used

The segmentation model is used to segment the first frame of the ice surface video with the river and the floating ice mask, and it is passed to the video object segmentation model for tracking mask detection of subsequent frames, avoiding manual annotation, and adding an attention mechanism to the segmentation model to adapt to small object detection from the perspective of the drone.

Benefits of technology

The efficiency of calculating the proportion of river ice-sealed area is improved, and the accurate estimate of the river ice-sealed area is achieved, which is suitable for ice-surface video segmentation tasks from the perspective of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228871B_ABST
    Figure CN119228871B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image processing, and discloses a method, system, medium and device for calculating the proportion of ice-covered area of ​​a river, comprising: dividing an ice surface video into a plurality of video segments, and extracting the first frame of each video segment; for each video segment, using a segmentation model to segment the first frame to obtain a floating ice mask and a river mask, using the floating ice mask and the river mask of the first frame as annotations, and using a video object segmentation model to segment all frames in the video segment to obtain a floating ice mask and a river mask of each frame; based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame, calculating an estimated interval of the proportion of ice-covered area; and adding an attention mechanism to the segmentation model to make it more suitable for small target object detection and segmentation under the perspective of a drone, and finally achieving an accurate video segmentation task, thereby accurately estimating the ice-covered area of ​​the river.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, system, medium and equipment for calculating the proportion of ice-covered area of ​​a river. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Cold weather in winter often freezes rivers, which affects tasks such as river navigation. According to the hydrological intelligent monitoring system, the distribution of floating ice on the river can be captured in real time, so as to determine the degree of ice cover in the river. According to the river image data at different times, the duration of ice cover can be estimated, and the long-term data can be used to determine seasonal changes, ice thickness, thawing trends and even climate changes. Severe river ice cover will affect transportation and infrastructure construction. Monitoring of river ice cover can quickly guide whether the river is navigable, and reasonably arrange ice-breaking plans and avoid various possible dangers and economic property losses. In addition, river ice cover may also affect the surrounding ecological environment and the inhabiting plants and animals; it may also affect the industrial and agricultural use of water resources. Therefore, intelligent monitoring of river ice cover is a current research focus.

[0004] In water conservancy systems, manual inspections may not be able to detect hydrological changes in a timely manner, and the inspection cycle is long and the inspection frequency is limited.

[0005] Artificial Intelligence (AI) is a branch of computer science that focuses on creating systems and machines that can perform tasks that normally require human intelligence. These tasks include learning, reasoning, problem solving, perception, language understanding, and decision making. AI systems are designed to mimic human cognitive functions and adapt to new situations, enabling them to perform tasks autonomously and efficiently.

[0006] The task of video object segmentation (VOS) is to identify and track objects from continuous video frames, segment the objects of interest completely and continuously, and calculate the proportion of river ice cover area.

[0007] AOT (Associating Objects with Transformers) is a commonly used model based on the Transformer (sequence-to-sequence model based on multi-head self-attention) architecture in the field of video object segmentation. It can effectively segment and track multiple objects in video frames. By combining the initial object annotation of the first frame with the self-attention mechanism, the model can effectively track and segment multiple objects over a period of time.

[0008] However, the AOT model requires the first frame of the image to mark the objects that need to be segmented and tracked, and cannot be directly and automatically applied to video data, resulting in the AOT model being insensitive to the appearance of new objects. However, due to the long river, the shapes of floating ice in different areas are different. If it is a long ice surface video, the first frame of the AOT model needs to be re-annotated after a period of time. Summary of the invention

[0009] In order to solve the above problems, the present invention provides a method, system, medium and equipment for calculating the proportion of ice-covered area of ​​a river. A segmentation model is used to segment the river and the floating ice mask in the first frame of the ice surface video, and the segmentation model is passed to the video object segmentation model. The video object segmentation model gives a tracking mask for subsequent frames, thereby avoiding manual labeling of the first frame of the ice surface image and improving the efficiency of calculating the proportion of ice-covered area of ​​the river. Moreover, an attention mechanism is added to the segmentation model to make it more suitable for small target object detection and segmentation from the perspective of a drone, thereby finally achieving accurate video segmentation tasks, thereby accurately estimating the ice-covered area of ​​the river.

[0010] In order to achieve the above object, the present invention adopts the following technical solution:

[0011] A first aspect of the present invention provides a method for calculating the proportion of river ice cover area, comprising:

[0012] Get video of the ice surface;

[0013] Divide the ice surface video into several video segments and extract the first frame of each video segment;

[0014] For each video segment, the segmentation model is used to segment the first frame to obtain the floating ice mask and the river mask. The floating ice mask and the river mask of the first frame are used as annotations. The video object segmentation model is used to segment all frames in the video segment to obtain the floating ice mask and the river mask of each frame.

[0015] Based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame, the estimated interval of the ice cover area ratio is calculated;

[0016] Among them, the segmentation model uses the channel attention mechanism to process the feature map of the first frame extracted by the deep learning network. After obtaining the feature map with channel attention, parallel void convolution layers with different expansion rates are used to capture multi-scale contextual information from the feature map with channel attention. After the multi-scale contextual information is fused, a fused feature map is obtained. The floating ice mask and river mask are detected based on the fused feature map.

[0017] Furthermore, the segmentation model fuses the fused feature maps of different sizes using a feature pyramid, uses a detection branch to predict the detection frame and the mask coefficient, uses a segmentation branch to predict the prototype mask, linearly combines the mask coefficient with the prototype mask, and obtains the floating ice mask and the river mask.

[0018] Furthermore, the calculation step of the ice-covered area ratio estimation interval includes:

[0019] For each frame, the ratio of the pixel area of ​​the floating ice mask to the pixel area of ​​the river mask is used as the estimated value of the ice-covered area ratio;

[0020] Calculate the estimated value of the ice-covered area percentage corresponding to all frames, calculate the standard error, and combine it with the given confidence level to calculate the estimated interval of the ice-covered area percentage.

[0021] Furthermore, the ice surface video includes visible light video and infrared video;

[0022] For each frame, the confidence levels of the estimated values ​​of the percentage of ice-covered area obtained from the two videos are compared, and the estimated value of the percentage of ice-covered area corresponding to the larger confidence level is selected.

[0023] A second aspect of the present invention provides a river ice cover area ratio calculation system, comprising:

[0024] A data acquisition module is configured to: acquire ice surface video;

[0025] A video segmentation module, which is configured to: divide the ice surface video into a number of video segments, and extract the first frame of each video segment;

[0026] The mask detection module is configured to: for each video segment, segment the first frame using the segmentation model to obtain a floating ice mask and a river mask, use the floating ice mask and the river mask of the first frame as annotations, and segment all frames in the video segment using the video object segmentation model to obtain a floating ice mask and a river mask of each frame;

[0027] A proportion calculation module is configured to: calculate an estimated interval of ice-covered area proportion based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame;

[0028] Among them, the segmentation model uses the channel attention mechanism to process the feature map of the first frame extracted by the deep learning network. After obtaining the feature map with channel attention, parallel void convolution layers with different expansion rates are used to capture multi-scale contextual information from the feature map with channel attention. After the multi-scale contextual information is fused, a fused feature map is obtained. The floating ice mask and river mask are detected based on the fused feature map.

[0029] Furthermore, the segmentation model fuses the fused feature maps of different sizes using a feature pyramid, uses a detection branch to predict the detection frame and the mask coefficient, uses a segmentation branch to predict the prototype mask, linearly combines the mask coefficient with the prototype mask, and obtains the floating ice mask and the river mask.

[0030] Furthermore, the calculation step of the ice-covered area ratio estimation interval includes:

[0031] For each frame, the ratio of the pixel area of ​​the floating ice mask to the pixel area of ​​the river mask is used as the estimated value of the ice-covered area ratio;

[0032] Calculate the estimated value of the ice-covered area percentage corresponding to all frames, calculate the standard error, and combine it with the given confidence level to calculate the estimated interval of the ice-covered area percentage.

[0033] Furthermore, the ice surface video includes visible light video and infrared video;

[0034] For each frame, the confidence levels of the estimated values ​​of the percentage of ice-covered area obtained from the two videos are compared, and the estimated value of the percentage of ice-covered area corresponding to the larger confidence level is selected.

[0035] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor. When the program is executed by the processor, the steps in the method for calculating the proportion of river ice cover area as described above are implemented.

[0036] A fourth aspect of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the steps in the method for calculating the proportion of river ice cover area as described above are implemented.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The present invention provides a method for calculating the proportion of ice-covered area of ​​a river. The method adopts a segmentation model to segment the river and the floating ice mask in the first frame of an ice surface video, and passes the segmentation model to a video object segmentation model. The video object segmentation model provides a tracking mask for subsequent frames, thereby avoiding manual labeling of the first frame of ice surface image and improving the efficiency of calculating the proportion of ice-covered area of ​​the river.

[0039] The present invention provides a method for calculating the proportion of ice-covered area of ​​a river. An attention mechanism is added to the segmentation model to make it more suitable for small target object detection and segmentation from the perspective of a drone, and ultimately achieve accurate video segmentation tasks, thereby accurately estimating the ice-covered area of ​​the river. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which constitute a part of the specification of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute limitations of the present invention.

[0041] Figure 1 It is a structural diagram of the backbone part of the segmentation model of the first embodiment of the present invention;

[0042] Figure 2 This is a schematic diagram of data enhancement in Embodiment 1 of the present invention;

[0043] Figure 3 It is a structural diagram of a traditional YOLACT according to the first embodiment of the present invention;

[0044] Figure 4 It is a structural diagram of the ASPP of the first embodiment of the present invention;

[0045] Figure 5 This is a structural diagram of the SENet module of the first embodiment of the present invention. DETAILED DESCRIPTION

[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0048] In the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. The present invention is further described below with reference to the accompanying drawings and embodiments.

[0049] Terminology explanation:

[0050] YOLACT (You Only Look At Coefficients) is a simple, fully convolutional model for real-time instance segmentation that achieves real-time processing speed while maintaining high accuracy.

[0051] Embodiment 1

[0052] The purpose of this first embodiment is to provide a method for calculating the proportion of river ice cover area.

[0053] Visible light cameras mainly rely on natural light or ambient light for imaging, and can capture the true color and details of objects, providing clear high-resolution images. However, the imaging effect of visible light cameras at night or in low-light environments is limited, and the image may become blurred. Infrared cameras have low resolution, but at night or in low-light environments, infrared cameras can obtain clear images and penetrate obstacles such as smoke and dust.

[0054] This embodiment provides a method for calculating the proportion of river ice cover area, wherein a group of cameras are mounted on a drone at the same time, including a visible light camera and an infrared camera, and the two cameras are controlled by a pod connected to the drone to collect ice surface videos. During the day when the visual conditions are good, the visible light camera mainly plays a role, and at night or under poor visual conditions, the infrared camera mainly plays a role.

[0055] The present embodiment provides a method for calculating the proportion of river ice cover area, comprising the following steps:

[0056] Step 1: Collect and annotate ice surface images, and then perform data enhancement to obtain a set of visible light river ice datasets and a set of infrared river ice datasets. The details are as follows:

[0057] Step 101: Integrate a visible light camera and an infrared camera on a drone. When a river freezes at a relatively low temperature, the drone is driven under good visibility conditions to photograph the river surface with the visible light camera, thereby forming a visible light river floating ice dataset, i.e., a visible light ice surface image dataset. When the visibility and illumination are poor, the drone is driven under poor visibility conditions to photograph the river surface with the infrared camera, thereby forming an infrared river floating ice dataset, i.e., an infrared ice surface image dataset.

[0058] Step 102: performing data enhancement on the captured ice surface image data, wherein the data enhancement method includes horizontal rotation, vertical flipping, brightness adjustment, and noise addition;

[0059] Step 103: Manually annotate the floating ice and the river surface on the captured visible light ice surface image and infrared ice surface image, where floating ice refers to the frozen part of the river surface, and the river surface refers to the entire river, including flowing water and ice. The annotation format of the COCO (Common Objects in Context) dataset is used to complete instance segmentation and annotation of the river and floating ice.

[0060] In step 102, data enhancement is a technique that artificially increases the size and diversity of a data set by performing various transformations on existing data samples. Data enhancement helps improve the generalization and robustness of machine learning models by expanding the data set with a modified version of the original data. The data enhancement methods used in this embodiment include horizontal rotation, vertical flipping, brightness adjustment, and noise addition. Among them, horizontal rotation is to flip the image once along the horizontal axis, vertical flipping is to flip the image once along the vertical axis, brightness adjustment is to brighten or dim the brightness of the image, and noise addition is to add noise to the image according to a certain method, thereby reducing its quality. Figure 2 As shown, the top is the captured ice surface image, the second row from left to right is the horizontal flip image, the vertical flip image, the brightness enhancement image, and the third row from left to right is the brightness reduction image and the noise addition image.

[0061] In step 103, manual labeling of the objects to be detected in the target detection task is a crucial part of the supervised learning task. In the supervised learning task, labeled data is needed to train the machine learning algorithm to identify and locate objects in the visual data. In this embodiment, it is necessary to label the data of two categories, ice surface (floating ice) and river.

[0062] Step 2: Build and improve YOLACT and add an attention mechanism to make it more suitable for small target object detection and segmentation from the perspective of drones. The details are as follows:

[0063] Step 201, establish the YOLACT network architecture, which is mainly divided into three parts: Backbone, neck and head. The backbone part is mainly responsible for feature extraction of the input ice surface image; the neck part is mainly responsible for further feature extraction and aggregation of features at different scales; the head mainly includes three parts, the detection branch, the prototype mask prediction branch and the instance segmentation result generation part. The detection branch is used to predict the detection box, category and k mask coefficients; the prototype mask prediction branch is used to predict k prototype masks of the entire ice surface image; the instance segmentation result generation part generates the final instance segmentation result by linearly combining the mask coefficients generated by the detection branch with the prototype mask.

[0064] like Figure 3As shown in the figure, YOLACT is a single-stage model that can perform target detection and instance segmentation on the input image at the same time. After the ice surface image is input, it first passes through a Backbone for feature extraction to obtain feature maps of different sizes. The Backbone can choose a deep residual network (ResNet), VGG (Very Deep Convolutional Networks for Large-Scale Visual Recognition) or MobileNet (a lightweight deep learning network), etc. The feature maps of each level are represented by C1, C2, C3, C4, and C5 respectively; then the neck part, the feature map passes through the feature pyramid (FPN), and the feature maps of different sizes are fused. The feature maps of each level are represented by P3, P4, P5, P6, and P7 respectively; finally, the head part includes a prediction head network and a prototype network; for each target body (floating ice and river surface), the prediction head network outputs the category, bounding box (detection box) information and k mask coefficients (mask coefficients); the prototype network outputs k prototype images for the current ice surface image, and finally the mask coefficients and the prototype image are added together, and then the instance segmentation mask result and the output image of the target are obtained after clipping and threshold filtering.

[0065] This example uses ResNet101 as the backbone, which includes five convolution modules and five BN modules. The parameters of each convolution layer represent the size of the convolution kernel, the convolution step size, and the number of output channels. The BN (Batch Normalization) layer can accelerate the convergence speed of the network. Then, a SENet module and an ASPP module are added.

[0066] Although YOLACT performs well in real-time, its accuracy may be poor in some complex scenarios. In particular, ice images are collected by drones, which have problems such as inaccurate detection of small targets, inaccurate bounding boxes, and inaccurate segmentation results at long distances. In addition, the algorithm is based on images, which ignores the connection between video frames.

[0067] Step 202: Add an Atrous Spatial Pyramid Pooling (ASPP) part to the end of the Backbone part of YOLACT.

[0068] The atrous convolution contained in the atrous spatial convolution pyramid is different from the traditional convolution method in that it can expand the receptive field of the convolution kernel. Atrous convolution kernels with different expansion rates can help the network extract features of different scales in parallel, so that the atrous spatial convolution pyramid can obtain feature fusion information at different scales and enrich the network's ability to extract targets of different sizes.

[0069] ASPP captures multi-scale context information from the input feature map by applying parallel dilated convolutional layers with different dilation rates, which enables the network to fuse context information of different spatial resolutions, thereby effectively capturing local and global context, and fusing multi-scale context information to obtain a fused feature map. ASPP expands the network's receptive field through convolution kernels with different dilation rates and obtains multi-scale object information.

[0070] like Figure 4 As shown in the figure, ASPP includes: a 1×1 convolution layer (Conv2d 1×1) is usually used as the first part of ASPP to reduce the dimension of the input feature map (input) or adjust the number of channels. The second part of ASPP is multiple hole convolution layers (Conv2d 3×3), each layer uses a different expansion rate rate, which can increase the receptive field without increasing the number of parameters and the amount of calculation, thereby capturing a wider range of information. The third part of ASPP is a pooling layer Pooling, which uses adaptive average pooling. This layer compresses the input feature map to 1×1 to extract global features, and then restores the feature map to its original size through bilinear interpolation upsampling. Finally, the output of the first three parts is concatenated (Concat) in the channel dimension. The concatenated feature map is fused and reduced in dimension through another 1×1 convolution layer (Conv2d 1×1), and the feature map of the ASPP module is output (output).

[0071] Step 203: Add a channel attention mechanism (SENet, Squeeze and Excitation Network) module to the Backbone part of YOLACT to enable the network to pay attention to the relationship between channels, which is beneficial for the network to understand the connection between different channels.

[0072] Add a SENet module before the dilated spatial convolutional pyramid added to the backbone of YOLACT. The input of this module is the feature map after the fifth BN layer, and its output is passed to the ASPP module.

[0073] like Figure 5 As shown, the SENet module contains the following steps:

[0074] (1) For the feature map X, whose length, width and number of channels are H', W', and C' respectively, a convolution operation F is performedtr Then the feature map U is obtained, whose length, width and number of channels are H, W and C respectively.

[0075] (2) Compress the feature map U and use the global average pooling of the channel to directly compress the W×H×C feature map containing global information into a 1×1×C feature map. The channel features of the C feature maps are compressed into a single value, generating the compressed input information z: , where c represents the channel, is the cth element of z, represents the i-th row and j-th column element of the c-th channel in the feature map U, Represents a compression operation function.

[0076] (3) Perform an excitation operation to increase or decrease the dimension of the compressed input information z: , where σ represents the Sigmoid function (S-type function), represents the ReLU function (an activation function), W1 and W2 represent the full connection operation, Represents the excitation operation function, and assigns the generated weight vector s to the feature map U to obtain the final required feature map : ,in, represents the weight vector corresponding to the cth channel, Represents a weight assignment operation.

[0077] The H×W data of each channel in the feature map U are multiplied by the weight of the corresponding channel in s, and finally the feature map with channel attention is Output to the next layer.

[0078] At this point, the modification of the YOLACT network architecture has been completed, and the neck and head parts have not been modified.

[0079] like Figure 1As shown in the figure, the Backbone part of the improved YOLACT includes the first convolution layer (the size of the convolution kernel is 7×7, the stride of the convolution is 2, and the output feature map dimension is 64), the first BN layer, the activation function layer, the maximum pooling layer, the second convolution layer (the size of the convolution kernel is 3×3, the stride of the convolution is 2, and the output feature map dimension is 64), the second BN layer, the third convolution layer (the size of the convolution kernel is 3×3, the stride of the convolution is 2, and the output feature map dimension is 64), the third BN layer, the fourth convolution layer (the size of the convolution kernel is 3×3, the stride of the convolution is 2, and the output feature map dimension is 256), the fourth BN layer, the fifth convolution layer (the size of the convolution kernel is 3×3, the stride of the convolution is 1, and the output feature map dimension is 512), the fifth BN layer, the SENet module and the ASPP part.

[0080] Step 3: Input the training set data in the visible light river ice dataset and the infrared river ice dataset into the improved YOLACT network respectively, train the improved YOLACT model, generate visible light prediction weight file and infrared prediction weight file, and obtain the trained YOLACT model.

[0081] The experiment selected the YOLACT model, trained for a total of 100 generations, with a batch size of 16, and the optimizer selected SGD (gradient descent method) with a learning rate of 0.001. The model selected ResNet101 (deep residual network) as the backbone, and improved it according to steps 202 and 203. In addition, FPN (Feature Pyramid Network) was selected as the neck part.

[0082] During the training phase of the YOLACT model, the classification loss is VFL loss (Varifocal Loss). This loss proposes an asymmetric weighted operation to address the imbalance of positive and negative samples. It is used to train dense target detectors to predict IoU-aware Classification Scores (IACS). Among them, IoU (Intersection over Union) represents the intersection over union ratio, which is used to describe the overlap between boxes. The main principle is to more accurately sort a large number of candidate detection boxes through IACS representation. The formula is as follows: , where p is the IACS predicted by the model and q is the target IoU score. This loss function aims to improve detection performance by learning a classification score (i.e., IACS), which combines the confidence of the object's existence and the positioning accuracy.

[0083] Step 4: Use the trained YOLACT model for the first frame of the test video (i.e., the ice surface video), and use it to give the actual mask shapes of the floating ice and the river, respectively. Combine the mask sizes of the floating ice and the river to estimate the proportion of the ice-covered area of ​​the river. The details are as follows:

[0084] Step 401: The YOLACT model predicts the detection frame label positions and instance segmentation mask results of the river and ice floes in the first frame of the visible light and infrared video, respectively.

[0085] Preferably, in the inference stage of the model, in order to reduce the time of model inference, Fast NMS (Fast Non-maximum Suppression) is used to select the final detection frame. Fast NMS changes the iterative calculation method of traditional NMS into a calculation method that uses matrix calculation to obtain the result at one time. The traditional NMS algorithm first sorts all frames from large to small according to the classification score, and then iterates. In each iteration, the frame with the highest classification score is retained first, and then the IoU between other frames and the frame is calculated. The frame with IoU greater than the threshold is deleted, and iterates repeatedly until there is no candidate frame. The Fast NMS algorithm first sorts all frames from large to small according to the classification score, and then calculates the IoU between all frames to obtain a symmetric matrix. Then the matrix is ​​triangulated, and the diagonal elements from the upper left to the lower right are also set to 0, and then the maximum IoU is taken from the matrix according to dimension 0, and then each IoU is judged to be greater than the filtering threshold. The frame greater than the threshold is filtered, and if the IoU of this frame exceeds the threshold, it is filtered out. Since the matrix is ​​an upper triangular matrix, the previous frame will not interfere with the subsequent frame when filtering it. This algorithm speeds up the filtering process of the detection frame.

[0086] Step 402: Calculate the pixel area enclosed by the segmentation mask of each floating ice instance, and finally add up the pixel areas covered by all floating ice as the pixel area of ​​the frozen river in the first frame of the ice surface video.

[0087] For visible light and infrared images, YOLACT can extract the bounding box of the ice floe category, whose main parameters include (x, y, h, w). Among them, (x, y) represents the position of the upper left corner of the bounding box in the image pixel coordinate system, and (h, w) represents the length and width of the bounding box. In addition, the pixel mask of each part of the ice floe in the image can be obtained, that is, the result of its instance segmentation. The total pixel area of ​​the ice floe can be obtained by adding the pixel values ​​enclosed by each part of the ice floe.

[0088] Step 403: Calculate the pixel area enclosed by the segmentation mask of each river instance, and finally add up the pixel areas covered by all rivers as the pixel area of ​​the frozen river in the first frame of the ice surface video.

[0089] For visible light and infrared images, YOLACT can extract the bounding box of the river category, whose main parameters include (x, y, h, w). Among them, (x, y) represents the position of the upper left corner of the bounding box in the image pixel coordinate system, and (h, w) represents the length and width of the bounding box. In addition, the pixel mask of each part of the river in the image can be obtained, that is, the result of its instance segmentation. The total pixel area of ​​the river can be obtained by adding the pixel values ​​enclosed by each part of the river.

[0090] Step 404: divide the total pixel area of ​​floating ice in the first frame of the ice surface video by the total pixel area of ​​the river as an estimated value of the area ratio of the frozen river, compare the confidence levels of the river instances in the visible light image and the infrared image, and take the estimated frozen river area with the larger confidence level as the final estimated value.

[0091] The area of ​​floating ice pixels in the image is divided by the total river pixel area as the estimated area ratio of the frozen river. The formula is: , among which, IceContour i and RiverContour j They represent the mask contours of the ith ice floe and the jth river obtained in step 402 and step 403, respectively, which are n and m in total. contourArea represents the function used to calculate the pixel area of ​​the region enclosed by the polygonal contour.

[0092] Finally calculate P v represents the estimated value of the proportion of ice-covered area to the total river area in the visible light image, P i It represents the estimated value of the proportion of ice-covered area to the total river area in the infrared image. Compare the confidence C of the river examples given by the large model in the visible light image and the infrared image v and C i , the one with the larger confidence level corresponds to the estimated river ice cover area, P v or P i as the final estimate.

[0093] Step 5: The ice surface video and the mask of the first frame are passed to the AOT model, which tracks and predicts the actual mask shapes of floating ice and rivers in the next period of time, and estimates the proportion of the ice-covered area of ​​the river in each frame based on the mask sizes of floating ice and rivers.

[0094] Step 501: The river and ice floe masks given by the YOLACT model in the first frame of the ice surface video are passed to the AOT model as annotations.

[0095] The AOT model is a multi-target VOS model based on the Transformer architecture. The model designs LSTT (Long Short Term Transformer) to achieve tracking and segmentation of video frame sequences. LSTT first uses a self-attention layer to learn the connection between the current frame targets. Then a long-term attention is introduced to aggregate the memory of distant frames, and a short-term attention is introduced to learn short-term temporal smoothing information. First, the source image of the first frame of the video is sent to AOT for encoding, and then all the masks calculated by the LSTT module and the improved YOLACT model are added and passed backward. Subsequent frames are also first encoded, and then the LSTT and ID (number) information from the previous frame are added and sent to the LSTT module. After decoding the output information of the LSTT module, the predicted current frame mask can be obtained. In this step, the river category mask of the first frame given by the YOLACT model is assigned IDs of 0, 1, ..., etc., and the ice category mask given by the YOLACT model is assigned IDs of 100, 101, ..., etc., and all river and ice category masks are integrated into a mask image.

[0096] Step 502: The AOT model tracks and segments subsequent consecutive frames of the video, and divides the pixel area of ​​the frozen river tracked by the image by the total pixel area of ​​the tracked river as an estimated value of the area ratio of the frozen river.

[0097] The source image, mask image, and ID information of the first frame are passed to the AOT model, and the AOT model provides tracking masks for subsequent frames. The same method as step 404 is used to calculate the estimated ice-covered area value for each subsequent frame.

[0098] Step 503: Take a frame of image as a new first frame at regular intervals, and execute step 4, step 501 and step 502 to improve the accuracy of tracking data.

[0099] In summary, based on a video, we can get the estimated value P of the ice-covered area ratio for each frame. i , i=1,2,...n.

[0100] Step 504: A confidence interval with a confidence level of 95% is comprehensively given based on the above estimated value data as the final estimated value of the proportion of river ice cover area.

[0101] According to the estimated value P i , i=1,2,...n, the final estimated interval of ice cover area is given by the following method. First, calculate P i The average : ;

[0102] Calculate the standard deviation of the data : ;

[0103] Calculate the standard error SE of the data: ;

[0104] Calculate the confidence interval CI of the data: ;

[0105] Among them, 1.96 is the confidence interval with a 95% confidence level for the standard normal distribution.

[0106] Step 6: After a certain period of time, reselect the first frame and execute steps 4 and 5. After the ice surface video ends, the estimated value of the ice-covered area ratio of the entire river is given by combining the situation of each frame.

[0107] It can be understood that the ice surface video is divided into several video segments, the first frame of each video segment is extracted, and steps 4 and 5 are performed on each video segment.

[0108] As an implementation method, the length of each video segment is adjusted according to the standard error corresponding to the previous video segment. Specifically, assuming that the ice surface video is divided into Q video segments (R1, R2, …, RQ), the length of the i-th video segment Ri is adjusted according to the standard error SE corresponding to the previous video segment Ri-1. The larger the standard error SE, the smaller the length of the i-th video segment Ri.

[0109] This embodiment adopts a segmentation model to segment the river and floating ice masks in the first frame of the ice surface video, and passes it to the video object segmentation model. The video object segmentation model provides tracking masks for subsequent frames, avoiding manual labeling of the first frame of ice surface image and improving the efficiency of calculating the proportion of river ice-covered area.

[0110] This embodiment adds an attention mechanism to the segmentation model to make it more suitable for small target object detection and segmentation from the perspective of the drone, and ultimately achieves accurate video segmentation tasks, thereby accurately estimating the ice-covered area of ​​the river.

[0111] Embodiment 2

[0112] The purpose of this second embodiment is to provide a river ice cover area ratio calculation system.

[0113] A data acquisition module is configured to: acquire ice surface video;

[0114] A video segmentation module, which is configured to: divide the ice surface video into a number of video segments, and extract the first frame of each video segment;

[0115] The mask detection module is configured to: for each video segment, segment the first frame using the segmentation model to obtain a floating ice mask and a river mask, use the floating ice mask and the river mask of the first frame as annotations, and segment all frames in the video segment using the video object segmentation model to obtain a floating ice mask and a river mask of each frame;

[0116] A proportion calculation module is configured to: calculate an estimated interval of ice-covered area proportion based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame;

[0117] Among them, the segmentation model uses the channel attention mechanism to process the feature map of the first frame. After obtaining the feature map with channel attention, parallel void convolution layers with different expansion rates are used to capture multi-scale contextual information from the feature map with channel attention. The multi-scale contextual information is fused to obtain a fused feature map. The floating ice mask and river mask are detected based on the fused feature map.

[0118] Among them, the segmentation model fuses the fused feature maps of different sizes using the feature pyramid, uses the detection branch to predict the detection frame and mask coefficient, uses the segmentation branch to predict the prototype mask, and linearly combines the mask coefficient with the prototype mask to obtain the floating ice mask and river mask.

[0119] The calculation steps for estimating the ice-covered area ratio include:

[0120] For each frame, the ratio of the pixel area of ​​the floating ice mask to the pixel area of ​​the river mask is used as the estimated value of the ice-covered area ratio;

[0121] Calculate the estimated value of the ice-covered area percentage corresponding to all frames, calculate the standard error, and combine it with the given confidence level to calculate the estimated interval of the ice-covered area percentage.

[0122] Among them, ice surface videos include visible light videos and infrared videos;

[0123] For each frame, the confidence levels of the estimated values ​​of the percentage of ice-covered area obtained from the two videos are compared, and the estimated value of the percentage of ice-covered area corresponding to the larger confidence level is selected.

[0124] It should be noted here that each module in this embodiment corresponds to each step in Example 1 one by one, and the specific implementation process is the same, which will not be repeated here.

[0125] Embodiment 3

[0126] This embodiment provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor. When the program is executed by the processor, the steps in the method for calculating the proportion of river ice cover area as described in the above-mentioned embodiment 1 are implemented.

[0127] Embodiment 4

[0128] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps in the method for calculating the proportion of river ice cover area as described in the above-mentioned embodiment 1 are implemented.

[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0130] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A method for calculating the proportion of river ice cover area, characterized in that: include: Acquire an ice surface video; the ice surface video includes a visible light video and an infrared video; Divide the ice surface video into several video segments and extract the first frame of each video segment; For each video segment, the first frame is segmented by the YOLACT segmentation model to obtain the floating ice mask and the river mask, and the floating ice mask and the river mask of the first frame are passed as annotations to the video object segmentation model, and all frames in the video segment are segmented by the video object segmentation model to obtain the floating ice mask and the river mask of each frame; the video object segmentation model adopts the AOT model; the YOLACT segmentation model is a simple, fully convolutional model for real-time instance segmentation; the AOT model is a multi-target video object segmentation model based on the Transformer architecture; Based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame, the estimated interval of the ice cover area ratio is calculated; The calculation steps of the ice-covered area ratio estimation interval include: For each frame, the ratio of the pixel area of ​​the floating ice mask to the pixel area of ​​the river mask is used as the estimated value of the ice-covered area ratio; Calculate the estimated value of the ice-covered area percentage corresponding to all frames, calculate the standard error, and calculate the estimated interval of the ice-covered area percentage based on the given confidence level; For each frame, the confidence levels of the estimated values ​​of the percentage of ice-covered area obtained from the two videos are compared, and the estimated value of the percentage of ice-covered area corresponding to the larger confidence level is selected; Among them, the segmentation model uses the channel attention mechanism to process the feature map of the first frame. After obtaining the feature map with channel attention, parallel void convolution layers with different expansion rates are used to capture multi-scale contextual information from the feature map with channel attention. The multi-scale contextual information is fused to obtain a fused feature map. The floating ice mask and river mask are detected based on the fused feature map.

2. A method for calculating the proportion of river ice cover area according to claim 1, characterized in that: The segmentation model fuses fused feature maps of different sizes using a feature pyramid, uses a detection branch to predict a detection frame and a mask coefficient, uses a segmentation branch to predict a prototype mask, linearly combines the mask coefficient with the prototype mask, and obtains an iceberg mask and a river mask.

3. A river ice cover area ratio calculation system, characterized in that: include: A data acquisition module is configured to: acquire ice surface video; the ice surface video includes visible light video and infrared video; A video segmentation module, which is configured to: divide the ice surface video into a number of video segments, and extract the first frame of each video segment; The mask detection module is configured as follows: for each video segment, the first frame is segmented by the YOLACT segmentation model to obtain the floating ice mask and the river mask, the floating ice mask and the river mask of the first frame are passed as annotations to the video object segmentation model, and all frames in the video segment are segmented by the video object segmentation model to obtain the floating ice mask and the river mask of each frame; the video object segmentation model adopts the AOT model; the YOLACT segmentation model is a simple, fully convolutional model for real-time instance segmentation; the AOT model is a multi-target video object segmentation model based on the Transformer architecture; A proportion calculation module is configured to: calculate an estimated interval of ice-covered area proportion based on the pixel area of ​​the floating ice mask and the pixel area of ​​the river mask corresponding to each frame; The calculation steps of the ice-covered area ratio estimation interval include: For each frame, the ratio of the pixel area of ​​the floating ice mask to the pixel area of ​​the river mask is used as the estimated value of the ice-covered area ratio; Calculate the estimated value of the ice-covered area percentage corresponding to all frames, calculate the standard error, and calculate the estimated interval of the ice-covered area percentage based on the given confidence level; For each frame, the confidence levels of the estimated values ​​of the percentage of ice-covered area obtained from the two videos are compared, and the estimated value of the percentage of ice-covered area corresponding to the larger confidence level is selected; Among them, the segmentation model uses the channel attention mechanism to process the feature map of the first frame. After obtaining the feature map with channel attention, parallel void convolution layers with different expansion rates are used to capture multi-scale contextual information from the feature map with channel attention. The multi-scale contextual information is fused to obtain a fused feature map. The floating ice mask and river mask are detected based on the fused feature map.

4. A river ice cover area ratio calculation system as claimed in claim 3, characterized in that: The segmentation model fuses fused feature maps of different sizes using a feature pyramid, uses a detection branch to predict a detection frame and a mask coefficient, uses a segmentation branch to predict a prototype mask, linearly combines the mask coefficient with the prototype mask, and obtains an iceberg mask and a river mask.

5. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor, characterized in that: When the program is executed by a processor, the steps in a method for calculating the proportion of river ice cover area as described in any one of claims 1-2 are implemented.

6. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for calculating the proportion of river ice cover area as described in any one of claims 1-2 are implemented.

Citation Information

Patent Citations

  • Object detection method, object detection device and electronic equipment

    CN111488776A

  • Full-convolution single-stage human body instance segmentation method in natural scene

    CN111597920A

  • Visible light and infrared light fused target recognition method

    CN111611905A

  • Photovoltaic panel hot spot detection and area proportion calculation method

    CN116433591A