Fluid leakage detection method, device and equipment and storage medium
By extracting the global long-term dependency features and local micro-deformation sensitive features in the video data, combined with the classification recognition model, the problem of low detection accuracy in pipeline fluid leakage detection is solved, and the accurate identification of serious and minor leakage is achieved.
Patent Information
- Application Number
- CN202311847959.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art detects pipeline fluid leakage, especially in the case of slight leakage, the detection accuracy is not high or is prone to missed detection, making it difficult to identify in time.
By acquiring video data within the preset time period, the first feature extraction model is used to extract the global long-time dependency feature and the second feature extraction model is used to extract local micro-deformation sensitive features, and combined with the classification recognition model for fusion recognition to detect long-term state changes and local slight changes in the pipeline.
It improves the detection accuracy of pipeline leakage, can identify serious and minor leakage situations, and achieves timely identification of fluid leakage.
Smart Images

Figure CN120236107A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a fluid leakage detection method, device, equipment and storage medium. Background Art
[0002] There are a large number of important equipment and pipelines in factories for storing or transporting fluids. Affected by various factors such as pipeline aging, component corrosion, thermal expansion and contraction, and untimely maintenance, liquid leakage may occur. The occurrence of fluid leakage will not only cause waste of resources and economic losses, but may also cause short circuits, ground settlement, breeding of harmful bacteria, etc., resulting in a series of serious problems such as equipment damage. Therefore, the timely detection and identification of pipeline fluid leakage has important practical significance.
[0003] At present, manual inspection is generally used to check for running, dripping and leaking. Some factories also adopt automated detection methods. The common method is to use cameras and image processing technology, such as applying clustering analysis methods to detect leakage phenomena in images, and using segmentation methods based on region growing for leakage localization.
[0004] However, due to the tiny leakage of fluids such as pipeline oil and water, the tiny leakage has characteristics such as slow change, no fixed shape, and fine liquid leakage at some leakage points. There are still problems such as low detection accuracy or missed detection in a large number of detection models in this regard. Summary of the Invention
[0005] This application proposes a fluid leakage detection method, device, equipment and storage medium, which can solve the technical problems of low detection accuracy or missed detection in the current detection of tiny leakage of liquids such as fluids in liquid-filled equipment and liquid pipelines.
[0006] The first aspect embodiment of this application proposes a fluid leakage detection method, including:
[0007] Obtain video data of a detection object within a preset time period, where the detection object is used for storing or transporting fluids;
[0008] Extract the global long-term dependence features in the video data through a first feature extraction model, where the global long-term dependence features are used to characterize the global state change of the detection object within the preset time period;
[0009] Extract the local micro-deformation sensitive features in the video data through a second feature extraction model, where the local micro-deformation sensitive features are used to characterize the local state change of the detection object within the preset time period;
[0010] Fuse the full - length long - term dependence feature and the local micro - deformation sensitive feature, and then input them into a classification and recognition model for recognition. Output the recognition result, which is used to characterize whether fluid leakage has occurred in the detection object during the preset time period.
[0011] An embodiment of the second aspect of the present application provides a fluid leakage detection device, including:
[0012] An acquisition module, configured to acquire video data of a detection object within a preset time period, where the detection object is used to store or transport fluid;
[0013] An extraction module, configured to extract the full - length long - term dependence feature in the video data through a first feature extraction model, where the full - length long - term dependence feature is used to characterize the global state change of the detection object within the preset time period;
[0014] The extraction module is further configured to extract the local micro - deformation sensitive feature in the video data through a second feature extraction model, where the local micro - deformation sensitive feature is used to characterize the local state change of the detection object within the preset time period;
[0015] A recognition module, configured to fuse the full - length long - term dependence feature and the local micro - deformation sensitive feature, and then input them into a classification and recognition model for recognition, and output a recognition result, where the recognition result is used to characterize whether fluid leakage has occurred in the detection object within the preset time period.
[0016] An embodiment of the third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the method described in the first aspect above.
[0017] An embodiment of the fourth aspect of the present application provides a computer - readable storage medium, on which a computer program is stored. The program is executed by a processor to implement the method described in the first aspect above.
[0018] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0019] The present application provides a leakage detection method, apparatus, device, and storage medium. The method includes: obtaining video data of a detection object within a preset time period; extracting global long-term dependence features in the video data through a first feature extraction model; extracting local micro-deformation sensitive features in the video data through a second feature extraction model; concatenating the global long-term dependence features and the local micro-deformation sensitive features, and inputting them into a classification and recognition model to obtain the leakage condition of the detection object within the preset time period. In the embodiments of the present application, by combining the global long-term dependence features and the local micro-deformation sensitive features in the video data, the long-term state changes and local tiny changes of pipeline leakage can be reflected, thereby improving the detection accuracy and detecting severe leakage conditions and tiny leakage conditions of the detection object.
[0020] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0022] In the drawings:
[0023] Figure 1 A flowchart of a fluid leakage detection method provided by an embodiment of the present application is shown;
[0024] Figure 2 A flowchart of extracting local micro-deformation sensitive features provided by an embodiment of the present application is shown;
[0025] Figure 3 A structural diagram of a Video Swin Transformer-tiny model provided by an embodiment of the present application is shown;
[0026] Figure 4 A schematic flowchart of a fluid leakage detection method provided by an embodiment of the present application is shown;
[0027] Figure 5 A structural diagram of a fluid leakage detection apparatus provided by an embodiment of the present application is shown;
[0028] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present application is shown;
[0029] Figure 7The figure shows a schematic diagram of a storage medium provided by an embodiment of the present application. Detailed implementation manners
[0030] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.
[0031] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present application belongs.
[0032] Continuing with the above background art, since the occurrence of fluid leakage will not only cause waste of resources and economic losses, but may also cause short circuits, ground settlement, breeding of harmful germs, etc., resulting in a series of serious problems such as equipment damage.
[0033] To solve the above problems, factories generally use manual inspections to check for running, dripping, and leaking, but manual inspections not only take a long time, are difficult to organize and analyze historical data, but also have poor real-time performance.
[0034] With the continuous development of artificial intelligence technology, deep learning models have begun to be widely used in the field of pipeline leakage detection. For example, by analyzing various pipeline data to detect whether there is leakage, or by installing image acquisition devices such as cameras, and deploying the trained deep learning model on the server to detect the status of industrial site equipment and pipelines, including oil leakage detection, gas leakage detection, etc.
[0035] However, since some liquid leakage is a slow and minute change process, there are still problems such as low detection accuracy or missed detection in a large number of detection models in this regard.
[0036] To solve the above problems, the embodiments of the present application provide a fluid leakage detection method, device, equipment, and storage medium. In the embodiments of the present application, by obtaining video data of a detection object within a preset time period; extracting the global long-term dependence features in the video data through a first feature extraction model; extracting the local micro-deformation sensitive features in the video data through a second feature extraction model; concatenating the global long-term dependence features and the local micro-deformation sensitive features and inputting them into a classification and recognition model to obtain the leakage situation of the detection object within the preset time period. The embodiments of the present application combine the global long-term dependence features and the local micro-deformation sensitive features in the video data, which can reflect the long-term state changes and local minute changes of pipeline leakage, that is, can detect the serious leakage situation and minute leakage situation of the detection object.
[0037] The fluid leakage detection method of the present application can be executed by a computing device, which can be a server, such as a single server, multiple servers, a server cluster, a cloud computing platform, etc. Optionally, the computing device can also be a terminal device, such as a mobile phone, a tablet computer, a game console, a portable computer, a desktop computer, an advertising machine, an all-in-one computer, etc. The present application does not limit the type and number of computing devices.
[0038] For the execution subject, in each embodiment of the present application, the computing device is taken as an example for illustration.
[0039] Next, a fluid leakage detection method proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0040] See Figure 1 , the method specifically includes the following steps:
[0041] S101. Obtain the video data of the detection object within a preset time period.
[0042] The detection object is used to store or transport fluid.
[0043] Among them, the preset time period can be flexibly set according to the actual situation. For example, if the acquisition frequency of the video data acquisition device is high, the preset time period can be a short time; if the sampling frequency of the video acquisition device is low, the preset time period can be a long time.
[0044] The detection object can be a device or pipeline for the user to store or transport fluid, etc.
[0045] The fluid can include gas and liquid.
[0046] The video data can be video images with a large number of frames. The number of frames of the video data is related to the preset time period and the acquisition frequency of the video data acquisition device. Therefore, the length of the preset time period is related to the preset number of frames and the acquisition frequency of the video data acquisition device. For example, if the sampling frequency is represented by the number of video frames sampled per second, then the preset number of frames is the product of the sampling frequency and the preset time period.
[0047] S102. Extract the global long-term dependence features in the video data through the first feature extraction model.
[0048] The global long-term dependence features are used to characterize the global state change of the detection object within the preset time period.
[0049] Among them, the first feature extraction model can be a Transformer model, a Long Short-Term Memory (LSTM) model, a Gated Recurrent Unit (GRU) model, a WaveNet model, a video swin transformer model, and so on.
[0050] Among them, for the video swin transformer model, Video Swin Transformer is a video understanding model based on the Transformer architecture, which can solve the processing efficiency and expression ability limitations of traditional convolutional neural networks in dealing with change features that can only be shown within a long time.
[0051] Through the first feature extraction model, the global long-term dependence features in the video data can be obtained, that is, the features representing the global state changes of the detection object within a preset time period can be obtained. These global long-term dependence features can be used to analyze the general leakage situation of the detection object. For example, they can be used to analyze the leakage situation visible to the naked eye.
[0052] S103. Extract the local micro-deformation sensitive features in the video data through the second feature extraction model.
[0053] The local micro-deformation sensitive features are used to represent the local state changes of the detection object within a preset time period.
[0054] Since the changes in many pipeline leaks, especially in the early stage of leakage, are very subtle and the first feature extraction model is not easy to capture such changes, in the model design of this application, a local micro-deformation sensitive feature extraction module is added to obtain the local micro-deformation sensitive features, that is, the features used to represent the local state changes of the detection object within a preset time period. These local micro-deformation sensitive features can be used to analyze the micro-leakage situation of the detection object. For example, they can be used to analyze leakage situations with slow changes, no fixed shape, and fine liquid leakage at some leakage points.
[0055] S104. Fuse the global long-term dependence features and the local micro-deformation sensitive features and then input them into the classification and recognition model for recognition, and output the recognition result.
[0056] The recognition result is used to represent whether fluid leakage occurs in the detection object within a preset time period.
[0057] Among them, the way of fusing the global long-term dependence feature and the local micro-deformation sensitive feature can be to connect the above two features together to form a more comprehensive feature representation. For example, assuming that the two features are represented in binary, the global long-term dependence feature is 01010100, and the local micro-deformation sensitive feature is 01110111. The fused feature can be 0101010001110111.
[0058] The classification and recognition model can include a Multilayer Perceptron (MLP for short), which is used to determine the recognition result of the detection object based on the fused feature.
[0059] The recognition result can be the leakage probability of the detection object. For example, the recognition result can be that the leakage probability is 70% and the non-leakage probability is 30%.
[0060] Furthermore, a leakage probability threshold can be set. When the leakage probability is greater than the leakage probability threshold, it is considered that the detection object has leaked within the preset time period.
[0061] In some embodiments, the recognition result can also be the leakage result of the detection object. For example, the recognition result can be that the detection object has leaked within the preset time period or the detection object has not leaked within the preset time period.
[0062] This application proposes a leakage detection method, which includes: obtaining video data of a detection object within a preset time period; extracting the global long-term dependence feature in the video data through a first feature extraction model; extracting the local micro-deformation sensitive feature in the video data through a second feature extraction model; concatenating the global long-term dependence feature and the local micro-deformation sensitive feature and inputting them into a classification and recognition model to obtain the leakage situation of the detection object within the preset time period. The embodiments of this application combine the global long-term dependence feature and the local micro-deformation sensitive feature in the video data, which can reflect the long-term state change and local micro-changes of pipeline leakage, thereby improving the detection accuracy and detecting both serious leakage and minor leakage situations of the detection object.
[0063] In some embodiments, extracting the local micro-deformation sensitive feature in the video data through the second feature extraction model includes: performing differential processing on any two adjacent frames of images in the video data to obtain a differential image; dividing the differential image into multiple sub-differential images; inputting the multiple sub-differential images into multiple convolutional units in the second feature extraction model one by one to obtain multiple sub-differential feature images; performing global average pooling operation on the multiple sub-differential feature images to obtain the local micro-deformation sensitive feature corresponding to the differential image.
[0064] Performing differential processing on any two adjacent frames of images in the video data can generally be implemented as:
[0065] Convert any two adjacent frames of images into grayscale images: If the input is a color image, it first needs to be converted into a grayscale image. This can be achieved by performing a weighted average of the pixel values of the red, green, and blue channels to obtain the grayscale image.
[0066] Calculate the difference image: Subtract the grayscale values of the previous frame image from the current frame image to obtain the difference image. Larger pixel values in the difference image indicate obvious changes between the two frames.
[0067] Thresholding: To further enhance the changed regions and reduce the influence of noise, thresholding can be performed on the difference image. Set a suitable threshold, set the pixels less than the threshold to 0, and the pixels greater than or equal to the threshold to 255 (or other maximum pixel values) to obtain a binary difference image.
[0068] It can be understood that the difference operation can be used for change detection between consecutive frames. Therefore, local micro-deformation sensitive features can be extracted based on the difference image after the difference operation.
[0069] Since local micro-deformation sensitive features are used to characterize the local state changes of the detection object within a preset time period, that is, for analyzing leakage situations such as slow changes, no fixed shape, and fine leakage at some leakage points, it is necessary to divide the difference image into multiple sub-difference images along a preset direction. For example, the difference image can be divided into multiple sub-images of uniform size along the horizontal direction, or the difference image can be divided into multiple sub-images of uniform size along the horizontal direction, etc. Further, the multiple sub-difference images are input into multiple convolutional units in the second feature extraction model one by one to obtain multiple sub-difference feature images. Compared with extracting the difference features of the difference image, that is, extracting local micro-deformation sensitive features through the difference image, the difference sub-feature images of multiple sub-difference images are respectively subjected to global average pooling operation to obtain the local micro-deformation sensitive features corresponding to the difference image. The receptive field of the model will be restricted within the pre-segmented region, and the attention of the model will shift to a smaller range to explore the features of a smaller region. So that more refined and accurate local micro-deformation sensitive features can be obtained.
[0070] Figure 2 Shows a flowchart of extracting local micro-deformation sensitive features provided by an embodiment of the present application. As Figure 2 shown, the difference image is divided into 4 sub-difference images of uniform size along the horizontal direction, and then the same convolution operation is performed on these 4 sub-difference images respectively, that is, through Figure 2Multiple convolutional units in it are used to obtain the differential sub - feature images of each differential sub - image. Finally, these 4 feature maps are concatenated to obtain the feature map of the original differential image, that is, the locally micro - deformation - sensitive feature.
[0071] In some embodiments, the convolutional unit includes multiple serially connected convolutional blocks, and each convolutional block includes multiple serially connected convolutional layers and pooling layers; inputting multiple differential sub - images into multiple convolutional units in the second feature extraction model one - to - one to obtain multiple differential sub - feature images, including: setting corresponding processing parameter values for each convolutional layer and each pooling layer respectively, and the processing parameter values include at least one of the following parameter values: input channels, output channels, feature size, and number of kernels; inputting multiple differential sub - images into multiple convolutional units in the second feature extraction model one - to - one, and performing serial processing through each convolutional layer and each pooling layer in the convolutional unit according to the corresponding processing parameter values respectively to obtain multiple differential sub - feature images.
[0072] To ensure the accuracy of feature extraction for each differential sub - image, generally, multiple serially connected convolutional blocks can be set for each convolutional unit, and each convolutional block is provided with multiple serially connected convolutional layers and a pooling layer serially connected to the convolutional layer. For example, 3 convolutional blocks can be set for the convolutional unit, and each convolutional block is provided with two serially connected convolutional layers and a pooling layer serially connected to the convolutional layer. Inputting multiple differential sub - images into multiple convolutional units in the second feature extraction model one - to - one, each convolutional unit performs convolution and pooling operations respectively through the convolutional layers and a pooling layer serially connected to the convolutional layer of each convolutional block, and further, each convolutional block performs an average pooling operation once, so as to perform serial processing through each convolutional layer and each pooling layer in the convolutional unit according to the corresponding processing parameter values respectively to obtain multiple differential sub - feature images.
[0073] Table 1 List of the structure and processing parameter values of the locally micro - deformation - sensitive feature extraction module
[0074]
[0075] As shown in Table 1, the convolutional unit consists of 3 convolutional blocks, named Block 1, Block 2, and Block 3 in sequence. Each convolutional block includes 2 convolutional layers and 1 pooling layer. For example, in Block 1, the two convolutional layers are Conv1 and Conv2 respectively, and the pooling layer can be the max - pooling layer MaxPool, and the same applies to Block 2 and Block 3. When each convolutional layer and pooling layer perform convolution and pooling operations, they perform convolution and pooling operations respectively according to the corresponding processing parameter values in Table 1. After the convolution and pooling operations of Block 3, a global average pooling operation is performed to obtain the differential sub - image feature corresponding to this convolutional unit.
[0076] In some embodiments, the global long-term dependence features in the video data are extracted by the first feature extraction model, including: obtaining video images of a preset number of frames from the video data; splitting the video images of the preset number of frames to obtain video vectors; and inputting the video vectors into a plurality of sequentially connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data.
[0077] Among them, the number of frames of the video images in the video data is related to the acquisition frequency of the acquisition device and the length of the preset time. In the process of the first feature extraction model extracting the global long-term dependence features in the video data, if the number of frames of the video images is too small, the obtained global long-term dependence features will be inaccurate. If the number of frames of the video images is too large, the resource consumption in the extraction process will be too large. Therefore, in order to avoid inaccurate global long-term dependence features, the number of frames of the video images in the video data will be greater than the preset number of frames. In order to reduce the resource consumption in the extraction process while ensuring the accuracy of the obtained global long-term dependence features, video images of the preset number of frames can be obtained from the video data. For example, if the acquisition frequency of the acquisition device is 30 Hz and the preset time is 3 seconds, then the number of frames of the video data is 90 frames. If the preset number of frames is 32 frames, 32 frames of video images can be obtained from the video data.
[0078] Figure 3 Fig. shows an architecture diagram of a Video Swin Transformer-tiny model provided by an embodiment of the present application. As Figure 3 shown, after obtaining the video images of the preset number of frames, the video images are input into the Figure 3 splitting unit in, and the video images are split by the splitting unit to obtain video vectors, and then input into a plurality of sequentially connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data.
[0079] The process of splitting the video images of the preset number of frames to obtain video vectors may include the following steps:
[0080] Among them, image preprocessing: For the sampled images, some preprocessing operations can be performed, such as adjusting the image size, cropping, rotating, denoising, etc., so as to better extract features.
[0081] Feature extraction: Through the feature extraction algorithm, the information in the image is extracted and converted into a one-dimensional vector. Common feature extraction methods include using a convolutional neural network (CNN) for feature extraction, using a hand-designed image feature extraction algorithm, etc.
[0082] Vectorization: Convert the extracted feature representation into a vector form. This can be achieved by simply flattening the features into a one-dimensional vector or using a dimensionality reduction algorithm to reduce the dimensionality of the features.
[0083] Normalization: To ensure the comparability of vectors between different video frames, the vectors can be normalized, such as scaling the vectors to unit norm, with a mean of 0 and a variance of 1, etc.
[0084] In some embodiments, if the first feature extraction model is a video swin transformer model, the obtained video vector can be a vector, where T represents the number of frames, and H and W represent the height and width of each frame of the picture respectively. In the embodiments of the present application, both H and W are taken as 224.
[0085] Further, the video vector is input into a plurality of sequentially connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data.
[0086] As Figure 3 shown, if the first feature extraction model is a Video Swin Transformer-tiny model, the video vector will be input into a plurality of sequentially connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data
[0087] In some embodiments, the first feature extraction unit includes a linear embedding subunit and a video understanding model, and the first feature unit is any one of the plurality of feature extraction units; the input vector is subjected to a fully connected process through the linear embedding subunit of the first feature extraction unit to obtain a feature vector; the feature vector is processed through the multi-head self-attention mechanism of the video understanding model of the first feature extraction unit to obtain an output feature.
[0088] As Figure 3 shown, each feature extraction unit includes a linear embedding subunit and a video understanding model. The video vector is subjected to a fully connected process through the linear embedding subunit to obtain a feature vector with a fixed length. Then, through the video understanding model and by adding a windowed multi-head self-attention mechanism, the feature interaction and extraction of the image patches extracted in the previous layer are performed. The combination of the linear embedding module and the Video Swin Transformer block is called a feature extraction unit. There are a total of 4 feature extraction units in the Video Swin Transformer-tiny model. Through the above steps, the global long-term dependence features in the video data can be obtained.
[0089] In some embodiments, the first feature extraction model, the second feature extraction model, and the classification and recognition model are pre-trained. The training process includes: obtaining a training set, where the training samples in the training set include multiple video data and the target recognition results corresponding to the target detection objects to which the multiple video data belong; inputting the target video data into the first feature extraction model and the second feature extraction model, and respectively outputting the predicted global long-term dependence features and the predicted local micro-deformation sensitive features corresponding to the target video data, where the target video data is any one of the multiple video data; fusing the predicted global long-term dependence features and the predicted local micro-deformation sensitive features and inputting the fused features into the classification and recognition model to output the predicted recognition result corresponding to the target detection object to which the target video data belongs; calculating the loss function value based on the target recognition result and the predicted recognition result; adjusting the parameters of the first feature extraction model, the second feature extraction model, and the classification and recognition model based on the loss function value, and continuing the training until the preset training completion condition is met, thereby obtaining the trained first feature extraction model, the second feature extraction model, and the classification and recognition model.
[0090] In some embodiments, the existing leakage images and video data in the factory can be integrated to construct a training set covering multiple scenarios such as indoor and outdoor pipelines and valves. The training samples in the training set include multiple video data and the target recognition results corresponding to the target detection objects to which the multiple video data belong.
[0091] Input the target video data into the first feature extraction model and the second feature extraction model, and respectively output the predicted global long-term dependence features and the predicted local micro-deformation sensitive features corresponding to the target video data.
[0092] Fuse the predicted global long-term dependence features and the predicted local micro-deformation sensitive features and input the fused features into the classification and recognition model to output the predicted recognition result corresponding to the target detection object to which the target video data belongs.
[0093] The loss function can be the cross-entropy loss function, which can be expressed by the following formula:
[0094]
[0095] loss is the cross-entropy loss during the network iteration process, p(x) is the target recognition result, and q(x) is the predicted recognition result.
[0096] Adjust the parameters of the first feature extraction model, the second feature extraction model, and the classification and recognition model based on the loss function value, and continue the training until the preset training completion condition is met, thereby obtaining the trained first feature extraction model, the second feature extraction model, and the classification and recognition model.
[0097] Among them, the preset training completion condition can be the preset number of training times, or the value of the loss function is less than the preset threshold. The preset threshold can be flexibly set based on the actual situation and will not be elaborated here.
[0098] In some embodiments, in addition to training the above three models based on the training set, a validation set and a test set can also be obtained. Among them, the validation set is used to evaluate the performance of the model, and the network model is adjusted according to the actual situation during the training process to improve the model performance. The test set can be used for the evaluation of the final performance of the model.
[0099] In some embodiments, the existing leakage images and video data of the factory can be integrated to construct a pipeline leakage detection dataset covering multiple scenarios such as indoor and outdoor pipelines and valves. After data cleaning, a training set, a validation set, and a test set can be obtained.
[0100] In some embodiments, the classification and recognition model includes a multi-layer perceptron. Each multi-layer perceptron includes an input layer, a hidden layer, and an output layer. The long-term global dependence features and local micro-deformation sensitive features are fused and then input into the classification and recognition model for recognition, and the recognition result is output, including: inputting the fused features into the hidden layer through the input layer; fully connecting the fused features through the hidden layer to obtain a connection vector; judging the connection vector through the activation function of the hidden layer to determine the leakage probability corresponding to the connection vector; determining the recognition result of the detection object based on the leakage probability through the output layer.
[0101] After fusing the long-term global dependence feature extraction module and the local micro-deformation sensitive features, the fused features can be input into the classification and recognition model. The input layer of the classification and recognition model receives the fused features and inputs the fused features into the hidden layer.
[0102] It should be noted that a multi-layer perceptron can include multiple hidden layers. Each hidden layer can be split into two parts: full connection and activation function. Fully connecting the fused features through the hidden layer to obtain a connection vector can be implemented as multiplying the input fused features by a preset weight to obtain a connection vector, and further judging the connection vector through the activation function of the hidden layer to determine the leakage probability corresponding to the connection vector. The activation function can be a Sigmoid function, a ReLU function, a Tanh function, etc.
[0103] Furthermore, the recognition result of the detection object is determined based on the leakage probability through the output layer. The recognition result is the leakage probability of the detection object, or whether the detection object has leaked.
[0104] If the recognition result is the leakage probability of the detection object, a leakage probability threshold can be set. When the leakage probability is greater than the leakage probability threshold, it is considered that the detection object has leaked within the preset time period.
[0105] In some embodiments, to completely describe the fluid leakage detection provided by the embodiments of the present application, the embodiments of the present application also provide a schematic flowchart of a fluid leakage detection method, as Figure 4 shown below:
[0106] First, obtain video images of a preset number of frames, and input the video images of the preset number of frames into a first feature extraction model and a second feature extraction model respectively. Among them, the first feature extraction model is used to extract the global long-term dependence features in the video data, and the second feature extraction model is used to extract the local micro-deformation sensitive features in the video data. Further, fuse the global long-term dependence features and the local micro-deformation sensitive features, and input the fused features into a classification recognition model for recognition, so that the classification recognition model outputs a recognition result for the detection object.
[0107] The embodiments of the present application also provide a fluid leakage detection device, which is used to execute the fluid leakage detection method provided in any of the above embodiments. As Figure 5 shown, the device includes: an acquisition module 501, an extraction module 502, and a recognition module 503.
[0108] The acquisition module 501 is used to acquire video data of a detection object within a preset time period, and the detection object is used to store or transport fluid;
[0109] The extraction module 502 is used to extract the global long-term dependence features in the video data through the first feature extraction model, and the global long-term dependence features are used to characterize the global state change of the detection object within the preset time period;
[0110] The extraction module 502 is also used to extract the local micro-deformation sensitive features in the video data through the second feature extraction model, and the local micro-deformation sensitive features are used to characterize the local state change of the detection object within the preset time period;
[0111] The recognition module 503 is used to fuse the global long-term dependence features and the local micro-deformation sensitive features and then input them into the classification recognition model for recognition, and output a recognition result, and the recognition result is used to characterize whether fluid leakage occurs in the detection object within the preset time period.
[0112] The present application provides a leakage detection device, which obtains video data of a detection object within a preset time period; extracts global long-term dependence features in the video data through a first feature extraction model; extracts local micro-deformation sensitive features in the video data through a second feature extraction model; concatenates the global long-term dependence features and the local micro-deformation sensitive features, inputs them into a classification and recognition model, and obtains the leakage condition of the detection object within the preset time period. In the embodiments of the present application, by combining the global long-term dependence features and the local micro-deformation sensitive features in the video data, the long-term state change and local micro-changes of pipeline leakage can be reflected, that is, the severe leakage condition and micro-leakage condition of the detection object can be detected.
[0113] In some embodiments, the extraction module 502 is specifically configured to:
[0114] Perform differential processing on any two adjacent images in the video data to obtain a differential image;
[0115] Divide the differential image into a plurality of sub-differential images;
[0116] Input the plurality of sub-differential images into a plurality of convolutional units in the second feature extraction model one by one to obtain a plurality of sub-differential feature images;
[0117] Stitch the plurality of sub-differential feature images to obtain local micro-deformation sensitive features corresponding to the differential image.
[0118] In some embodiments, the extraction module 502 is specifically configured to:
[0119] Average-divide the differential image into a plurality of sub-differential images along a preset direction.
[0120] In some embodiments, the convolutional unit includes a plurality of cascaded convolutional blocks, and each convolutional block includes a plurality of cascaded convolutional layers and pooling layers; the extraction module 502 is further specifically configured to:
[0121] Set corresponding processing parameter values for each convolutional layer and each pooling layer respectively, and the processing parameter values include at least one of the following parameter values: input channel, output channel, feature size, and number of kernels;
[0122] Input the plurality of sub-differential images into a plurality of convolutional units in the second feature extraction model one by one, and perform serial processing on each convolutional layer and each pooling layer in the convolutional unit according to the corresponding processing parameter values respectively to obtain a plurality of sub-differential feature images.
[0123] In some embodiments, the extraction module 502 is further specifically configured to:
[0124] Obtain video images of a preset number of frames from the video data;
[0125] Split the video images of the preset number of frames to obtain video vectors;
[0126] Input the video vectors into a series of connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data.
[0127] In some embodiments, the first feature extraction unit includes a linear embedding subunit and a video understanding model, and the first feature unit is any one of the multiple feature extraction units;
[0128] Perform a fully connected process on the input vector through the linear embedding subunit of the first feature extraction unit to obtain a feature vector;
[0129] Process the feature vector through the multi-head self-attention mechanism of the video understanding model of the first feature extraction unit to obtain an output feature.
[0130] In some embodiments, the first feature extraction model, the second feature extraction model, and the classification and recognition model are pre-trained, and the training process includes:
[0131] Obtain a training set, where the training samples in the training set include multiple video data and the target recognition results corresponding to the target detection objects to which the multiple video data belong;
[0132] Input the target video data into the first feature extraction model and the second feature extraction model, and respectively output the predicted global long-term dependence features and the predicted local micro-deformation sensitive features corresponding to the target video data. The target video data is any one of the multiple video data;
[0133] Fuse the predicted global long-term dependence features and the predicted local micro-deformation sensitive features and input them into the classification and recognition model to output the predicted recognition result corresponding to the target detection object to which the target video data belongs;
[0134] Calculate the loss function value based on the target recognition result and the predicted recognition result;
[0135] Based on the loss function value, adjust the parameters of the first feature extraction model, the second feature extraction model, and the classification and recognition model, and continue training until the preset training completion condition is met, to obtain the trained first feature extraction model, second feature extraction model, and classification and recognition model.
[0136] In some embodiments, the classification and recognition model includes a multi-layer perceptron, and each multi-layer perceptron includes an input layer, a hidden layer, and an output layer. The recognition module 503 is specifically used for:
[0137] Input the fused features into the hidden layer through the input layer;
[0138] Perform a fully connected process on the fused features through the hidden layer to obtain a connection vector;
[0139] The activation function of the hidden layer is used to judge the connection vector to determine the leakage probability corresponding to the connection vector;
[0140] The output layer determines the recognition result of the detection object based on the leakage probability.
[0141] The fluid leakage detection device provided by the embodiment of the present application and the fluid leakage detection method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0142] The embodiment of the present application also provides an electronic device to execute the above-mentioned fluid leakage detection method. Please refer to Figure 6 It shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 6 shown, the electronic device 7 includes: a processor 700, a memory 701, a bus 702, and a communication interface 703. The processor 700, the communication interface 703, and the memory 701 are connected through the bus 702; a computer program that can run on the processor 700 is stored in the memory 701, and when the processor 700 runs the computer program, it executes the fluid leakage detection method provided by any one of the foregoing embodiments of the present application.
[0143] Among them, the memory 701 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 703 (which can be wired or wireless), the communication connection between the device network element and at least one other network element can be realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0144] The bus 702 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 701 is used to store programs. After the processor 700 receives the execution instruction, it executes the program. The fluid leakage detection method disclosed in any one of the foregoing embodiments of the present application can be applied to the processor 700 or implemented by the processor 700.
[0145] The processor 700 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 700 or the instructions in the form of software. The above-mentioned processor 700 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 701, and the processor 700 reads the information in the memory 701 and combines its hardware to complete the steps of the above method.
[0146] The electronic device provided in the embodiments of the present application and the fluid leakage detection method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.
[0147] The embodiments of the present application also provide a computer-readable storage medium corresponding to the fluid leakage detection method provided in the foregoing embodiments. Please refer to Figure 7 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the fluid leakage detection method provided in any of the foregoing embodiments.
[0148] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0149] The computer-readable storage medium provided by the above embodiments of the present application and the fluid leakage detection method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0150] It should be noted that:
[0151] In the specification provided herein, a large number of specific details are set forth. However, it is understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail in order not to obscure the understanding of this specification.
[0152] Similarly, it should be understood that in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present application.
[0153] In addition, those skilled in the art will appreciate that although some of the embodiments herein include certain features included in other embodiments but not others, the combination of features of different embodiments is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0154] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A fluid leakage detection method, characterized in that, Including: Obtain video data of a detection object within a preset time period, where the detection object is used for storing or transporting fluid; Extract the global long-term dependence features in the video data through a first feature extraction model, where the global long-term dependence features are used to characterize the global state change of the detection object within the preset time period; Extract the local micro-deformation sensitive features in the video data through a second feature extraction model, where the local micro-deformation sensitive features are used to characterize the local state change of the detection object within the preset time period; Fuse the global long-term dependence features and the local micro-deformation sensitive features and then input them into a classification and recognition model for recognition, and output a recognition result, where the recognition result is used to characterize whether fluid leakage occurs in the detection object within the preset time period.
2. The method according to claim 1, wherein The extracting the local micro-deformation sensitive features in the video data through the second feature extraction model includes: Perform differential processing on any two adjacent images in the video data to obtain a differential image; Divide the differential image into multiple sub-differential images; Input the multiple sub-differential images into multiple convolutional units in the second feature extraction model one by one to obtain multiple sub-differential feature images; Stitch the multiple sub-differential feature images to obtain the local micro-deformation sensitive features corresponding to the differential image.
3. The method according to claim 2, wherein The convolutional unit includes multiple serially connected convolutional blocks, and each convolutional block includes multiple serially connected convolutional layers and pooling layers; the inputting the multiple sub-differential images into multiple convolutional units in the second feature extraction model one by one to obtain multiple sub-differential feature images includes: Set corresponding processing parameter values for each convolutional layer and each pooling layer respectively, where the processing parameter values include at least one of the following parameter values: input channel, output channel, feature size, and number of kernels; Input the multiple sub-differential images into multiple convolutional units in the second feature extraction model one by one, and perform serial processing on each convolutional layer and each pooling layer in the convolutional unit according to the corresponding processing parameter values respectively to obtain multiple sub-differential feature images.
4. The method according to claim 2, characterized in that The dividing the differential image into multiple sub-differential images includes: Evenly divide the differential image into multiple sub-differential images along a preset direction.
5. The method according to claim 1, wherein The extracting the global long-term dependence features in the video data through the first feature extraction model includes: Obtain video images of a preset number of frames from the video data; Perform splitting processing on the video images of the preset number of frames to obtain video vectors; Input the video vectors into multiple sequentially connected feature extraction units in the first feature extraction model to obtain the global long-term dependence features corresponding to the video data.
6. The method according to claim 5, wherein The first feature extraction unit includes a linear embedding subunit and a video understanding model, and the first feature unit is any one of the multiple feature extraction units; Perform a fully connected process on the input vector through the linear embedding subunit of the first feature extraction unit to obtain a feature vector; Process the feature vector through the multi-head self-attention mechanism of the video understanding model of the first feature extraction unit to obtain an output feature.
7. The method according to any one of claims 1-6, characterized in that, The first feature extraction model, the second feature extraction model, and the classification and recognition model are pre-trained. The training process includes: Obtain a training set, where the training samples in the training set include multiple video data and the corresponding target recognition results of the target detection objects to which the multiple video data belong; Input the target video data into the first feature extraction model and the second feature extraction model, and respectively output the predicted global long-term dependence features and the predicted local micro-deformation sensitive features corresponding to the target video data. The target video data is any one of the multiple video data; Fuse the predicted global long-term dependence features and the predicted local micro-deformation sensitive features and input them into the classification and recognition model to output the predicted recognition result corresponding to the target detection object to which the target video data belongs; Based on the target recognition result and the predicted recognition result, calculate the loss function value; Based on the loss function value, adjust the parameters of the first feature extraction model, the second feature extraction model, and the classification and recognition model, and continue training until the preset training completion condition is met, to obtain the trained first feature extraction model, the second feature extraction model, and the classification and recognition model.
8. The method according to claim 1, wherein The classification and recognition model includes a multi-layer perceptron. Each multi-layer perceptron includes an input layer, a hidden layer, and an output layer. The step of fusing the global long-term dependence features and the local micro-deformation sensitive features and inputting them into the classification and recognition model for recognition to output the recognition result includes: Input the fused features into the hidden layer through the input layer; Fully connect the fused features through the hidden layer to obtain a connection vector; Judge the connection vector through the activation function of the hidden layer to determine the leakage probability corresponding to the connection vector; Determine the recognition result of the detection object through the output layer based on the leakage probability.
9. A fluid leakage detection device, characterized in that, It includes: An acquisition module, configured to acquire video data of a detection object within a preset time period, where the detection object is used for storing or transporting fluid; An extraction module, configured to extract the global long-term dependence features in the video data through a first feature extraction model, where the global long-term dependence features are used to characterize the global state change of the detection object within the preset time period; The extraction module is further configured to extract the local micro-deformation sensitive features in the video data through a second feature extraction model, where the local micro-deformation sensitive features are used to characterize the local state change of the detection object within the preset time period; A recognition module, configured to fuse the global long-term dependence features and the local micro-deformation sensitive features and input them into a classification and recognition model for recognition to output a recognition result, where the recognition result is used to characterize whether fluid leakage occurs in the detection object within the preset time period.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-7.