Video target detection method and system suitable for industrial fluid liquid level tracking
By combining ResNet50 and Transformer's video object detection methods, the problems of complex environment, scarcity and missed data in industrial liquid level monitoring are solved, and higher detection accuracy and stability are achieved, and suitable for complex dynamic industrial scenarios.
Patent Information
- Application Number
- CN202510118799.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In industrial production, liquid level monitoring faces complex environments, scarce data, insufficient timing information modeling and missed detection problems, resulting in the limitation of the accuracy and stability of traditional methods in dynamic industrial scenarios.
A video object detection method is adopted, combining ResNet50 convolutional neural network and Transformer attention mechanism, a query mechanism based on object category and front frame position prediction is designed, and a long-term feature memory module and Kalman filter are introduced to improve the accuracy and stability of detection through multiple rounds of training and data augmentation, and the missed detection problem is alleviated through a slow exit mechanism.
It significantly improves the detection accuracy and stability of industrial fluid level tracking, effectively alleviates the interference of missed detection and misdetection on liquid level tracking, and ensures the robustness of the system in complex scenarios.
Smart Images

Figure CN120107846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video target recognition of computer vision, and in particular to a video target detection method and system suitable for industrial fluid level tracking. Background Art
[0002] In the industrial production process, liquid level monitoring is an important basic task. It is of great significance to accurately track and monitor the liquid level in real time, especially in pipeline transportation and tank management. Traditional liquid level detection methods usually rely on sensors such as float level gauges, capacitive level gauges, and ultrasonic level gauges. Although these methods can provide relatively accurate measurement data in static environments, in dynamic industrial scenarios, such as rapid fluid flow, pipeline vibration, light interference or other complex environments, the accuracy and stability of sensors are easily limited.
[0003] In recent years, with the development of computer vision technology, video target detection technology has gradually been applied to the field of industrial liquid level tracking, and non-contact liquid level tracking is achieved through image processing and target detection algorithms. However, fluid monitoring in industrial scenarios faces many challenges, including:
[0004] Environmental complexity: Industrial sites may have complex lighting conditions, reflection interference, and equipment occlusion, which may affect the accuracy of video target detection.
[0005] Data scarcity: It is difficult to obtain labeled data for industrial liquid level tracking scenarios. Insufficient data leads to poor generalization ability of trained detection models.
[0006] Insufficient modeling of temporal information: The dynamic changes and continuity of fluids require that video target detection models be able to make full use of temporal information. However, existing methods usually only focus on single-frame detection and lack effective use of contextual information, resulting in instability in detection results.
[0007] Missed detection problem: Since target detection models are prone to missed detection in complex industrial scenarios, it may cause intermittent liquid level tracking or detection failure, thus affecting the stability and reliability of the system. Summary of the invention
[0008] In order to solve the above problems, this paper proposes a video target detection method and system suitable for industrial fluid level tracking.
[0009] The object of the present invention is achieved through the following technical solutions: A video target detection method suitable for industrial fluid level tracking, comprising:
[0010] S1, obtain the data set and perform annotation and preprocessing;
[0011] S2. Construct an industrial fluid level target detection network, including a feature extraction network, a long-term feature memory module and a target detection network, and use a Kalman filter to predict the next frame after target detection, wherein the query in the target detection network is a query based on the object category and a query based on the previous frame position prediction obtained by the Kalman filter;
[0012] S3, using the pre-processed real data set and the mixed data set to perform multiple rounds of training on the network to obtain a trained industrial fluid level target detection network;
[0013] S4. Use the trained target detection model and Kalman filter to perform target detection and next frame prediction on the pipes and liquid parts in the video; introduce a slow exit mechanism for targets that do not appear, and exit when the linearly predicted fluid exits the screen or is not detected for a long time.
[0014] Furthermore, the acquisition of the data set includes: acquiring real data and pseudo data, extracting key frames of the real data, randomly inserting pseudo data into the real data set by inserting frames, and generating a mixed data set;
[0015] The real data is video data of some time periods of some points to be detected acquired at industrial sites, which is obtained by annotating the collected videos. The pseudo data is generated through network search or computer simulation. The computer simulation uses computer simulation technology to generate simulated pictures that are consistent with the actual pipeline diameter and fluid flow characteristics.
[0016] Furthermore, the data preprocessing includes: performing a size normalization operation on all data, and using a basic data enhancement operation for different image data in the same video sequence.
[0017] Furthermore, the feature extraction network specifically uses the residual module in the ResNet50 network as the basic module, uses its pre-trained weights on Imagenet for initialization, and uses skip connections to solve the gradient vanishing problem in deep neural network training.
[0018] Furthermore, the long-term feature memory module is composed of the feature maps of the backbone of the previous T frames, and is a fixed first-in-first-out queue structure that stores the feature maps of the previous T frames. In long-term feature processing, all data and the current feature map are globally pooled, dimensionally reduced, and max-pooled, and several frames of feature maps with the smallest similarity to the current feature are selected through similarity measurement, and the corresponding features are aggregated into the current feature map.
[0019] Furthermore, the feature aggregation includes: performing maximum pooling on the stacked features, using a cross-attention mechanism to fuse the features of the current frame and the historical frames, and then fusing the current feature map, the stacked feature map after maximum pooling, and the feature map after attention set to obtain the final aggregated feature.
[0020] Furthermore, in the target detection network, the encoder module of Transformer is used to globally model the fused feature map, and the query part is obtained by fusing the object query based on the object category and the query based on the position prediction of the previous frame; wherein the object query based on the object category represents the detection target of the corresponding category, and the category label value corresponding to the embedding is used as the query value for initialization during the initialization process;
[0021] The query based on the previous frame position prediction is specifically: generating a corresponding probability map for the upper left corner and the lower right corner of the target box in the Kalman filter prediction result in the previous frame, and obtaining a prediction query through an encoder;
[0022] Fusion formula includes: Q final =β·Q cls +(1-β)·Q pred
[0023] Where Q final is the query obtained by the final fusion, Q cls is the object query based on the object category; Q pred is the query based on the previous frame position prediction, and β is the weight coefficient.
[0024] Furthermore, a slow exit mechanism is introduced for non-appearing targets, which will be exited only when the linear prediction fluid exits the screen or is not detected for a long time; the specific steps are:
[0025] For undetected objects, the Kalman filter is used to find out their corresponding coordinate information and motion information, and time series prediction is performed based on the coordinate information and motion information. If the target area is close to 0 or the predicted target position is close to the edge of the screen, it can be confirmed that it is not detected; otherwise, it is considered to be blocked or missed within a time series of length t, and the parameters in the filter are used as the motion standard. If it is still not detected after t time, it is considered to have disappeared.
[0026] On the other hand, the present invention also provides a video target detection system suitable for industrial fluid level tracking, comprising:
[0027] Data processing module, used to obtain real data and mixed data, and to label and preprocess the data;
[0028] The model building module uses part of the Resnet50 convolutional neural network to build the backbone, uses the encoder-decoder module of the Transformer self-attention mechanism structure to build the target recognition core module, uses queries based on object categories and previous frame position predictions to replace the original location-based queries, and introduces a long-term feature memory module in the network to provide richer contextual information for target detection;
[0029] Model training module, which initializes the model, sets corresponding parameters and performs training;
[0030] The post-processing module performs post-processing on the undetected object situations through the slow exit mechanism.
[0031] On the other hand, the present invention specification also provides a video target detection device suitable for industrial fluid level tracking, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it implements the video target detection method suitable for industrial fluid level tracking.
[0032] Beneficial effects of the present invention: The innovation of the present invention is that it targets the needs of industrial scenarios, combines the ResNet50 convolutional neural network and the Transformer attention mechanism, designs a query mechanism based on object category and previous frame position prediction, and detects and tracks fluids in combination with a long-term feature memory module, which significantly improves the accuracy and stability of detection. In addition, by introducing a slow exit mechanism, it can effectively alleviate the interference of missed detections and false detections on liquid level tracking, ensuring the robustness of the system in complex scenarios. At the same time, through targeted solutions to the problems of data scarcity and missed detections in complex scenarios, it demonstrates its wide application potential and practical value.
[0033] In the field of industrial fluid level tracking, the present invention has significant technical advantages and economic benefits compared to traditional sensor-based monitoring technology, and provides a new solution for level monitoring in complex dynamic scenarios, with high application prospects and industrialization potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A structural diagram of a video object detection network model provided by an embodiment of the present invention;
[0035] Figure 2 A detailed structural diagram of a long-term feature memory module provided in an embodiment of the present invention;
[0036] Figure 3 A schematic diagram of a slow exit mechanism introduced in an embodiment of the present invention;
[0037] Figure 4A schematic diagram of a video target detection device suitable for industrial fluid level tracking provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0039] Combine the following Figure 1-Figure 3 The present invention describes a target detection method and system for industrial fluid level tracking scenarios.
[0040] Example
[0041] S1. Dataset construction and annotation. The dataset includes real data acquisition and pseudo data. The real data is the video data of some time periods at some points to be detected acquired at the industrial site. By deploying cameras to collect image data of liquid flow in industrial pipelines and annotating the collected videos, a basic real dataset is formed. At the same time, in order to solve the problem of sample scarcity in industrial data scenarios, a pseudo dataset is generated.
[0042] There are two methods for constructing pseudo data sets: Obtained through web search: Use web search to obtain relevant images of pipes and fluids from public images in the field of industrial fluids, such as pipe fluid images obtained through web crawlers and fluid images under the same pipe diameter obtained through computer simulation. Computer simulation generation: Use computer simulation technology to generate simulated images consistent with the actual pipe diameter and fluid flow characteristics to supplement data diversity. In order to enhance the robustness and accuracy of detection, the generated pseudo data is not directly used for training. Instead, the key frames of the real data are extracted from the real data, and the pseudo data is randomly inserted into the real data set by interpolation to generate a mixed data set. Key frame extraction can optimize the redundant frames in the video, shorten the video length using key frame extraction technology, and retain key action frames to reduce the waste of computing resources, especially the invalid reasoning training caused by the lack of changes in short-term video data in industrial scenarios.
[0043] The data is preprocessed and resized, and the resolution of all data is adjusted to a uniform size of 1024×1024 to ensure the consistency of input data during model training. Data augmentation is used on the dataset to expand the generalization of the data. For different image data in the same video sequence, basic data augmentation operations are used, such as color transformation (such as brightness and contrast adjustment), color noise injection, image horizontal flipping, color jittering, etc. Through the above data augmentation operations, a small range of noise is injected to improve the robustness of the model to different environments. For different video sequences, more powerful data augmentation methods can be adopted.
[0044] S2. Build an industrial fluid level target detection network. This method uses part of the ResNet50 network as a feature extraction and downsampling network, uses its pre-trained weights on Imagenet for initialization, uses the encoder and decoder of the transformer architecture for target recognition, and uses the Kalman filter to predict the next frame after target detection. The residual module in the ResNet50 network is mainly used as the basic module, and the jump connection is used to solve the gradient disappearance problem in the deep neural network training. The specific method is as follows:
[0045] S21. Select part of ResNet50 as the backbone to extract features and perform downsampling. Use the weights of ResNet50 on Imagenet as pre-training parameters. Input the image to generate a feature map through the backbone. Send the feature map after long-term feature processing to the Transformer structure encoder for encoding and further feature extraction.
[0046] S22. After obtaining the feature map through the backbone, the long-term feature memory module is processed. The long-term memory feature module is composed of the feature map of the backbone of the previous T frames, and is a fixed first-in-first-out queue structure that stores the feature map of the previous T frames. In the long-term feature processing, all data and the current feature map are globally pooled. Through the similarity measurement, several frames of feature maps with the smallest similarity to the current feature can be selected, and the corresponding features are aggregated to the current feature map through the feature aggregation module, introducing richer context information and timing information to the feature map;
[0047] S23. For the query part, a new query is composed of queries based on object categories and queries based on the previous frame position prediction. For object queries based on object categories, it is defined as a set of learnable vectors related to the target category. Each object query represents a query of the corresponding category and position. The initial vector is obtained by random initialization + embedding operation of the corresponding number of the category label during initialization. For queries based on the previous frame position prediction, in the video sequence, it can be assumed that the movement of the target is linear between frames. The Kalman filter and the Hungarian matching algorithm can provide the possible future position information for the object predicted in the previous frame. The corresponding heat map can be generated by the Gaussian function, and the predicted query can be generated by a simple encoder structure. The two queries are combined to generate the real query. In order to ensure the convergence of the training, the real label data of the previous frame with mixed Gaussian noise is added to the query based on the previous frame position prediction in the early training, that is, the heat map of the prediction result and the result of the mixed Gaussian noise heat map are used as the next frame prediction query.
[0048] S24. Generate the result graph through the multi-layer decoder.
[0049] Preferably, the specific implementation process in this embodiment includes:
[0050] The long-term memory module is used to enhance the temporal context information of target detection on the feature map generated by the Resnet50 network. The process is as follows:
[0051] For the long-term feature memory module, a time constant T is defined, representing the previous T frames. First, the feature maps extracted by Backbone from the previous T frames are stored in a fixed-length first-in-first-out (FIFO) queue:
[0052] M t =[F t-1 ,F t-2 ,…,F t-T ]
[0053] Among them, M t is the feature map set stored in the long-term memory module, F t-i Represents the feature map of the previous i-th frame.
[0054] like Figure 2 As shown in the figure, for the feature maps in the queue, all feature maps are first subjected to adaptive pooling, and then dimension reduction and maximum pooling are performed. For the features of the i-th frame,
[0055] g(F t-i ) = F′ t-i =MaxPolling(Reduce dim(AdaptivePooling(F t-i ) i ))
[0056] After the above treatment, F′ t-i From BXCXWXH to BXCX1X1, it can be transformed into BXC through dimensionality reduction.
[0057] The processed feature graph is concat-operated, and the similarity between the current feature graph and the historical feature graph is calculated through similarity measurement. The similarity calculation uses the cosine function as the distance measurement function.
[0058]
[0059] Where <·,·> represents the inner product of the eigenvector, and ||·|| represents the L2 norm of the vector;
[0060] The similarity between the current feature map and the historical feature map can be obtained by similarity measurement, and the index of the 3 frames and the current feature map can be found from them.
[0061]
[0062] in is the feature vector after long-term feature memory processing
[0063] The three feature maps with the largest gap with the current feature are obtained through similarity measurement, and the selected feature maps are stacked through concat.
[0064] The historical features and current features are fused by feature aggregation to make up for the defects of the current feature map. Here, feature aggregation is composed of parallel maximum pooling structure, deep separable convolution, cross attention mechanism and jump connection, such as Figure 2 As shown in
[0065] First, the stacked feature F stacked Perform maximum pooling to extract salient features of different regions
[0066] F pool =MaxPooling(F stacked )
[0067] F stacked Perform a depth-separable convolution operation so that the number of channels corresponds to the feature map. The cross-attention mechanism fuses the features of the current frame and the historical frame. Q, V are the current frame features, and K is the historical frame stacked feature F stacked After the output of the depth-wise separable convolution, the attention aggregation formula is:
[0068]
[0069] The features output by the above modules are combined to obtain
[0070] F fused =F t +F atten +F pool
[0071] The encoder module of Transformer is used to further globally model the fused feature map. Through the self-attention mechanism, the correlation information between distant pixels in the image can be captured. Assume that the feature map input is F fused , and its self-attention calculation formula is as follows:
[0072]
[0073] Where Q, K, and V are query, key, and value respectively. k is the scaling factor of the feature dimension.
[0074] In order to further improve the temporal modeling capability of the target detection model in industrial liquid level tracking, the present invention improves the query part in the target detection network. By using object queries based on object categories and queries based on previous frame position prediction, the robustness and temporal consistency of the model for target detection are significantly enhanced.
[0075] Object query Q based on object category cls It is a set of learnable vectors related to the target category. Each query represents the detection target of the corresponding category. In this way, the network can explicitly learn the specific representation of each category. The object query is parameterized so that it converges to a suitable value during the training process. The initialized query is composed of a combination of random initialization and the category label value corresponding to the embedding. That is, the category label value corresponding to the embedding is used as the query value for initialization in some dimensions of the query.
[0076] In a video sequence, the motion of a target can be approximately assumed to be a linear movement between frames. Therefore, the target position of the previous frame can be used to predict the target position of the next frame, and a corresponding query can be generated. The present invention uses a Kalman filter to estimate the motion state, which includes the following steps:
[0077] The working principle of Kalman filter is as follows: it includes two parts: prediction and update.
[0078] Prediction Phase
[0079]
[0080] in, is the predicted current state vector, F k is the transfer matrix, which describes how the system evolves from time k-1 to time k. Through the above process, the predicted state at the current moment can be obtained. k|k-1 is the prediction error covariance matrix of the current state, indicating the uncertainty of the prediction value. k|k-1 is the state error at the previous moment, Q k is the process noise covariance, which indicates the uncertainty in the system model. This formula updates the prediction error covariance, taking into account the uncertainty in state transition and the influence of process noise. In the present invention, the target position information of the next frame image can be obtained by predicting the result of the previous frame image, and then feedback is performed through query.
[0081] renew:
[0082]
[0083] P k|k =(IK k H k ) k|k-1
[0084] Where K k is the Kalman gain, which measures the impact of the measurement value on the state estimation. k is the observation matrix, describing how the state is mapped to the measurement space. k is the measurement noise covariance, which represents the measurement uncertainty. k is the measured value at the current moment, The prediction state is mapped to the measurement space. is the state vector after the predicted value and the observed value are balanced. k|k is the updated error covariance matrix. By updating the covariance and correcting the uncertainty of the state in combination with the measured value, the covariance can be gradually converged.
[0085] In the present invention, we define the state vector as follows:
[0086] For detected objects, the general output is the position coordinates and the corresponding width and height as well as the category. Therefore, the position coordinates and the corresponding width and height of the target object are used as state variables:
[0087]
[0088] where x c (k),y c (k) is the current object center coordinate, w(k), h(k) are the width and height of the corresponding object, is the first-order derivative of the corresponding variable, that is, the rate.
[0089] Similarly, the corresponding observed variable is defined as:
[0090]
[0091] Noise Matrix Q k ,R k Use the following definition
[0092]
[0093] where σ p , σ v , σ m is the corresponding process noise factor and measurement noise factor, which can be initialized according to the actual operation conditions. The initialization setting σ p =σ m , σ v =0.1σp More reasonable.
[0094] According to the above definitions of state vector and observation variables, the corresponding transfer matrix F and observation matrix H are reasonably deduced.
[0095]
[0096] The Kalman filter model can be constructed through the above method. During the training process, the model finally outputs the location and category information of the current target. The matching status of the detected target result and the current prediction model can be obtained through Hungarian matching. For targets that cannot be matched, a corresponding Kalman filter can be reinitialized.
[0097] The Kalman filter is used to predict the position of the detected target in the next frame, and the post-processing mechanism is used to remove objects that are out of the picture. The Kalman filter prediction result is converted into a probability map
[0098] The coordinate information can be used to generate a probability map through the Gaussian kernel function. The generating function is as follows:
[0099]
[0100] where x c ,y c is the predicted target center, is the variance of the Gaussian distribution, which is proportional to the width w and height h of the target box. G(x,y) is the probability value of each pixel position in the image space.
[0101] Similarly, the above operation can be used to generate corresponding probability maps for the upper left and lower right corners of the target box. 2 Determined by the aspect ratio and the target box area.
[0102] The above operations can generate a prediction probability map, and a simple encoding operation structure can generate the corresponding prediction queryQ pre d;
[0103] Finally, the fusion query is generated, and the fusion formula is:
[0104] Q final =β·Q cls +(1-β)·Q pred ,
[0105] Among them, β is the weight coefficient, which is used to control the relative importance of category query and prediction query.
[0106] S3. Use the preprocessed real data set and mixed data set to train the network, and obtain the target detection network model through multiple rounds of training.
[0107] Through Q based on the last fused query final Perform multi-layer decoder operations to obtain the final result. And continue to iterate training; during the training process, in order to ensure the validity of the query, the predicted probability map generated in the early stage is integrated with the low-noise processed object position probability map of the previous frame. When the training gradually converges, the real noise probability map of the object position of the previous frame is no longer added.
[0108] Specifically, continuous training is performed on a server that meets the hardware requirements, generally on an RTX4090, with 200 training rounds, and validation is performed on the validation set every 40 rounds and the corresponding model weight file is saved. The initial learning rate is set to 3e-3, and the cosine annealing adjustment strategy is gradually reduced. At the same time, the AdamW optimizer is used to accelerate convergence and prevent overfitting. After training, the model with better performance on the validation set is selected as the recognition model.
[0109] S4. By detecting the pipes and fluids in the video sequence, the corresponding liquid level can be calculated and tracked. We introduced a slow exit mechanism in the use of model detection. Figure 3 Explain the slow exit mechanism
[0110] In the detection of long video sequences, noise disturbance may cause missed detection. Therefore, we add a Kalman filter to the detection process, and its specific working conditions are as described above. For undetected objects, the Kalman filter is used to predict the next movement of the target. It will exit only when the area of the next target approaches 0 or its target frame approaches the edge of the screen. Otherwise, it is considered to be blocked or missed in the time sequence of length t, and the parameters in the filter are used as the motion standard and predicted. If it is still not detected after t time, it is considered to have disappeared.
[0111] For the above network structure, the depth of the backbone and the time parameter T of the long-term feature memory module can be adjusted according to actual requirements to reduce the amount of calculation and parameters, thereby achieving lightweight network.
[0112] Corresponding to the aforementioned embodiment of a video target detection method applicable to industrial fluid level tracking, the present invention also provides an embodiment of a video target detection device applicable to industrial fluid level tracking.
[0113] See also Figure 4 A video target detection device suitable for industrial fluid level tracking provided in an embodiment of the present invention includes a memory and one or more processors. The memory stores executable code. When the processor executes the executable code, it is used to implement a video target detection method suitable for industrial fluid level tracking in the above embodiment.
[0114] An embodiment of a video target detection device suitable for industrial fluid level tracking provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 4 As shown, it is a hardware structure diagram of a video target detection device for industrial fluid level tracking provided by the present invention, in which any device with data processing capability is located, except Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0115] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0116] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0117] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a video target detection method suitable for industrial fluid level tracking in the above embodiment is implemented.
[0118] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0119] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the video target detection method suitable for industrial fluid level tracking is implemented.
[0120] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0121] It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. The present application is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the attached claims.
Claims
1. A video target detection method suitable for industrial fluid level tracking, characterized in that: include: S1, obtain the data set and perform annotation and preprocessing; S2. Construct an industrial fluid level target detection network, including a feature extraction network, a long-term feature memory module and a target detection network, and use a Kalman filter to predict the next frame after target detection, wherein the query in the target detection network is a query based on the object category and a query based on the previous frame position prediction obtained by the Kalman filter; S3, using the pre-processed real data set and the mixed data set to perform multiple rounds of training on the network to obtain a trained industrial fluid level target detection network; S4, using the trained target detection model and Kalman filter to perform target detection and next frame prediction on the pipe and liquid parts in the video; A slow exit mechanism is introduced for targets that do not appear, which exits when the linear prediction fluid exits the screen or is not detected for a long time.
2. A video target detection method suitable for industrial fluid level tracking according to claim 1, characterized in that: The acquisition of the data set includes: acquiring real data and pseudo data, extracting key frames of the real data, randomly inserting pseudo data into the real data set by inserting frames, and generating a mixed data set; The real data is video data of some time periods of some points to be detected acquired at industrial sites, which is obtained by annotating the collected videos. The pseudo data is generated through network search or computer simulation. The computer simulation uses computer simulation technology to generate simulated pictures that are consistent with the actual pipeline diameter and fluid flow characteristics.
3. The video target detection method for industrial fluid level tracking according to claim 1, characterized in that: The data preprocessing includes: performing a size normalization operation on all data, and using a basic data enhancement operation for different image data in the same video sequence.
4. The video target detection method for industrial fluid level tracking according to claim 1, characterized in that: The feature extraction network specifically uses the residual module in the ResNet50 network as the basic module, uses its pre-trained weights on Imagenet for initialization, and uses skip connections to solve the gradient vanishing problem in deep neural network training.
5. The video target detection method suitable for industrial fluid level tracking according to claim 1, characterized in that: The long-term feature memory module is composed of the feature maps of the backbone of the previous T frames, and is a fixed first-in-first-out queue structure that stores the feature maps of the previous T frames. In long-term feature processing, all data are globally pooled, dimensionally reduced, and max-pooled with the current feature map. Several frames of feature maps with the smallest similarity to the current feature are selected through similarity measurement, and the corresponding features are aggregated into the current feature map.
6. A video target detection method suitable for industrial fluid level tracking according to claim 5, characterized in that: The feature aggregation includes: performing maximum pooling on the stacked features, using a cross-attention mechanism to fuse the features of the current frame and the historical frames, and then fusing the current feature map, the stacked feature map after maximum pooling, and the feature map after attention collection to obtain the final aggregated feature.
7. The video target detection method for industrial fluid level tracking according to claim 1, characterized in that: The target detection network uses a Transformer encoder module to globally model the fused feature map, and the query part is obtained by fusing the object query based on the object category and the query based on the previous frame position prediction; The object query based on the object category represents the detection target of the corresponding category. During the initialization process, the category label value corresponding to the embedding is used as the query value for initialization; The query based on the previous frame position prediction is specifically: generating a corresponding probability map for the upper left corner and the lower right corner of the target box in the Kalman filter prediction result in the previous frame, and obtaining a prediction query through an encoder; Fusion formula includes: Q final =β·Q cls +(1-β)·Q pred Where Q final is the query obtained by the final fusion, Q cls is the object query based on the object category; Q pred is the query based on the previous frame position prediction, and β is the weight coefficient.
8. The video target detection method for industrial fluid level tracking according to claim 1, characterized in that: For targets that do not appear, a slow exit mechanism is introduced, which will be used when the linear prediction fluid exits the screen or is not detected for a long time. The specific steps are as follows: For undetected objects, the Kalman filter is used to find out their corresponding coordinate information and motion information, and time series prediction is performed based on the coordinate information and motion information. If the target area is close to 0 or the predicted target position is close to the edge of the screen, it is confirmed that it is not detected; Otherwise, it is considered to be blocked or missed within a time series of length t, and the parameters in the filter are used as the motion standard. If it is still not detected after t time, it is considered to have disappeared.
9. A video target detection system suitable for industrial fluid level tracking for implementing the method according to any one of claims 1 to 8, characterized in that: include: Data processing module, used to obtain real data and mixed data, and to label and preprocess the data; The model building module uses part of the Resnet50 convolutional neural network to build the backbone, uses the encoder-decoder module of the Transformer self-attention mechanism structure to build the target recognition core module, uses queries based on object categories and previous frame position predictions to replace the original location-based queries, and introduces a long-term feature memory module in the network to provide richer contextual information for target detection; Model training module, which initializes the model, sets corresponding parameters and performs training; The post-processing module performs post-processing on the undetected object situations through the slow exit mechanism.
10. A video target detection device suitable for industrial fluid level tracking, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a video target detection method suitable for industrial fluid level tracking according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Weak and small airspace target detection method based on super-resolution feature enhancement
CN113223059A
Target tracking algorithm based on attention mechanism
CN114463375A
Key video data extraction method based on multi-dimensional semantic information
WO2024109308A1