Metering field operation management method based on machine vision

Through machine vision technology, real-time monitoring and object detection of the metrology site is solved, and the problem of inability to monitor and track accurately in traditional methods is achieved, efficient identification and management of on-site dynamics is achieved, and management efficiency and safety of the metrology site are improved.

CN120259964APending Publication Date: 2025-07-04MARKETING SERVICE CENT OF STATE GRID HEILONGJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315873.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional methods cannot monitor dynamic changes in the metrology site in real time, cannot efficiently and accurately identify and track multiple targets, there is a risk of missed and missed detection, and cannot effectively analyze the interactive behavior between targets. Background information interference affects the monitoring effect, resulting in inefficient management.

Method used

The measurement field operation management method based on machine vision is used to obtain video frame sequences through camera monitoring, perform background frame filtering, and use pre-trained object detection model, Kalman filtering, and Hungarian matching algorithm to track targets and judge interaction behaviors, and extract keyframes and send them to the management equipment.

Benefits of technology

It realizes real-time acquisition of on-site dynamic information, accurately identifying metrological terminal equipment and staff, improves target tracking accuracy, identify interactive behaviors, optimizes management decisions, reduces human errors, and improves on-site operation management efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259964A_ABST
    Figure CN120259964A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of field operation monitoring, in particular to a metering field operation management method based on machine vision, which comprises the following steps: carrying out video monitoring on a metering field based on a preset camera to obtain a metering field video, and obtaining a metering field video frame sequence of the metering field video, performing background frame filtering on the metering field video frame sequence to obtain a foreground video frame sequence; and performing target detection on each foreground video frame in the foreground video frame sequence based on a pre-trained target detection model to obtain a target detection result. According to the invention, through real-time video monitoring of a metering site and subsequent target detection and tracking, dynamic information of the site can be obtained in real time, and the target detection and target tracking technology can ensure that position information and motion tracks of metering terminal equipment and workers can be captured in time. And comprehensive understanding and effective management of field operation conditions are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of on-site operation monitoring, and particularly relates to a metering on-site operation management method based on machine vision. Background Art

[0002] Traditional methods usually rely on manual inspections or regular manual checks, and cannot monitor all dynamic changes on-site in real time. Such methods cannot capture the real-time information of the metering site in a timely manner, resulting in the inability to respond quickly to emergencies; moreover, traditional methods may only rely on manual records or simple fixed monitoring devices, and cannot efficiently and accurately identify and track multiple targets on-site. Especially in the case of limited human resources, there is often a risk of missed inspections and misidentifications; furthermore, traditional methods cannot effectively analyze the interaction behaviors between targets, such as the operations or interactions between staff and equipment. Usually relying on manual observation and records, important behavior patterns between targets may be missed, leading to low decision-making efficiency; and the monitoring videos of traditional methods often contain a large amount of irrelevant background information, making it difficult for staff to concentrate on important targets. The interference of background information will affect the monitoring effect, making it difficult for managers to discover abnormal situations in a timely manner. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a metering on-site operation management method based on machine vision.

[0004] The technical solution adopted to solve the above technical problem is: A metering on-site operation management method based on machine vision, including:

[0005] Performing video monitoring on the metering site based on a preset camera to obtain a metering site video, acquiring a metering site video frame sequence of the metering site video, and filtering background frames from the metering site video frame sequence to obtain a foreground video frame sequence;

[0006] Performing target detection on each foreground video frame in the foreground video frame sequence based on a pre-trained target detection model to obtain a target detection result, and performing target tracking on the detected targets through the S0RT tracking algorithm according to the target detection result to obtain a first target tracking result;

[0007] Performing tracking prediction on the detected targets through Kalman filtering according to the first target tracking result to obtain a second target tracking result, and associating and matching the first target tracking result and the second target tracking result based on the Hungarian matching algorithm to obtain a third target tracking result;

[0008] Based on the third target tracking result, determine whether there is an interaction behavior between the detected targets. If there is an interaction behavior between the detected targets, obtain the subsequence of foreground video frames corresponding to the detected targets with interaction behavior.

[0009] Extract key frames from the subsequence of foreground video frames to obtain the key video frames of the detected targets with interaction behavior, and send the key video frames of the detected targets with interaction behavior to the intelligent device of the management personnel.

[0010] Preferably, filtering the background frames of the metering site video frame sequence to obtain a foreground video frame sequence, including:

[0011] Traverse the metering site video frame sequence, and model the background of the initial metering site video frame in the metering site video frame sequence based on the Gaussian distribution to obtain the background model corresponding to the initial metering site video frame.

[0012] When the current metering site video frame is traversed, use the background model to compare each pixel point of the current metering site video frame one by one to classify the pixel points of the current metering site video frame to determine the type of the pixel points, where the type of the pixel points includes background points and foreground points.

[0013] Count the number of foreground points determined in the current metering site video frame to obtain the number of foreground points, and calculate the retention coefficient of the current metering site video frame based on the number of foreground points to obtain the retention coefficient of the current metering site video frame.

[0014] Compare the retention coefficient of the current metering site video frame with a preset retention threshold. If the retention coefficient of the current metering site video frame is greater than the preset retention threshold, add the current metering site video frame to the foreground video frame sequence.

[0015] Update the background model based on the current metering site video frame to obtain a new background model, and repeat the above operations until all video frames in the metering site video frame sequence are traversed to obtain the foreground video frame sequence corresponding to the metering site video frame sequence.

[0016] Preferably, the background model is as follows:

[0017]

[0018] where x j,t represents the pixel value of the j-th pixel point at time t in the metering site video frame, P represents the background distribution of the pixel point, represents the weight value of the i-th Gaussian distribution at time t in the background model. represents the mean value of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame, represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame, where, and represents the average pixel value of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space, where, and represents the standard deviation of the pixel values of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space, η represents the probability density function of the Gaussian distribution, where,

[0019]

[0020] The calculation formula of the retention coefficient is as follows:

[0021]

[0022] where, τ represents the retention coefficient, P represents the background distribution of the pixel, and m*n represents the total number of pixels in the metering site video frame;

[0023] The update formula of the background model is as follows:

[0024]

[0025] where, M i,t represents whether the pixel is a foreground point, M i,t =1 indicates that the pixel is a foreground point, M i,t =0 indicates that the pixel is a background point, X t represents the pixel value of the pixel, and α and ρ represent preset weight coefficients.

[0026] Preferably, the target detection model includes an appearance feature extraction module and a classification module. The appearance feature extraction module sequentially includes 5 feature extraction blocks and 4 feature fusion blocks. Among them, the feature extraction block is sequentially composed of a convolutional layer, a batch normalization layer, a ReLU activation function, and a downsampling operation. The multi-level feature fusion block includes a convolutional layer and a feature fusion layer. The feature fusion layer consists of two parallel branches. Among them, the first parallel branch is composed of a global average pooling layer, a convolutional layer, a ReLU activation function, and a convolutional layer. The second parallel branch is composed of a convolutional layer, a ReLU activation function, and a convolutional layer. The classification module sequentially includes a global average pooling layer, a fully connected layer, and a Softmax activation function.

[0027] Preferably, the five feature extraction blocks sequentially perform shallow feature extraction on the foreground video frame to obtain a first shallow feature, a second shallow feature, a third shallow feature, a fourth shallow feature, and a fifth shallow feature. The multi-level feature fusion block is used to perform a convolution operation on the first shallow feature, the second shallow feature, the third shallow feature, the fourth shallow feature, and the fifth shallow feature based on a convolutional layer to obtain a first global shallow feature, a second global shallow feature, a third global shallow feature, a fourth global shallow feature, and a fifth global shallow feature, and perform a preliminary fusion on the first global shallow feature and the second global shallow feature through the feature fusion layer to obtain a first shallow fusion feature, perform a preliminary fusion on the first shallow fusion feature and the third global shallow feature through the feature fusion layer to obtain a second shallow fusion feature, perform a preliminary fusion on the second shallow fusion feature and the fourth global shallow feature through the feature fusion layer to obtain a third shallow fusion feature, and perform a preliminary fusion on the third shallow fusion feature and the fifth global shallow feature through the feature fusion layer to obtain an appearance feature.

[0028] Preferably, the expression of the object detection result is as follows:

[0029] R = {r1, r2,..., r t};

[0030] where R represents the object detection result, and r t represents the detected object at time t, where, represents the jth detected object at time t, where, represents the category of the jth detected object at time t, and the categories of the detected objects include metering terminal devices and staff members, represents the recognition probability of the jth detected object at time t, represents the upper left corner coordinates of the detection box of the jth detected object at time t, and represent the width and height of the detection box of the jth detected object at time t;

[0031] The expression of the first object tracking result is as follows:

[0032] OT = {ot1, ot2,..., ot t};

[0033] where OT represents the first object tracking result, and ot t represents the detected object tracked at time t, where, Denote the j-th detected target tracked by the target at time t, where, Denote the tracking number of the j-th detected target tracked by the target at time t;

[0034] The expression of the second target tracking result is as follows:

[0035] UT = {ut2,..., ut t+1};

[0036] Where, UT represents the second target tracking result, and ut t+1 Denote the detected target tracked by the Kalman filter at time t+1, where, Denote the j-th detected target tracked by the Kalman filter at time t+1, where,

[0037] Preferably, judging whether there is an interaction behavior between the detected targets based on the third target tracking result includes:

[0038] Calculating the area of the first detection frame of the detected target of the category of metering terminal equipment at adjacent moments according to the third target tracking result;

[0039] Calculating the overlapping area of the first detection frame of the detected target of the category of metering terminal equipment at adjacent moments according to the third target tracking result;

[0040] Calculating the first detection coincidence degree based on the area of the first detection frame and the overlapping area of the first detection frame at the adjacent moments, and calculating the first spatio-temporal feature of the detected target of the category of metering terminal equipment based on the first detection coincidence degree;

[0041] If the first spatio-temporal feature of the detected target of the category of metering terminal equipment is greater than a preset first threshold, it is judged that the metering terminal equipment is in a stationary state, otherwise, it is judged that the metering terminal equipment is in a non-stationary state;

[0042] Calculating the area of the second detection frame of the detected target of the category of staff according to the third target tracking result;

[0043] Calculating the overlapping area of the second detection frame of the detected target of the category of metering terminal equipment and staff according to the third target tracking result;

[0044] Calculating the second detection coincidence degree of the metering terminal equipment and the staff based on the area of the second detection frame and the overlapping area of the second detection frame, and calculating the second spatio-temporal feature of the metering terminal equipment and the staff based on the second detection coincidence degree;

[0045] If the second spatio-temporal feature of the metering terminal device and the staff is greater than a preset second threshold, then determine the non-contact position state between the metering terminal device and the staff; otherwise, determine the contact position state between the metering terminal device and the staff;

[0046] If the metering terminal device is in a stationary state and the contact position state exists between the metering terminal device and the staff, then determine that there is an interaction behavior between the detection targets; otherwise, determine that there is no interaction behavior between the detection targets.

[0047] Preferably, the calculation formula of the first detection coincidence degree is as follows:

[0048]

[0049] Wherein, represents the first detection coincidence degree between the (t - 1)th moment and the tth moment, S(k1) and S(k2) represent the first detection frame areas at the (t - 1)th moment and the tth moment, and S(O) represents the first detection frame overlapping area at the (t - 1)th moment and the tth moment;

[0050] The calculation formula of the first spatio-temporal feature is as follows:

[0051]

[0052] Wherein, R k (t) represents the first spatio-temporal feature at the tth moment, and t s represents the first target detection moment of the metering terminal device;

[0053] The calculation formula of the second detection coincidence degree is as follows:

[0054]

[0055] Wherein, represents the second detection coincidence degree of the metering terminal device and the staff at the tth moment, S(a) represents the second detection frame overlapping area at the tth moment, and S(i) represents the second detection frame area at the tth moment;

[0056] The calculation formula of the second spatio-temporal feature is as follows:

[0057]

[0058] Wherein, represents the second spatio-temporal feature at the tth moment, and t u represents the first target detection moment of the staff.

[0059] Preferably, key frame extraction is performed on the foreground video frame subsequence to obtain key video frames of the detection target with interaction behavior, including:

[0060] Performing key point extraction on the foreground video frames in the foreground video frame subsequence based on a pre-trained key point extraction model to obtain a set of key points of the foreground video frames;

[0061] Obtaining an optical flow field between the foreground video frame subsequences, and extracting the optical flow vector of each key point in the optical flow field between adjacent foreground video frames in the foreground video frame subsequence;

[0062] Summing all the optical flow vectors in the key points to obtain an average optical flow vector of the key points, and calculating a unit vector of the average optical flow vector of the key points to obtain a unit vector of the optical flow vector of the key points;

[0063] Calculating an average modulus length of the optical flow vector of the key points, and multiplying the average modulus length of the optical flow vector of the key points by the unit vector to obtain a global motion vector, where the expression of the global motion vector is as follows:

[0064]

[0065] Among them, represents the global motion vector of the key point, represents the optical flow vector of the key point, represents the L2 norm of the optical flow vector of the key point, represents the set of key points;

[0066] Removing the global motion vector in the optical flow vector of the key points to obtain a standard optical flow vector of the key points, where the expression of the standard optical flow vector of the key points is as follows:

[0067]

[0068] Among them, represents the standard optical flow vector of the key point.

[0069] Preferably, key frame extraction is performed on the foreground video frame subsequence to obtain key video frames of the detection target with interaction behavior, further including:

[0070] Setting a sliding window, and sliding the sliding window in the foreground video frame subsequence to obtain a standard optical flow vector of the key points of the foreground video frame subsequence within the sliding window;

[0071] Calculate the unit vector of the standard optical flow vector of the key points of the foreground video frame subsequence within the sliding window to obtain the main direction optical flow vector. The expression of the main direction optical flow vector is as follows:

[0072]

[0073] Among them, represents the main direction optical flow vector;

[0074] Calculate the magnitude of the main direction optical flow vector. The calculation formula for the magnitude of the main direction optical flow vector is as follows:

[0075]

[0076] Among them, represents the main direction optical flow vector, and B max represents the set of all optical flow vectors to which the main direction within the sliding window belongs;

[0077] Based on the magnitude of the main direction optical flow vector and the main direction optical flow vector, obtain the main vector of the main direction. Project the main vector of the main direction from the Cartesian coordinate system to the polar coordinate system to obtain the spatial feature of the main direction of the key point. The expression of the main vector of the main direction is as follows:

[0078]

[0079] Among them, represents the main vector of the main direction;

[0080] Repeat the above operations until the sliding window moves to the end of the foreground video frame subsequence to obtain the spatial features of the main directions of all key points;

[0081] Sort the spatial features of the main directions of all key points in time series to obtain the spatial feature sequence corresponding to the foreground video frame subsequence;

[0082] Calculate the degree of change between adjacent spatial features in the spatial feature sequence, and use the foreground video frames with the degree of change greater than the preset change degree threshold as key video frames.

[0083] The beneficial effects of the present invention are as follows: (1) Through real-time video monitoring of the metering site and subsequent object detection and tracking, the present invention can obtain the dynamic information of the site in real time. Object detection and object tracking technologies can ensure timely capture of the position information and movement trajectories of metering terminal devices and staff, guarantee comprehensive understanding and effective management of on-site operation conditions. Moreover, through pre-trained object detection models and object tracking algorithms, different objects in the metering site (such as metering terminal devices and staff) can be accurately identified, and the dynamic changes of these objects can be tracked. Combining the Kalman filter and the Hungarian matching algorithm can improve the accuracy and stability of object tracking, thereby enabling efficient differentiation of interaction behaviors between objects and identifying whether there is interaction between objects (such as operations or interactions between staff and devices), further optimizing management decisions; (2) By filtering background frames from the foreground video frame sequence, the present invention can effectively remove background information, thus enabling clearer focus on the active objects at the site. In addition, the key frame extraction technology can accurately capture the critical moments of interaction behaviors and send these key video frames to management personnel. This not only ensures timely monitoring of important events but also facilitates management personnel to make decisions quickly or handle abnormal situations. And management personnel can receive key video frames through intelligent devices, enabling timely grasp of on-site dynamics. Especially when there are interaction behaviors or abnormal operations, management personnel can directly view the relevant video frames to quickly take corresponding measures. This not only improves the efficiency of on-site operation management but also avoids human errors and ensures the smooth progress of on-site operations; (3) Through real-time monitoring and behavior analysis of the metering site operation process, the present invention helps to discover potential safety hazards or non-standard operations (such as improper staff operations, equipment failures, etc.). By timely capturing relevant events and key behaviors, management personnel can take effective preventive measures or adjust the operation process, thereby enhancing on-site safety and compliance. Brief Description of the Drawings

[0084] Figure 1 It is a schematic flowchart of the steps of the overall method in an embodiment proposed by the present invention. Detailed Embodiments

[0085] Embodiment 1, as Figure 1 shown, a method for managing metering site operations based on machine vision proposed by the present invention includes:

[0086] S1. Conduct video monitoring of the metering site based on a preset camera to obtain a metering site video, acquire the metering site video frame sequence of the metering site video, and filter the background frames from the metering site video frame sequence to obtain a foreground video frame sequence;

[0087] S2. Perform object detection on each foreground video frame in the foreground video frame sequence based on a pre-trained object detection model to obtain object detection results, and perform object tracking on the detected objects using the SORT tracking algorithm based on the object detection results to obtain a first object tracking result;

[0088] S3. Perform tracking prediction on the detected objects through Kalman filtering based on the first object tracking result to obtain a second object tracking result, and perform correlation matching on the first object tracking result and the second object tracking result based on the Hungarian matching algorithm to obtain a third object tracking result;

[0089] S4. Determine whether there is an interaction behavior between the detected objects based on the third object tracking result. If there is an interaction behavior between the detected objects, obtain the subsequence of foreground video frames corresponding to the detected objects with the interaction behavior;

[0090] S5. Extract key frames from the subsequence of foreground video frames to obtain the key video frames of the detected objects with the interaction behavior, and send the key video frames of the detected objects with the interaction behavior to the intelligent device of the management personnel.

[0091] In the present invention, the preset camera refers to the camera installed at the metering site. By performing real-time video monitoring on the site, dynamic images or video information of the site are obtained. These cameras are usually fixed in position and monitor the site conditions at a set angle; SORT is a commonly used object tracking algorithm for tracking detected objects in a video. The SORT algorithm performs data association on the object detection results to update the position and trajectory of the object in real time. It is an efficient tracking method suitable for fast processing; Kalman filtering is a mathematical tool commonly used for estimation and prediction in dynamic systems. It can predict and correct the state of an object through sensor data to obtain more accurate results. In object tracking, Kalman filtering is used to predict the future position of an object, especially in the case of object loss or occlusion, and can effectively perform smoothing and prediction; The Hungarian algorithm is a classic optimization algorithm usually used to solve the optimal matching problem of a bipartite graph. In object tracking, the Hungarian algorithm is used to match different detection results of an object (such as the object positions at different time points) to ensure that the tracking results of the same object in different frames can be correctly associated; Key frame extraction refers to selecting representative and important frames from a video frame sequence. Key frames usually contain moments of significant changes in the video and can effectively represent the important content in the video. In the case of object interaction, extracting these frames can help better analyze the interaction behavior.

[0092] Embodiment 2, a method for managing metrology field operations based on machine vision proposed by the present invention, compared with embodiment 1, this embodiment further includes: filtering background frames of the metrology field video frame sequence to obtain a foreground video frame sequence, including:

[0093] A1. Traversing the metering scene video frame sequence, modeling the background of the initial metering scene video frame in the metering scene video frame sequence based on Gaussian distribution, so as to obtain a background model corresponding to the initial metering scene video frame;

[0094] A2. When traversing to the current metering scene video frame, the background model is compared with the pixel points of the current metering scene video frame one by one to classify the pixel points of the current metering scene video frame to determine the type of the pixel points, wherein the type of the pixel points includes background points and foreground points;

[0095] A3. Counting the number of foreground points in the current metering scene video frame to obtain the number of foreground points, and calculating the retention coefficient of the current metering scene video frame based on the number of foreground points to obtain the retention coefficient of the current metering scene video frame;

[0096] A4, comparing the retention coefficient of the current metering scene video frame with a preset retention threshold, if the retention coefficient of the current metering scene video frame is greater than the preset retention threshold, adding the current metering scene video frame to the foreground video frame sequence;

[0097] A5. Update the background model based on the current metering scene video frame to obtain a new background model, and repeat the above operation until all video frames in the metering scene video frame sequence are traversed to obtain the foreground video frame sequence corresponding to the metering scene video frame sequence.

[0098] In this embodiment, Gaussian distribution is a common probability distribution used to describe the distribution of continuous random variables. It usually takes the form of a bell-shaped curve and is widely used in signal processing and background modeling. Background modeling is a key step in video analysis. It creates a model that reflects the background content in the video by analyzing the image data of consecutive frames. The background model helps distinguish the foreground (such as an object) from the background (such as a static environment) in the video.

[0099] In an optional embodiment, the background model is as follows:

[0100]

[0101] Among them, x j,t represents the pixel value of the jth pixel at time t in the video frame of the measurement site, P represents the background distribution of the pixel, represents the weight value of the i-th Gaussian distribution at time t in the background model, represents the mean value of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame, represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame, where, and represents the average pixel value of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space, where, and represents the standard deviation of the pixel values of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space, η represents the probability density function of the Gaussian distribution, where,

[0102]

[0103] The calculation formula of the retention coefficient is as follows:

[0104]

[0105] where, τ represents the retention coefficient, P represents the background distribution of the pixel, and m*n represents the total number of pixels in the metering site video frame;

[0106] The update formula of the background model is as follows:

[0107]

[0108] where, M i,t represents whether the pixel is a foreground point, M i,t =1 indicates that the pixel is a foreground point, M i,t =0 indicates that the pixel is a background point, X t represents the pixel value of the pixel, and α and ρ represent preset weight coefficients.

[0109] In an optional embodiment, the target detection model includes an appearance feature extraction module and a classification module. The appearance feature extraction module sequentially includes 5 feature extraction blocks and 4 feature fusion blocks. Among them, the feature extraction block is sequentially composed of a convolutional layer, a batch normalization layer, a ReLU activation function, and a downsampling operation. The multi-level feature fusion block includes a convolutional layer and a feature fusion layer. The feature fusion layer is composed of two parallel branches. Among them, the first parallel branch is composed of a global average pooling layer, a convolutional layer, a ReLU activation function, and a convolutional layer. The second parallel branch is composed of a convolutional layer, a ReLU activation function, and a convolutional layer. The classification module sequentially includes a global average pooling layer, a fully connected layer, and a Softmax activation function.

[0110] In an optional embodiment, five feature extraction blocks sequentially perform shallow feature extraction on the foreground video frame to obtain the first shallow feature, the second shallow feature, the third shallow feature, the fourth shallow feature, and the fifth shallow feature. The multi-level feature fusion block is used to perform a convolution operation on the first shallow feature, the second shallow feature, the third shallow feature, the fourth shallow feature, and the fifth shallow feature based on a convolutional layer to obtain the first global shallow feature, the second global shallow feature, the third global shallow feature, the fourth global shallow feature, and the fifth global shallow feature, and initially fuse the first global shallow feature and the second global shallow feature through a feature fusion layer to obtain the first shallow fusion feature, initially fuse the first shallow fusion feature and the third global shallow feature through a feature fusion layer to obtain the second shallow fusion feature, initially fuse the second shallow fusion feature and the fourth global shallow feature through a feature fusion layer to obtain the third shallow fusion feature, and initially fuse the third shallow fusion feature and the fifth global shallow feature through a feature fusion layer to obtain the appearance feature.

[0111] In an optional embodiment, the expression of the object detection result is as follows:

[0112] R = {r1, r2,..., r t};

[0113] where R represents the object detection result, and r t represents the detected object at time t, where represents the jth detected object at time t, where represents the category of the jth detected object at time t, where the categories of the detected objects include metering terminal devices and staff members, represents the recognition probability of the jth detected object at time t, represents the upper left corner coordinates of the detection box of the jth detected object at time t, and represent the width and height of the detection box of the jth detected object at time t;

[0114] The expression of the first object tracking result is as follows:

[0115] OT = {ot1, ot2,..., ot t};

[0116] where OT represents the first object tracking result, and ot t represents the detected object tracked at time t, where represents the jth detected object tracked at time t, where Indicates the tracking number of the j-th detected target tracked by the target at time t;

[0117] The expression of the second target tracking result is as follows:

[0118] UT = {ut2,..., ut t+1};

[0119] Wherein, UT represents the second target tracking result, and ut t+1 Indicates the detected target tracked by the Kalman filter at time t + 1, where Indicates the j-th detected target tracked by the Kalman filter at time t + 1, where

[0120] It should be noted that the metering terminal device refers to a type of device, such as an intelligent meter or a monitoring device; the staff refers to the object recognized as a person in the image, usually the staff in the working environment.

[0121] In an optional embodiment, determining whether there is an interaction behavior between detected targets based on the third target tracking result includes:

[0122] B1. Calculate the area of the first detection frame of the detected target whose category is the metering terminal device at adjacent moments according to the third target tracking result;

[0123] B2. Calculate the overlapping area of the first detection frame of the detected target whose category is the metering terminal device at adjacent moments according to the third target tracking result;

[0124] B3. Calculate the first detection coincidence degree based on the area of the first detection frame and the overlapping area of the first detection frame at adjacent moments, and calculate the first spatio-temporal feature of the detected target whose category is the metering terminal device based on the first detection coincidence degree;

[0125] B4. If the first spatio-temporal feature of the detected target whose category is the metering terminal device is greater than a preset first threshold, it is determined that the metering terminal device is in a stationary state; otherwise, it is determined that the metering terminal device is in a non-stationary state;

[0126] B5. Calculate the area of the second detection frame of the detected target whose category is the staff according to the third target tracking result;

[0127] B6. Calculate the overlapping area of the second detection frame of the detected target whose category is the metering terminal device and the staff according to the third target tracking result;

[0128] B7. Calculate the second detection coincidence degree of the metering terminal device and the staff based on the second detection frame area and the second detection frame overlap area, and calculate the second spatio-temporal feature of the metering terminal device and the staff based on the second detection coincidence degree;

[0129] B8. If the second spatio-temporal feature of the metering terminal device and the staff is greater than the preset second threshold, then judge the non-contact position state between the metering terminal device and the staff, otherwise, judge the contact position state between the metering terminal device and the staff;

[0130] B8. If the metering terminal device is in a stationary state and the state between the metering terminal device and the staff is a contact position state, then judge that there is an interaction behavior between the detection targets, otherwise, judge that there is no interaction behavior between the detection targets.

[0131] In an alternative embodiment, the calculation formula for the first detection coincidence degree is as follows:

[0132]

[0133] Wherein, represents the first detection coincidence degree between time t - 1 and time t, S(k1) and S(k2) represent the first detection frame areas at time t - 1 and time t, and S(O) represents the first detection frame overlap area at time t - 1 and time t;

[0134] The calculation formula for the first spatio-temporal feature is as follows:

[0135]

[0136] Wherein, R k (t) represents the first spatio-temporal feature at time t, and t s represents the first target detection time of the metering terminal device;

[0137] The calculation formula for the second detection coincidence degree is as follows:

[0138]

[0139] Wherein, represents the second detection coincidence degree of the metering terminal device and the staff at time t, S(a) represents the second detection frame overlap area at time t, and S(i) represents the second detection frame area at time t;

[0140] The calculation formula for the second spatio-temporal feature is as follows:

[0141]

[0142] Wherein, represents the second spatio-temporal feature at time t, and tu Indicates the first target detection moment of the staff member.

[0143] In an optional embodiment, key frame extraction is performed on the foreground video frame subsequence to obtain key video frames of the detection target with interaction behavior, including:

[0144] C1. Perform key point extraction on the foreground video frames in the foreground video frame subsequence based on a pre-trained key point extraction model to obtain a set of key points of the foreground video frames;

[0145] C2. Obtain the optical flow field between the foreground video frame subsequences, and extract the optical flow vectors of each key point in the optical flow field between adjacent foreground video frames in the foreground video frame subsequence;

[0146] C3. Sum all the optical flow vectors in the key points to obtain the average optical flow vector of the key points, and calculate the unit vector of the average optical flow vector of the key points to obtain the unit vector of the optical flow vector of the key points;

[0147] C4. Calculate the average modulus length of the optical flow vectors of the key points, and multiply the average modulus length of the optical flow vectors of the key points by the unit vector to obtain the global motion vector, where the expression of the global motion vector is as follows:

[0148]

[0149] Where, represents the global motion vector of the key point, represents the optical flow vector of the key point, represents the L2 norm of the optical flow vector of the key point, represents the set of key points;

[0150] C5. Remove the global motion vector in the optical flow vectors of the key points to obtain the standard optical flow vector of the key points, where the expression of the standard optical flow vector of the key points is as follows:

[0151]

[0152] Where, represents the standard optical flow vector of the key point.

[0153] It should be noted that the pre-trained key point extraction model refers to a model that has been trained on a large amount of image data and is used to automatically identify key points in images or video frames. Common models include SIFT, SURF, ORB, etc.; optical flow refers to the movement of objects or pixels in a scene in an image, usually represented by the change in pixel positions between adjacent frames; the optical flow vector is a vector that describes the movement of each pixel or key point in the optical flow field; the global motion vector refers to the average motion representation of the optical flow vectors of a group of key points.

[0154] In an optional embodiment, for key frame extraction of the foreground video frame subsequence to obtain the key video frames of the detection target with interaction behavior, it further includes:

[0155] C6. Set a sliding window, and slide the sliding window in the foreground video frame subsequence to obtain the standard optical flow vectors of the key points of the foreground video frame subsequence within the sliding window;

[0156] C7. Calculate the unit vectors of the standard optical flow vectors of the key points of the foreground video frame subsequence within the sliding window to obtain the main direction optical flow vectors, where the expression of the main direction optical flow vectors is as follows:

[0157]

[0158] Among them, represents the main direction optical flow vector;

[0159] C8. Calculate the modulus length of the main direction optical flow vectors, where the calculation formula of the modulus length of the main direction optical flow vectors is as follows:

[0160]

[0161] Among them, represents the main direction optical flow vector, and B max represents the set of all optical flow vectors to which the main direction within the sliding window belongs;

[0162] C9. Based on the modulus length of the main direction optical flow vectors and the main direction optical flow vectors, obtain the main vectors of the main direction, project the main vectors of the main direction from the Cartesian coordinate system to the polar coordinate system to obtain the spatial features of the main direction of the key points, where the expression of the main vectors of the main direction is as follows:

[0163]

[0164] Among them, represents the main vectors of the main direction;

[0165] C10. Repeat the above operations until the sliding window moves to the end of the foreground video frame subsequence to obtain the spatial features of the main direction of all key points;

[0166] C11. Sort the spatial features of the main directions of all key points in time series to obtain the spatial feature sequence corresponding to the foreground video frame subsequence;

[0167] C12. Calculate the degree of change between adjacent spatial features in the spatial feature sequence, and use the foreground video frames with a degree of change greater than the preset change degree threshold as key video frames.

[0168] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.

Claims

1. A method for managing on-site metering operations based on machine vision, characterized in that, Including: Performing video monitoring on the metering site based on a preset camera to obtain a metering site video, acquiring a metering site video frame sequence of the metering site video, and filtering background frames from the metering site video frame sequence to obtain a foreground video frame sequence; Performing object detection on each foreground video frame in the foreground video frame sequence based on a pre-trained object detection model to obtain an object detection result, and performing object tracking on the detected object through the SORT tracking algorithm according to the object detection result to obtain a first object tracking result; Performing tracking prediction on the detected object through Kalman filtering according to the first object tracking result to obtain a second object tracking result, and associating and matching the first object tracking result and the second object tracking result based on the Hungarian matching algorithm to obtain a third object tracking result; Judging whether there is an interaction behavior between the detected objects based on the third object tracking result. If there is an interaction behavior between the detected objects, acquiring a subsequence of foreground video frames corresponding to the detected objects with the interaction behavior; Performing key frame extraction on the subsequence of foreground video frames to obtain key video frames of the detected objects with the interaction behavior, and sending the key video frames of the detected objects with the interaction behavior to the intelligent device of the management personnel.

2. The metrology field operation management method based on machine vision according to claim 1, characterized in that, Filtering background frames from the metering site video frame sequence to obtain a foreground video frame sequence, including: Traversing the metering site video frame sequence, and modeling the background of the initial metering site video frame in the metering site video frame sequence based on a Gaussian distribution to obtain a background model corresponding to the initial metering site video frame; When traversing to the current metering site video frame, comparing each pixel point of the current metering site video frame with the background model one by one to classify the pixel points of the current metering site video frame to judge the type of the pixel points, where the types of the pixel points include background points and foreground points; Counting the number of foreground points judged in the current metering site video frame to obtain the number of foreground points, and calculating a retention coefficient of the current metering site video frame based on the number of foreground points to obtain the retention coefficient of the current metering site video frame; Comparing the retention coefficient of the current metering site video frame with a preset retention threshold. If the retention coefficient of the current metering site video frame is greater than the preset retention threshold, adding the current metering site video frame to the foreground video frame sequence; Updating the background model based on the current metering site video frame to obtain a new background model, and repeating the above operations until all video frames in the metering site video frame sequence are traversed to obtain the foreground video frame sequence corresponding to the metering site video frame sequence.

3. The metrology on-site operation management method based on machine vision according to claim 2, characterized in that, The background model is as follows: Among them, x j,t represents the pixel value of the j-th pixel at time t in the metering site video frame, P represents the background distribution of the pixel, represents the weight value of the i-th Gaussian distribution at time t in the background model, represents the mean value of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame, represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel at time t in the metering site video frame. Among them, and represent the average pixel value of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space. Among them, and represent the standard deviation of the pixel values of the R, G, and B components of the j-th pixel at time t in the metering site video frame in the RGB color space. η represents the probability density function of the Gaussian distribution. Among them, The calculation formula of the retention coefficient is as follows: Where τ represents the retention coefficient, P represents the background distribution of the pixel points, and m*n represents the total number of pixel points in the metering site video frame; The update formula of the background model is as follows: Among them, M i,t indicates whether the pixel is a foreground point. M i,t = 1 indicates that the pixel is a foreground point. M i,t = 0 indicates that the pixel is a background point. X t represents the pixel value of the pixel, and α and ρ represent preset weight coefficients.

4. A method for on-site operation management of metrology based on machine vision according to claim 1, characterized in that The target detection model includes an appearance feature extraction module and a classification module. The appearance feature extraction module sequentially includes five feature extraction blocks and four multi-level feature fusion blocks. Among them, each feature extraction block is sequentially composed of a convolutional layer, a batch normalization layer, a ReLU activation function, and a downsampling operation. The multi-level feature fusion block includes a convolutional layer and a feature fusion layer. The feature fusion layer consists of two parallel branches. Among them, the first parallel branch is composed of a global average pooling layer, a convolutional layer, a ReLU activation function, and a convolutional layer. The second parallel branch is composed of a convolutional layer, a ReLU activation function, and a convolutional layer. The classification module sequentially includes a global average pooling layer, a fully connected layer, and a Softmax activation function.

5. A method for on-site operation management of metrology based on machine vision according to claim 4, characterized in that, The five feature extraction blocks sequentially perform shallow feature extraction on the foreground video frame to obtain the first shallow feature, the second shallow feature, the third shallow feature, the fourth shallow feature, and the fifth shallow feature. The multi-level feature fusion block is used to perform a convolution operation on the first shallow feature, the second shallow feature, the third shallow feature, the fourth shallow feature, and the fifth shallow feature based on the convolutional layer to obtain the first global shallow feature, the second global shallow feature, the third global shallow feature, the fourth global shallow feature, and the fifth global shallow feature, and preliminarily fuse the first global shallow feature and the second global shallow feature through the feature fusion layer to obtain the first shallow fusion feature. The first shallow fusion feature and the third global shallow feature are preliminarily fused through the feature fusion layer to obtain the second shallow fusion feature. The second shallow fusion feature and the fourth global shallow feature are preliminarily fused through the feature fusion layer to obtain the third shallow fusion feature. The third shallow fusion feature and the fifth global shallow feature are preliminarily fused through the feature fusion layer to obtain the appearance feature.

6. The method for managing on-site metering operations based on machine vision according to claim 5, wherein, The expression of the target detection result is as follows: R={r1,r2,...,r t}; Among them, R represents the target detection result, and r t represents the detected target at time t. Among them, represents the j-th detected target at time t. Among them, represents the category of the j-th detected target at time t. The categories of the detected targets include metering terminal devices and staff members, represents the recognition probability of the j-th detected target at time t, represents the upper left corner coordinates of the detection box of the j-th detected target at time t, and represents the width and height of the detection box of the j-th detected target at time t; The expression of the first target tracking result is as follows: OT = {ot1, ot2,..., ot t}; Among them, OT represents the first target tracking result, and ot t represents the detected target tracked by the target at time t. Among them, represents the j-th detected target tracked by the target at time t. Among them, represents the tracking number of the j-th detected target tracked by the target at time t; The expression of the second target tracking result is as follows: UT = {ut2,..., ut t+1}; Among them, UT represents the second target tracking result, and ut t+1 represents the detected target tracked by the Kalman filter at time t + 1. Among them, represents the j-th detected target tracked by the Kalman filter at time t + 1. Among them, 7. A method for managing on-site metering operations based on machine vision according to claim 6, characterized in that, Judging whether there is an interaction behavior between the detected targets based on the third target tracking result includes: Calculating the area of the first detection box of the detected target of the category of metering terminal equipment at adjacent moments according to the third target tracking result; Calculating the overlapping area of the first detection box of the detected target of the category of metering terminal equipment at adjacent moments according to the third target tracking result; Calculating the first detection coincidence degree based on the area of the first detection box and the overlapping area of the first detection box at the adjacent moments, and calculating the first spatio-temporal feature of the detected target of the category of metering terminal equipment based on the first detection coincidence degree; If the first spatio-temporal feature of the detected target of the category of metering terminal equipment is greater than a preset first threshold, it is determined that the metering terminal equipment is in a stationary state, otherwise, it is determined that the metering terminal equipment is in a non-stationary state; Calculating the area of the second detection box of the detected target of the category of staff according to the third target tracking result; Calculating the overlapping area of the second detection box of the detected target of the category of metering terminal equipment and staff according to the third target tracking result; Calculate the second detection coincidence degree of the metering terminal device and the staff based on the area of the second detection box and the overlapping area of the second detection box, and calculate the second spatio-temporal feature of the metering terminal device and the staff based on the second detection coincidence degree; If the second spatio-temporal feature of the metering terminal device and the staff is greater than a preset second threshold, then judge the non-contact position state between the metering terminal device and the staff, otherwise, judge the contact position state between the metering terminal device and the staff; If the metering terminal device is in a stationary state and the metering terminal device and the staff are in a contact position state, then judge that there is an interaction behavior between the detection targets, otherwise, judge that there is no interaction behavior between the detection targets.

8. A method for on-site operation management of metrology based on machine vision according to claim 7, characterized in that The calculation formula of the first detection coincidence degree is as follows: Among them, represents the first detection overlap degree between the (t - 1)th moment and the tth moment, S(k1) and S(k2) represent the areas of the first detection boxes at the (t - 1)th moment and the tth moment, and S(O) represents the overlapping area of the first detection boxes at the (t - 1)th moment and the tth moment; The calculation formula of the first spatio-temporal feature is as follows: Among them, R k (t) represents the first spatio-temporal feature at time t, where t s represents the first target detection time of the metering terminal device; The calculation formula of the second detection coincidence degree is as follows: Among them, represents the second detection coincidence degree of the metering terminal device and the staff at time t, S(a) represents the overlapping area of the second detection frame at time t, and S(i) represents the area of the second detection frame at time t; The calculation formula of the second spatio-temporal feature is as follows: Among them, represents the second spatio-temporal feature at time t, where t u represents the first target detection time of the staff member.

9. A method for on-site operation management of measurement based on machine vision according to claim 8, characterized in that, Perform key frame extraction on the foreground video frame subsequence to obtain the key video frames of the detection targets with interaction behavior, including: Perform key point extraction on the foreground video frames in the foreground video frame subsequence based on a pre-trained key point extraction model to obtain the key point set of the foreground video frames; Obtain the optical flow field between the foreground video frame subsequences, and extract the optical flow vectors of each key point in the optical flow field between adjacent foreground video frames in the foreground video frame subsequences; Sum all the optical flow vectors in the key points to obtain the average optical flow vector of the key points, and calculate the unit vector of the average optical flow vector of the key points to obtain the unit vector of the optical flow vector of the key points; Calculate the average modulus length of the optical flow vectors of the key points, and multiply the average modulus length of the optical flow vectors of the key points by the unit vector to obtain the global motion vector, where the expression of the global motion vector is as follows: Among them, represents the global motion vector of the key point, represents the optical flow vector of the key point, represents the L2 norm of the optical flow vector of the key point, represents the set of key points; Remove the global motion vector in the optical flow vectors of the key points to obtain the standard optical flow vector of the key points, where the expression of the standard optical flow vector of the key points is as follows: Among them, represents the standard optical flow vector of the key point.

10. A method for on-site operation management of measurement based on machine vision according to claim 9, characterized in that, Performing key frame extraction on the foreground video frame subsequence to obtain the key video frames of the detection targets with interaction behavior further includes: Set a sliding window, and slide the sliding window in the foreground video frame subsequence to obtain the standard optical flow vectors of the key points of the foreground video frame subsequence within the sliding window; Calculate the unit vector of the standard optical flow vectors of the key points of the foreground video frame subsequence within the sliding window to obtain the main direction optical flow vector, where the expression of the main direction optical flow vector is as follows: Among them, represents the main direction optical flow vector; Calculate the modulus length of the main direction optical flow vector, where the calculation formula of the modulus length of the main direction optical flow vector is as follows: Among them, represents the main direction optical flow vector, and B max represents the set of all optical flow vectors to which the main direction within the sliding window belongs; Obtain the main vector of the main direction based on the modulus length of the main direction optical flow vector and the main direction optical flow vector, and project the main vector of the main direction from the Cartesian coordinate system to the polar coordinate system to obtain the spatial feature of the main direction of the key points, where the expression of the main vector of the main direction is as follows: Among them, represents the main vector of the main direction; Repeat the above operation until the sliding window moves to the end of the foreground video frame subsequence to obtain the spatial features of the main directions of all key points; Sort the spatial features of the main directions of all key points in time series to obtain the spatial feature sequence corresponding to the foreground video frame subsequence; Calculate the degree of change between adjacent spatial features in the spatial feature sequence, and use the foreground video frames with the degree of change greater than the preset change degree threshold as key video frames.