Visual target tracking method based on machine learning

By adopting a visual target tracking method based on machine learning in the video surveillance system, the problems of waste of computing resources, low target detection accuracy and multi-camera coordination in the prior art are solved, efficient target detection and large-scale real-time monitoring are achieved, and the safety of public places is improved.

CN120071252AActive Publication Date: 2025-05-30SHENZHEN LIGUAN DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510151146.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art has problems in video surveillance that waste computing resources, low target detection accuracy, inability to effectively express the spatial and temporal relationship between targets in the monitoring area, and inability to effectively coordinate multi-camera information.

Method used

Using a visual target tracking method based on machine learning, video is collected through multiple surveillance cameras, video frame segmentation and target screening is performed, object detection is used to use a pre-trained object detection model, and monitoring diagram structure is constructed, and an abnormality detection model is finally used to use a pre-trained anomaly detection model.

Benefits of technology

Effectively eliminate unrelated frames, improve computing efficiency and target detection accuracy, clearly represent the relationship between targets, realize large-scale real-time monitoring and automated abnormal detection, and improve the safety of public places or important areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071252A_ABST
    Figure CN120071252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of anomaly monitoring, in particular to a visual target tracking method based on machine learning, which comprises the following steps: performing video acquisition on a plurality of monitoring areas based on a plurality of monitoring cameras to obtain a plurality of monitoring videos, and framing the monitoring videos to obtain a monitoring video frame sequence corresponding to the monitoring videos; performing target screening on the monitoring video frame sequence to obtain a monitoring video frame set; and performing target detection on the monitoring video frame set based on a pre-trained target detection model to obtain a target detection result of each monitoring video frame. According to the invention, by screening the monitoring video frames, irrelevant or irrelevant frames can be effectively eliminated, the calculation amount is reduced, the subsequent processing process is optimized, and by extracting the effective frame sequence related to target tracking, the system can process important information more intensively, so that the tracking efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly monitoring, and particularly to a visual target tracking method based on machine learning. Background Art

[0002] Traditional methods usually process all video frames, lacking the screening of irrelevant or unimportant frames. This leads to the need to process a large amount of unimportant information, wasting computing resources and reducing the overall efficiency of the system. In large-scale monitoring scenarios, the processing time and computing burden may increase significantly, affecting real-time performance and accuracy. Target detection in traditional methods often relies on simple image processing algorithms or manually designed features, and is easily affected by complex environmental conditions (such as light changes, occlusion, rapid movement of targets, etc.), resulting in inaccurate target detection results, easy missed detection or false detection. Lacking the support of deep learning and advanced target detection models, the detection accuracy of traditional methods is low and cannot meet the requirements of modern monitoring systems for accuracy and robustness. Traditional methods usually do not have an effective way to express the spatial and temporal relationships between targets in the monitoring area. Even if target detection can be performed, it is difficult to comprehensively understand the interactions and relationships between multiple targets, resulting in scattered analysis of monitoring data and difficult to provide in-depth target behavior analysis and prediction. Traditional methods usually only rely on a single camera or limited monitoring devices, with a small monitoring range and unable to effectively coordinate multi-camera information. In this way, it is difficult to perform real-time monitoring and comprehensive analysis of a large area, unable to fully utilize the synergy of multi-cameras, thus affecting the applicability and response speed of the system in large-scale monitoring. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a visual target tracking method based on machine learning.

[0004] The technical solution adopted to solve the above technical problem is: A visual target tracking method based on machine learning, comprising:

[0005] Performing video acquisition on multiple monitoring areas based on multiple monitoring cameras to obtain multiple monitoring videos, and performing frame division on the monitoring videos to obtain a monitoring video frame sequence corresponding to the monitoring videos;

[0006] Performing target screening on the monitoring video frame sequence to obtain a monitoring video frame set;

[0007] Performing target detection on the monitoring video frame set based on a pre-trained target detection model to obtain a target detection result for each monitoring video frame;

[0008] Constructing a monitoring graph structure corresponding to the multiple monitoring areas based on the target detection results of each monitoring video frame;

[0009] Perform anomaly detection on the monitored graph structure based on a pre-trained anomaly detection model to obtain the anomaly detection result corresponding to the monitored area.

[0010] Preferably, the target detection result set includes the detection box coordinates of the detection target, the confidence of the target detection result, and the target category, where the target category includes objects, people, and the environment. The monitored graph structure includes a node set and an edge set. The nodes in the node set correspond to the monitored areas, and the edges in the edge set correspond to the spatial adjacency relationship, where the spatial adjacency relationship is used to represent the transfer of the detection target from one monitored area to another.

[0011] Preferably, perform target screening on the monitored video frame sequence to obtain a monitored video frame set, including:

[0012] Traverse the monitored video frame sequence, and model the background of the initial monitored video frame in the monitored video frame sequence based on the Gaussian distribution to obtain the background model corresponding to the initial monitored video frame;

[0013] When traversing to the current monitored video frame, compare the pixel points of the current monitored video frame with the background model one by one to classify the pixel points of the current monitored video frame to determine the type of the pixel points, where the types of the pixel points include background points and foreground points;

[0014] Count the number of foreground points determined in the current monitored video frame to obtain the number of foreground points, and calculate the retention coefficient of the current monitored video frame based on the number of foreground points to obtain the retention coefficient of the current monitored video frame;

[0015] Compare the retention coefficient of the current monitored video frame with a preset retention threshold. If the retention coefficient of the current monitored video frame is greater than the preset retention threshold, add the current monitored video frame to the monitored video frame set;

[0016] Update the background model based on the current monitored video frame to obtain a new background model, and repeat the above operations until all the monitored video frames in the monitored video frame sequence are traversed to obtain the monitored video frame set corresponding to the monitored video frame sequence.

[0017] Preferably, the background model is as follows:

[0018]

[0019] where x j,trepresents the pixel value of the j-th pixel at time t in the monitored video frame, and P represents the background distribution of the pixel represents the weight value of the i-th Gaussian distribution at time t in the background model represents the mean value of the i-th Gaussian distribution of the j-th pixel at time t in the monitored video frame represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel at time t in the monitored video frame, where and represent the average pixel value of the R, G, and B components of the j-th pixel at time t in the monitored video frame in the RGB color space, where and represent the standard deviation of the pixel values of the R, G, and B components of the j-th pixel at time t in the monitored video frame in the RGB color space, and η represents the probability density function of the Gaussian distribution, where

[0020] The calculation formula of the retention coefficient is as follows:

[0021]

[0022] where τ represents the retention coefficient, P represents the background distribution of the pixel, and m*n represents the total number of pixels in the monitored video frame

[0023] The update formula of the background model is as follows:

[0024]

[0025] where M i,t represents whether the pixel is a foreground point, M i,t =1 indicates that the pixel is a foreground point, M i,t =0 indicates that the pixel is a background point, X t represents the pixel value of the pixel, and α and ρ represent preset weight coefficients

[0026] Preferably, the target detection model extracts features of different scales of the monitored video frame through 3 efficient convolutional modules, which are respectively represented as C 1 、C 2 and C 3 , and uses multi-scale feature fusion to perform target recognition on features of different scales. Among them, for the smallest-scale feature C 3 , first perform a feature transformation on the smallest-scale feature C 3 through a 1×1 convolutional operation, and then perform an upsampling operation to obtain a feature of the same size as C 2 , and then combine this feature with C 2are fused and input into an efficient convolution module to extract the fused features. Then, through an upsampling operation, the fused features are mapped to the same size as C 1 and concatenated with C 1 at the channel level. Finally, an efficient convolution module is used to extract the final features, and these features are input into a classification module to obtain the set of target detection results for the monitored video frame.

[0027] Preferably, the efficient convolution module uses two different branches to process the input features. Among them, the first branch first extracts the inter-channel dependence features of the input through a 3×3 depthwise convolution kernel, and then extracts the spatial dependence features of the input through a 1×1 pointwise convolution operation. Among them, the second branch first extracts the spatial dependence features of the input through a 1×1 pointwise convolution operation, then uses a 3×3 depthwise convolution operation to extract the inter-channel dependence features of the input features, and then weights each channel through an attention module to extract the weighted features. Finally, a 1×1 pointwise convolution operation is used to extract the spatial dependence features of the weighted features. The features of the two branches are integrated in a channel-level concatenation manner, and the features of the two branches are integrated through a 3×3 depthwise convolution operation and a 1×1 pointwise convolution operation.

[0028] Preferably, the anomaly detection model uses a graph neural network. The anomaly detection model includes a graph convolutional network, a convolutional network, a hybrid module, and an anomaly judgment module. The graph convolutional network is used to perform a convolutional operation on the monitored graph structure to obtain the temporal correlation features of the monitored graph structure. The convolutional network is used to perform a convolutional operation on the state feature matrix of the monitored graph structure in the preset time period to obtain the periodic features of the monitored graph structure. The hybrid module is used to fuse the temporal correlation features and periodic features of the monitored graph structure to obtain an anomaly probability result. The anomaly judgment module is used to judge whether the monitored graph structure is abnormal according to the anomaly probability result.

[0029] Preferably, the graph convolutional network sequentially includes an input layer, a first graph convolutional layer, a second graph convolutional layer, a first fully connected layer, and an output layer, and the convolutional network sequentially includes a first convolutional layer, a spatial attention layer, a second convolutional layer, a pooling layer, and a second fully connected layer. Among them, the spatial attention layer performs average pooling and max pooling on the input features to obtain an average pooling vector and a max pooling vector, processes the average pooling vector and the max pooling vector through a convolutional layer and a Sigmoid activation function to obtain a spatial attention output weight, and multiplies the spatial attention output weight element-wise with the input features to obtain a spatial attention vector. The hybrid module concatenates the temporal correlation feature and the periodic feature to obtain a concatenated vector feature, extracts the temporal feature of the concatenated vector feature through a fully connected layer with one neuron, and outputs the anomaly probability of the temporal feature through a Sigmoid activation function.

[0030] The beneficial effects of the present invention are as follows: (1) By screening the surveillance video frames, the present invention can effectively eliminate irrelevant or unimportant frames, reduce the computational amount and optimize the subsequent processing process. By extracting the effective frame sequence related to target tracking, the system can more centrally process important information, thereby improving the tracking efficiency and accuracy. Moreover, by using a pre-trained object detection model, the surveillance video frames can be accurately detected, ensuring the accurate identification and positioning of the targets in each surveillance video frame. The object detection results help to construct the subsequent surveillance graph structure and form the dynamic representation of the targets, which is helpful for further target analysis and behavior recognition; (2) Based on the object detection results of each surveillance video frame, the present invention constructs a surveillance graph structure that can better represent the relationships between the targets in the surveillance area. This structured representation form can clearly show the spatial and temporal associations between the targets. Moreover, by performing anomaly detection on the surveillance graph structure based on a pre-trained anomaly detection model, the system can automatically identify events or behaviors that do not conform to normal behaviors. For example, in public security surveillance, abnormal behaviors can be automatically detected and a real-time alarm function can be provided; (3) By combining multi-camera surveillance, object detection, and anomaly detection, the present invention can achieve real-time surveillance and automated anomaly detection in a large-scale surveillance area, timely discover potential security threats and respond quickly, which is crucial for improving the security of public places or important areas. Especially in the case of unattended, it can greatly improve the efficiency and accuracy of security surveillance. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic flowchart of the steps of the overall method in an embodiment proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] Embodiment 1, as Figure 1As shown in the figure, a vision target tracking method based on machine learning proposed by the present invention includes:

[0033] S1. Video acquisition is performed on multiple monitoring areas based on multiple monitoring cameras to obtain multiple monitoring videos, and the monitoring videos are framed to obtain the monitoring video frame sequences corresponding to the monitoring videos;

[0034] S2. Target screening is performed on the monitoring video frame sequences to obtain a set of monitoring video frames;

[0035] S3. Target detection is performed on the set of monitoring video frames based on a pre-trained target detection model to obtain the target detection results of each monitoring video frame;

[0036] S4. A monitoring graph structure corresponding to multiple monitoring areas is constructed based on the target detection results of each monitoring video frame;

[0037] S5. Anomaly detection is performed on the monitoring graph structure based on a pre-trained anomaly detection model to obtain the anomaly detection results corresponding to the monitoring areas.

[0038] In the present invention, a monitoring camera is a device for real-time video acquisition, which is commonly used in fields such as public security monitoring and traffic management. They can capture video signals in a specified area; a monitoring area refers to a specific spatial range covered by a monitoring camera. Each camera usually monitors one or more areas, such as shopping malls, streets, parking lots, etc.; target screening is a process of identifying and classifying different objects in video frames. The purpose is to screen out the targets of interest from the video frames to provide effective data for subsequent analysis; a monitoring graph structure is a graphical data structure constructed based on the targets, objects, or events detected in a monitoring area. It represents the relationship between the monitoring area and the targets in the form of nodes and edges. Each node represents a target or a monitoring area, and each edge represents the relationship between different nodes; after performing anomaly detection, the system will output the anomaly detection results of each monitoring area, usually in the form of a mark or a report, indicating whether an abnormal event has occurred in a certain monitoring area and the possible types or severities of these events.

[0039] Embodiment 2. A vision target tracking method based on machine learning proposed by the present invention. Compared with Embodiment 1, this embodiment further includes: The target detection result set includes the detection box coordinates of the detected target, the confidence of the target detection result, and the target category. Among them, the target category includes objects, people, and the environment. The monitoring graph structure includes a node set and an edge set. Among them, the nodes in the node set correspond to the monitoring areas, and the edges in the edge set correspond to the spatial adjacency relationship. Among them, the spatial adjacency relationship is used to represent the transfer of the detected target from one monitoring area to another monitoring area.

[0040] In this embodiment, the coordinates of the target detection box refer to the position of the target object recognized by the target detection algorithm in the image, and usually a rectangular box (i.e., the detection box) is used to represent the target, and the coordinates of the box represent the positions of the four vertices of the rectangular box; the confidence level of the target detection result is the probability value output by the target detection model, indicating the confidence level of detecting a certain target. For example, if the confidence level is 0.95, it means that the model believes that the probability that the object in the box belongs to a certain category is 95%. A high confidence level usually indicates that the target detection result is relatively reliable; the target category: refers to the category or type to which the detected target belongs. For example, the target can be an "object" (such as a car, a suitcase, etc.), a "person" (such as a pedestrian, a specific person, etc.) or an "environment" (such as a wall, a ceiling, etc.), and these categories help the system to distinguish and process different types of targets; the spatial adjacency relationship is used to represent the spatial relationship between different monitoring areas. Specifically, it describes the situation where the detected target transfers from one monitoring area to another. For example, if two areas covered by a monitoring camera (such as two rooms, two streets) are connected to each other, then there is a spatial adjacency relationship between them. When the target moves from one area to another, the spatial adjacency relationship helps to understand the path and behavior of the target; this adjacency relationship not only focuses on physical connection, but may also include the behavior of the target in dynamic monitoring. For example, a person is detected in monitoring area A and then appears in the adjacent area B. Then, the dynamic process of this target transferring from area A to area B can be described by the spatial adjacency relationship.

[0041] In an alternative embodiment, target screening is performed on the monitoring video frame sequence to obtain a monitoring video frame set, including:

[0042] A1. Traverse the monitoring video frame sequence, and model the background of the initial monitoring video frame in the monitoring video frame sequence based on the Gaussian distribution to obtain the background model corresponding to the initial monitoring video frame;

[0043] A2. When the current monitoring video frame is traversed, use the background model to compare each pixel point of the current monitoring video frame one by one to classify the pixel points of the current monitoring video frame to determine the type of the pixel points, where the types of pixel points include background points and foreground points;

[0044] A3. Count the number of foreground points determined in the current monitoring video frame to obtain the number of foreground points, and calculate the retention coefficient of the current monitoring video frame based on the number of foreground points to obtain the retention coefficient of the current monitoring video frame;

[0045] A4. Compare the retention coefficient of the current monitoring video frame with a preset retention threshold. If the retention coefficient of the current monitoring video frame is greater than the preset retention threshold, add the current monitoring video frame to the monitoring video frame set;

[0046] A5. Update the background model based on the current monitored video frame to obtain a new background model, and repeat the above operations until all the monitored video frames in the monitored video frame sequence are traversed to obtain the monitored video frame set corresponding to the monitored video frame sequence.

[0047] It should be noted that the Gaussian distribution: is a probability distribution in statistics, also called the normal distribution. Background modeling uses the Gaussian distribution to describe the background pixels in the video frame. These background pixels usually follow certain statistical laws, and the range of their pixel values can be estimated through the Gaussian distribution; Background modeling is used to extract background information from the video frame sequence and then distinguish the dynamic foreground (i.e., changing or moving objects). This method is usually used to separate dynamic targets from a static environment, such as moving pedestrians or vehicles; Background points refer to the pixel points in the image that belong to the background area, and these pixels remain stable and do not change in multiple frames of the monitored video; Foreground points refer to the pixel points in the image that belong to the foreground area, and these pixel points are usually dynamically changing, representing the moving objects or events in the image; The retention coefficient is a coefficient that determines the influence degree of the current video frame on the background model according to the number of foreground points. A frame with a higher retention coefficient indicates that the current frame contains more foreground information (for example, someone or an object passes by), and vice versa, it may only contain background information. The retention coefficient is used to adjust the weight of the video frame for updating the background model; Background model update refers to adjusting the background model according to the content of the current video frame (especially the change of foreground points) to make it more suitable for the current environment. Usually, the background model is gradually updated to adapt to changes in the environment (such as changes in lighting, seasons, or long-term scene changes).

[0048] In an optional embodiment, the background model is as follows:

[0049]

[0050] where x j,t represents the pixel value of the j-th pixel point at time t in the monitored video frame, P represents the background distribution of the pixel point, represents the weight value of the i-th Gaussian distribution at time t in the background model, represents the average value of the i-th Gaussian distribution of the j-th pixel point at time t in the monitored video frame, represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel point at time t in the monitored video frame, where, and represent the average pixel values of the R, G, and B components of the j-th pixel point at time t in the monitored video frame in the RGB color space, where, and denotes the standard deviation of the pixel values of the R, G, and B components of the j-th pixel at time t in the RGB color space of the monitored video frame. η represents the probability density function of the Gaussian distribution, where,

[0051] The calculation formula for the retention coefficient is as follows:

[0052]

[0053] where τ represents the retention coefficient, P represents the background distribution of the pixel, and m*n represents the total number of pixels in the monitored video frame;

[0054] The update formula for the background model is as follows:

[0055]

[0056] where M i,t represents whether the pixel is a foreground point. M i,t = 1 indicates that the pixel is a foreground point, and M i,t = 0 indicates that the pixel is a background point. X t represents the pixel value of the pixel, and α and ρ represent preset weight coefficients.

[0057] In an alternative embodiment, the object detection model extracts features of different scales of the monitored video frame through 3 efficient convolutional modules, respectively denoted as C 1 , C 2 and C 3 . Multi-scale feature fusion is used to perform object recognition on features of different scales. Among them, for the smallest-scale feature C 3 , first, a 1×1 convolutional operation is performed on the smallest-scale feature C 3 to perform feature transformation, and then an upsampling operation is performed to obtain a feature of the same size as C 2 . Then, this feature is fused with C 2 , and the fused feature is input into an efficient convolutional module to extract the fused feature. Next, the fused feature is mapped to the same size as C 1 through an upsampling operation, and it is concatenated with C 1 at the channel level. Finally, an efficient convolutional module is used to extract the final feature, and this feature is input into a classification module to obtain the object detection result set of the monitored video frame.

[0058] It should be noted that the efficient convolution module refers to the design of the convolution layer that improves the computational efficiency by optimizing the convolution operation. The efficient convolution module usually improves the performance of the traditional convolution layer by reducing the computational amount or increasing the computational speed. For example, methods such as depthwise separable convolution, group convolution, and attention mechanism can significantly reduce the computational amount and improve the efficiency of the model while ensuring performance. Multi-scale feature fusion refers to combining information from different scales (i.e., features of different resolutions or different levels) to form a unified feature representation. Upsampling is an operation that increases the resolution of an image or feature map, usually used to restore the details of a smaller-sized image. Channel Concatenation refers to the operation of splicing two or more feature maps along the channel (i.e., depth) dimension.

[0059] In an optional embodiment, the efficient convolution module uses two different branches to process the input features. Among them, the first branch first extracts the input channel interdependence features through a 3×3 depthwise convolution kernel, and then extracts the input spatial dependence features through a 1×1 pixelwise convolution operation. Among them, the second branch first extracts the input spatial dependence features through a 1×1 pixelwise convolution operation, then uses a 3×3 depth convolution operation to extract the channel interdependence features of the input features, and then weights each channel through an attention module to extract the weighted features. Finally, a 1×1 pixelwise convolution operation is used to extract the spatial dependence features of the weighted features. The features of the two branches are integrated by means of channel concatenation, and the features of the two branches are integrated through a 3×3 depth convolution operation and a 1×1 pixelwise convolution operation.

[0060] In an optional embodiment, the anomaly detection model uses a graph neural network. The anomaly detection model includes a graph convolutional network, a convolutional network, a hybrid module, and an anomaly judgment module. The graph convolutional network is used to perform a convolution operation on the monitored graph structure to obtain the temporal correlation features of the monitored graph structure. The convolutional network is used to perform a convolution operation on the state feature matrix of the monitored graph structure in a preset time period to obtain the periodic features of the monitored graph structure. The hybrid module is used to perform feature fusion on the temporal correlation features and periodic features of the monitored graph structure to obtain an anomaly probability result. The anomaly judgment module is used to judge whether the monitored graph structure is abnormal according to the anomaly probability result.

[0061] It should be noted that graph neural networks (GNNs) are a class of neural network models for processing graph-structured data. In traditional neural networks, data is usually a fixed two-dimensional matrix (such as pixel values in an image), while graph neural networks can process graph data composed of nodes and edges. In a monitoring system, the monitored graph structure may represent the relationships between different monitoring points. Through GNNs, the network can learn the relationships between nodes and their dynamic changes in order to extract valuable features from the graph; graph convolutional networks are a type of convolutional neural network specifically for graph data. Different from traditional convolutional neural networks (CNNs), graph convolutional networks capture the structural information of the graph by performing convolutional operations on the nodes and their neighboring nodes in the graph. GCNs propagate information through the adjacency relationships of nodes, thereby learning the features of the nodes in the graph and being able to effectively extract the spatial structure features of the graph data. Graph convolutional networks are used to perform convolutional operations on the monitored graph structure to capture the temporal correlation features of the monitored graph structure; the anomaly probability result refers to a probability value calculated by the model, indicating whether a certain monitored graph structure is abnormal.

[0062] In an optional embodiment, the graph convolutional network sequentially includes an input layer, a first graph convolutional layer, a second graph convolutional layer, a first fully connected layer, and an output layer. The convolutional network sequentially includes a first convolutional layer, a spatial attention layer, a second convolutional layer, a pooling layer, and a second fully connected layer. Among them, the spatial attention layer performs average pooling processing and max pooling processing on the input features to obtain an average pooling vector and a max pooling vector. The average pooling vector and the max pooling vector are processed through a convolutional layer and a Sigmoid activation function to obtain the spatial attention output weights. The spatial attention output weights are multiplied element-wise with the input features to obtain the spatial attention vector. The mixing module concatenates the temporal correlation features and the periodic features to obtain a concatenated vector feature. The temporal features of the concatenated vector feature are extracted based on a fully connected layer of a neuron, and the anomaly probability of the temporal features is output through a Sigmoid activation function.

[0063] It should be noted that the fully connected layer is one of the most common layers in neural networks. Each neuron is connected to all neurons in the previous layer, so it is called "fully connected". Its role is to perform a linear transformation on the features from the previous layer and perform a non-linear mapping through an activation function, and finally output the activated features; the output of the spatial attention layer is a weight vector representing different spatial positions (image regions or nodes in the graph). These weight values are used to emphasize the important parts in the input feature map. It determines the features that the model should focus on at each position and strengthens the feature expression of the important regions by multiplying with the original features.

[0064] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art to which the present invention pertains.

Claims

1. A visual target tracking method based on machine learning, characterized in that: include: Based on multiple surveillance cameras, multiple surveillance areas are video-captured to obtain multiple surveillance videos, and the surveillance videos are framed to obtain surveillance video frame sequences corresponding to the surveillance videos; Performing target screening on the surveillance video frame sequence to obtain a surveillance video frame set; Performing target detection on the surveillance video frame set based on a pre-trained target detection model to obtain a target detection result for each surveillance video frame; Constructing a monitoring graph structure corresponding to the multiple monitoring areas based on the target detection result of each monitoring video frame; Anomaly detection is performed on the monitoring graph structure based on a pre-trained anomaly detection model to obtain an anomaly detection result corresponding to the monitoring area.

2. The method for visual target tracking based on machine learning according to claim 1, characterized in that: The target detection result set includes the detection box coordinates of the detection target, the target detection result confidence and the target category, wherein the target category includes objects, people and environments, and the monitoring graph structure includes a node set and an edge set, wherein the nodes in the node set correspond to the monitoring areas, and the edges in the edge set correspond to the spatial adjacent relationships, wherein the spatial adjacent relationships are used to indicate that the detection target is transferred from one monitoring area to another monitoring area.

3. The method for visual target tracking based on machine learning according to claim 2, characterized in that: The monitoring video frame sequence is subjected to target screening to obtain a monitoring video frame set, including: The monitoring video frame sequence is traversed, and the background of the initial monitoring video frame in the monitoring video frame sequence is modeled based on Gaussian distribution to obtain a background model corresponding to the initial monitoring video frame; When traversing to the current monitoring video frame, the background model is compared with the pixel points of the current monitoring video frame one by one to classify the pixel points of the current monitoring video frame to determine the type of the pixel points, wherein the type of the pixel points includes background points and foreground points; Counting the number of foreground points in the current monitoring video frame to obtain the number of foreground points, and calculating the retention coefficient of the current monitoring video frame based on the number of foreground points to obtain the retention coefficient of the current monitoring video frame; Comparing the retention coefficient of the current monitoring video frame with a preset retention threshold, if the retention coefficient of the current monitoring video frame is greater than the preset retention threshold, adding the current monitoring video frame to the monitoring video frame set; The background model is updated based on the current monitoring video frame to obtain a new background model, and the above operation is repeated until all monitoring video frames in the monitoring video frame sequence are traversed to obtain a monitoring video frame set corresponding to the monitoring video frame sequence.

4. The method for visual target tracking based on machine learning according to claim 3, characterized in that: The background model is as follows: Among them, x j,t represents the pixel value of the jth pixel at time t in the surveillance video frame, P represents the background distribution of the pixel, represents the weight value of the i-th Gaussian distribution at time t in the background model, represents the average value of the i-th Gaussian distribution of the j-th pixel at time t in the surveillance video frame, represents the covariance matrix of the i-th Gaussian distribution of the j-th pixel at time t in the surveillance video frame, where and represents the average pixel value of the R, G, and B components of the j-th pixel in the RGB color space at time t in the surveillance video frame, where and represents the standard deviation of the pixel values ​​of the R, G, and B components of the j-th pixel in the RGB color space at time t in the surveillance video frame, and η represents the probability density function of the Gaussian distribution, where The calculation formula of the retention coefficient is as follows: Among them, τ represents the retention coefficient, P represents the background distribution of pixels, and m*n represents the total number of pixels in the surveillance video frame; The updating formula of the background model is as follows: Among them, M i,t Indicates whether the pixel is a foreground point, M i,t =1 means the pixel is a foreground point, M i,t =0 means the pixel is a background point, X t represents the pixel value of the pixel point, and α and ρ represent the preset weight coefficients.

5. The method for visual target tracking based on machine learning according to claim 2, characterized in that: The target detection model extracts features of different scales of the surveillance video frame through three efficient convolution modules, which are represented as C1, C2 and C3 respectively, and uses multi-scale feature fusion to identify targets for features of different scales. For the minimum scale feature C3, the minimum scale feature C3 is first transformed by a 1×1 convolution operation, and then an upsampling operation is performed to obtain a feature of the same size as C2. The feature is then fused with C2 and input into an efficient convolution module to extract the fused feature. The fused feature is then mapped to the same size as C1 through an upsampling operation, and is channel-level concatenated with C1. Finally, an efficient convolution module is used to extract the final feature, and the feature is input into a classification module to obtain a set of target detection results for the surveillance video frame.

6. A visual target tracking method based on machine learning according to claim 5, characterized in that: The efficient convolution module uses two different branches to process input features, wherein the first branch first extracts the input channel dependency features through a 3×3 depth-level convolution kernel, and then extracts the input spatial dependency features through a 1×1 pixel-level convolution operation, wherein the second branch first extracts the input spatial dependency features through a 1×1 pixel-level convolution operation, and then uses a 3×3 depth-level convolution operation to extract the channel dependency features of the input features, and then uses an attention module to weight each channel to extract the weighted features, and finally uses a 1×1 pixel-level convolution operation to extract the spatial dependency features of the weighted features, and integrates the features of the two branches through channel-level cascade, and integrates the features of the two branches through a 3×3 depth-level convolution operation and a 1×1 pixel-level convolution operation.

7. The method for visual target tracking based on machine learning according to claim 1, characterized in that: The anomaly detection model adopts a graph neural network, and the anomaly detection model includes a graph convolutional network, a convolutional network, a hybrid module and an anomaly judgment module. The graph convolutional network is used to perform a convolution operation on the monitoring graph structure to obtain the time correlation characteristics of the monitoring graph structure. The convolutional network is used to perform a convolution operation on the state feature matrix of the monitoring graph structure of the preset time period to obtain the periodic characteristics of the monitoring graph structure. The hybrid module is used to perform feature fusion on the time correlation characteristics and periodic characteristics of the monitoring graph structure to obtain an anomaly probability result. The anomaly judgment module is used to judge whether the monitoring graph structure is abnormal based on the anomaly probability result.

8. The method for visual target tracking based on machine learning according to claim 7, characterized in that: The graph convolutional network includes an input layer, a first graph convolutional layer, a second graph convolutional layer, a first fully connected layer and an output layer in sequence. The convolutional network includes a first convolutional layer, a spatial attention layer, a second convolutional layer, a pooling layer and a second fully connected layer in sequence. The spatial attention layer performs average pooling and maximum pooling on the input features to obtain an average pooling vector and a maximum pooling vector. The average pooling vector and the maximum pooling vector are processed by a convolutional layer and a Sigmoid activation function to obtain a spatial attention output weight. The spatial attention output weight is multiplied by the same-position element with the input feature to obtain a spatial attention vector. The mixing module splices the time correlation feature and the periodic feature to obtain a spliced ​​vector feature. The fully connected layer based on one neuron extracts the time series feature of the spliced ​​vector feature, and the abnormal probability of the time series feature is output by the Sigmoid activation function.

Citation Information

Patent Citations

  • Non-overlapping vision field multi-camera monitoring network topology self-adaptation learning method

    CN104010168A

  • Multi-target identification monitoring method and system based on intelligent learning

    CN116524422A

  • Traffic prediction method and system based on adaptive hypergraph

    CN116739170A

  • Information real-time interaction system and method based on intelligent data processing

    CN118609059A

  • Subway compartment passenger behavior analysis method and device based on deep learning

    CN119028026A

Cited By

  • Target identification method based on monitoring video

    CN121053606A

  • Gift distribution method and gift distribution system based on robot vision

    CN121280077A

  • Gift dispensing method and gift dispensing system based on robot vision

    CN121280077B