Target identification method and system based on video calculation
By combining the dynamic adaptive frame extraction algorithm and the spatiotemporal attention feature extraction model with the YOLOv8-DANN algorithm and the MPO-MOGRPO algorithm, the problems of insufficient recognition accuracy and poor real-time performance of existing target recognition technologies under complex backgrounds and environmental changes are solved, and efficient and robust target recognition and early warning are achieved.
Patent Information
- Application Number
- CN202510717641.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing target recognition technology has insufficient recognition accuracy, poor real-time performance, insufficient robustness, and lacks an effective early warning mechanism under complex backgrounds, lighting changes, occlusion, etc., making it difficult to meet real-time and adaptability requirements.
A dynamic adaptive frame extraction algorithm and an improved DBSCAN clustering algorithm are used, combined with the spatiotemporal attention feature extraction model and the target recognition model to generate a warning strategy. The spatiotemporal attention fusion features of the key frames are extracted through the spatiotemporal attention feature extraction model, the YOLOv8-DANN algorithm is used for target recognition, and the warning strategy is generated through the MPO-MOGRPO algorithm.
It improves the accuracy and robustness of target recognition, reduces computational complexity, enhances the adaptability to complex backgrounds and environmental changes, achieves real-time and early warning effectiveness, and improves the application capability in resource-constrained environments.
Smart Images

Figure CN120655971A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition, and in particular relates to a target recognition method and system based on video computing. Background Art
[0002] Object recognition is an advanced field based on computer vision and deep learning. It can identify objects in images or videos and is widely used in various fields. With the widespread application of video surveillance technology, real-time and efficient object recognition in video data has become an important research direction.
[0003] The defects of existing technologies in target recognition mainly include:
[0004] 1) Insufficient recognition accuracy: Existing methods often experience significant declines in recognition accuracy when dealing with complex backgrounds, lighting changes, occlusions, and the need to distinguish similar targets. For example, in crowded public places, where objects occlude each other or when lighting conditions change dramatically, many algorithms struggle to accurately identify targets. They often rely on artificially designed features that struggle to fully capture the essential attributes of the target, resulting in poor recognition performance in the face of posture changes and deformation. Existing methods often struggle to effectively identify small targets in images due to low resolution and limited feature information, resulting in a high rate of missed detections.
[0005] 2) Poor real-time performance: Existing methods, especially those based on deep learning, have high recognition accuracy but also suffer from increased computational complexity, resulting in slow processing speeds and difficulty meeting real-time requirements, especially in applications with high-definition video or high frame rates. Given the massive amount of video data, efficient processing of video streams and extraction of key information are challenges faced by existing methods. Some methods require complete processing of every frame, resulting in low efficiency and inability to meet the needs of real-time applications.
[0006] 3) Insufficient robustness: Existing technologies may introduce noise during the acquisition and transmission of video data. Some algorithms are sensitive to noise, resulting in unstable recognition results and false or missed detections. Environmental changes, such as lighting, weather, and scene changes, can also affect target recognition. Existing technologies often struggle to adapt to these changes, resulting in reduced recognition performance.
[0007] 4) Lack of effective early warning mechanism: Existing methods only stay at the target identification level and lack an active early warning mechanism. They can only identify the target after it appears, and are unable to predict and warn of potential risks. Even if some methods have early warning functions, their early warning strategies are often relatively simple and difficult to adapt to different scenarios and targets, resulting in poor early warning effects. Summary of the Invention
[0008] In order to solve the problems of insufficient recognition accuracy, poor real-time performance, insufficient robustness and lack of effective early warning mechanism in the existing technology, the purpose of the present invention is to provide a target recognition method and system based on video computing.
[0009] The technical solution adopted in the present invention is:
[0010] A target recognition method based on video computing includes the following steps:
[0011] Collect real-time video data and use a dynamic adaptive frame extraction algorithm to extract several real-time key frames of the real-time video stream;
[0012] Use the spatiotemporal attention feature extraction model to extract real-time spatiotemporal attention fusion features of several real-time key frames;
[0013] According to the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results;
[0014] According to the real-time target recognition results, the early warning strategy generation model is used to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy.
[0015] Furthermore, real-time video data is collected, and a dynamic adaptive frame extraction algorithm is used to extract several real-time key frames of the real-time video stream, including the following steps:
[0016] Collect real-time video data, convert the real-time video data into a real-time video stream, and obtain a real-time motion energy distribution map of the real-time video stream;
[0017] According to the real-time motion energy distribution map, the improved DBSCAN clustering algorithm is used to divide the real-time video stream into several real-time motion areas;
[0018] According to the real-time motion area, the real-time motion area density of the real-time video stream is obtained, and according to the real-time motion area density, the preset frame extraction frequency is dynamically adjusted to obtain the real-time frame extraction frequency;
[0019] According to the real-time frame extraction frequency, the real-time video stream is dynamically and adaptively extracted to obtain several real-time key frames.
[0020] Furthermore, collecting real-time video data, converting the real-time video data into a real-time video stream, and obtaining a real-time motion energy distribution map of the real-time video stream include the following steps:
[0021] Collect real-time video data, convert the real-time video data into a real-time video stream, and perform frame interception on the real-time video stream to obtain a number of real-time frame images;
[0022] Using the optical flow method, the real-time motion vector field of two consecutive frames of real-time frame images in the real-time video stream is obtained;
[0023] According to the real-time motion vector field, the real-time motion amplitude of each real-time pixel point in the real-time frame image is obtained;
[0024] Performing statistical analysis on the real-time motion amplitudes of all real-time pixel points in the real-time frame image to obtain an initial real-time motion energy distribution map of the real-time video stream;
[0025] The initial real-time motion energy distribution map is normalized to obtain a final real-time motion energy distribution map of the real-time video stream.
[0026] Furthermore, based on the real-time motion energy distribution map, an improved DBSCAN clustering algorithm is used to divide the real-time video stream into several real-time motion regions, including the following steps:
[0027] Use the NMS algorithm to extract several real-time feature points of the real-time motion energy distribution map, and extract the local real-time motion energy features around the real-time feature points as the corresponding real-time descriptors;
[0028] Using the improved DBSCAN algorithm based on the spatiotemporal constraint mechanism, several real-time feature points are clustered and several real-time feature points with spatiotemporal proximity are divided into the same motion region to obtain several initial real-time motion regions.
[0029] Several initial real-time motion regions are uniquely marked to obtain several final real-time motion regions of the real-time video stream.
[0030] Furthermore, according to the real-time motion area, the real-time motion area density of the real-time video stream is obtained, and according to the real-time motion area density, the preset frame frequency is dynamically adjusted to obtain the real-time frame frequency, which includes the following steps:
[0031] Obtaining the real-time motion region density of each real-time motion region;
[0032] Performing statistical analysis on the real-time motion region density of all real-time motion regions to obtain the real-time motion region density distribution;
[0033] According to the density distribution of the real-time motion area, the preset frame frequency is dynamically adjusted to obtain the real-time frame frequency.
[0034] Furthermore, the spatiotemporal attention feature extraction model is constructed based on the ST-Attention-CNN algorithm, and the spatiotemporal attention feature extraction model includes a spatial attention feature extraction channel constructed based on the CBAM algorithm, a temporal attention feature extraction channel constructed based on the 3D-CNN-Attention algorithm, and a feature fusion layer constructed based on the CNN algorithm. The spatial attention feature extraction channel and the temporal attention feature extraction channel are both connected to the feature fusion layer, and the spatial attention feature extraction channel includes a channel attention module and a spatial attention module connected in sequence, and the temporal attention feature extraction channel includes a 3D convolution module and a temporal self-attention module connected in sequence.
[0035] The target recognition model is built based on the YOLOv8-DANN algorithm, and the target recognition model includes a target recognition module built based on the YOLOv8 algorithm and a domain adversarial training module built based on the DANN algorithm, which are connected in sequence;
[0036] The early warning strategy generation model is constructed based on the MPO-MOGRPO algorithm, and the early warning strategy generation model includes a meta-strategy optimization module constructed based on the MPO algorithm and a warning strategy generation module constructed based on the MOGRPO algorithm. The early warning strategy generation module includes an objective function set, a policy network, an experience replay pool and an intelligent agent. The intelligent agent is connected to the objective function set and the policy network respectively, and the meta-strategy optimization module is connected to the policy network.
[0037] Furthermore, a spatiotemporal attention feature extraction model is used to extract real-time spatiotemporal attention fusion features of several real-time key frames, including the following steps:
[0038] Use the spatial attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time spatial attention features for each real-time keyframe;
[0039] Use the temporal attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time temporal attention features of several consecutive real-time key frames;
[0040] The feature fusion layer of the spatiotemporal attention feature extraction model is used to fuse the real-time spatial attention features and the real-time temporal attention features to obtain the real-time spatiotemporal attention fusion features of each real-time key frame.
[0041] Furthermore, based on the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results, including the following steps:
[0042] The real-time spatiotemporal attention fusion features of each real-time keyframe are sequentially input into the target recognition model;
[0043] According to the real-time spatiotemporal attention fusion features, the target recognition module of the target recognition model is used to perform target recognition and obtain the real-time recognition target;
[0044] The real-time spatiotemporal attention fusion features of all real-time key frames are traversed, and the SORT algorithm is used to perform single target tracking on the same real-time recognition target to obtain the real-time target recognition result.
[0045] Furthermore, based on the real-time target recognition result, the early warning strategy generation model is used to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy, including the following steps:
[0046] Based on the real-time target recognition results, the meta-strategy optimization module of the early warning strategy generation model is used to update the strategy network of the early warning strategy generation module to obtain an updated strategy network;
[0047] According to the real-time target recognition results, the state space of the intelligent agent of the early warning strategy generation module is updated to obtain an updated state space;
[0048] Randomly extract several historical early warning strategy generation experiences from the experience replay pool, and update the action space of the intelligent agent of the early warning strategy generation module based on the several historical early warning strategy generation experiences to obtain an updated action space;
[0049] A real-time objective function is selected from the objective function set. Based on the real-time objective function, according to the updated state space and the updated action space, the intelligent agent is used to control the updated policy network to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy.
[0050] A target recognition system based on video computing is used to implement a target recognition method. The system includes a key frame extraction unit, a fusion feature extraction unit, a target recognition unit and an early warning strategy generation unit which are connected in sequence.
[0051] The beneficial effects of the present invention are:
[0052] The present invention discloses a target recognition method and system based on video computing. Through the spatiotemporal attention feature extraction model, it can better capture the spatiotemporal context information of the target, enhance the model's adaptability to complex backgrounds, lighting changes, occlusions and the like, and thus improve recognition accuracy; the spatiotemporal attention mechanism can automatically focus on the key areas of the target and extract more discriminative features. Compared with traditional manually designed features, it can better reflect the essential attributes of the target and effectively improve the recognition effect under conditions of posture changes and deformations; combined with the dynamic adaptive frame extraction algorithm and the improved DBSCAN clustering algorithm, it can more effectively detect and track small targets and reduce the missed detection rate; the dynamic adaptive frame extraction algorithm dynamically adjusts the frame extraction frequency according to the video content, reduces the amount of data to be processed, reduces the computational complexity, and improves the processing speed; by identifying and tracking real-time key frames, the inefficiency of frame-by-frame processing is avoided, and the data processing is improved The algorithm can improve the processing efficiency and better meet the real-time requirements; the spatiotemporal attention mechanism can help the model focus on the target area, suppress noise interference, and improve the robustness of recognition; by detecting and tracking the real-time motion area, it can better adapt to environmental changes such as lighting, weather, scene switching, etc., to ensure the stability of recognition performance; the warning strategy generation model based on reinforcement learning can actively predict potential risks based on target recognition results and historical data, and realize advance warning rather than passive recognition; through reinforcement learning, the warning strategy generation model can adaptively adjust the warning strategy according to different scenarios and targets, thereby improving the accuracy and effectiveness of the warning; by learning more generalized spatiotemporal features, it can better adapt to different scenarios and data and improve the generalization ability of the model; the dynamic adaptive frame extraction algorithm and efficient recognition and tracking algorithm reduce the algorithm's demand for computing resources, enabling it to be applied in resource-constrained environments.
[0053] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flowchart of the target recognition method based on video computing in the present invention.
[0055] Figure 2 It is a structural block diagram of the target recognition system based on video computing in the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0057] Example 1:
[0058] like Figure 1 As shown, this embodiment provides a target recognition method based on video computing, including the following steps:
[0059] S1: Collect real-time video data and use a dynamic adaptive frame extraction algorithm to extract several real-time key frames from the real-time video stream, including the following steps:
[0060] S1-1: Collecting real-time video data, converting the real-time video data into a real-time video stream, and obtaining a real-time motion energy distribution map of the real-time video stream, including the following steps:
[0061] S1-1-1: Collect real-time video data, convert the real-time video data into a real-time video stream, and perform frame interception on the real-time video stream to obtain several real-time frame images;
[0062] In this embodiment, the frame rate of the real-time video stream is F frames / second;
[0063] S1-1-2: Use the optical flow method to obtain two consecutive frames I in the real-time video stream t , I t-1 The real-time motion vector field of the real-time frame image, where t is the frame indicator;
[0064] The formula is:
[0065]
[0066] Where u(x,y) is the real-time motion vector field, that is, the real-time displacement vector of the pixel point (x,y); (x,y) is the coordinate of the real-time pixel point; u is the real-time displacement vector; is the real-time frame image I of the tth frame t The gradient at position (x+u); is the real-time frame image I of the t-1 frame t-1 The gradient at position x;
[0067] S1-1-3: Obtain the real-time motion amplitude of each real-time pixel in the real-time frame image according to the real-time motion vector field;
[0068] The formula is:
[0069] M(x,y)=||u(x,y)||
[0070] Where M(x,y) is the real-time motion amplitude of the real-time pixel point (x,y);
[0071] S1-1-4: Statistically analyze the real-time motion amplitudes of all real-time pixel points in the real-time frame image to obtain an initial real-time motion energy distribution map of the real-time video stream;
[0072] The formula is:
[0073]
[0074] Where E(x,y,t) is the initial real-time motion energy distribution of the real-time pixel point (x,y) in the tth frame; M(x,y,τ) is the real-time motion amplitude of the real-time pixel point (x,y) in the τth frame; τ is the frame indicator;
[0075] S1-1-5: normalizing the initial real-time motion energy distribution map to obtain a final real-time motion energy distribution map of the real-time video stream;
[0076] The formula is:
[0077]
[0078] Where E'(x,y,t) is the final real-time motion energy distribution map of the real-time pixel point (x,y) in the tth frame;
[0079] S1-2: Based on the real-time motion energy distribution map, an improved density-based spatial clustering of applications with noise (DBSCAN) clustering algorithm is used to divide the real-time video stream into several real-time motion regions, including the following steps:
[0080] S1-2-1: Use the non-maximum suppression (NMS) algorithm to extract several real-time feature points from the real-time motion energy distribution map, and extract the local real-time motion energy features around the real-time feature points as the corresponding real-time descriptors;
[0081] By searching for local maxima, it identifies significant feature points in the image and suppresses non-maximum points in the neighborhood, thereby avoiding interference from redundant information. In tasks such as target detection and feature extraction, NMS can remove overlapping or repeated feature points and retain the most representative points, thereby improving the accuracy and efficiency of the results.
[0082] Real-time descriptors can characterize the local motion energy distribution around feature points, such as the intensity, direction, and change trend of motion. Through descriptors, local information of motion feature points can be more comprehensively portrayed, improving the stability and distinguishability of feature points in subsequent analysis. Real-time descriptors provide auxiliary judgment for subsequent motion area division, ensuring the real-time and accuracy of motion areas.
[0083] S1-2-2: Use the improved DBSCAN algorithm based on the spatiotemporal constraint mechanism to cluster several real-time feature points, divide several real-time feature points with spatiotemporal proximity into the same motion region, and obtain several initial real-time motion regions;
[0084] Based on the traditional DBSCAN algorithm, spatiotemporal constraints are introduced, that is, the proximity relationship of feature points in time and space is considered. Temporal constraint: if two feature points are temporally adjacent (i.e., the frame interval is less than a threshold), they are considered temporally adjacent. Spatial constraint: if the spatial distance between two feature points is less than a threshold, they are considered spatially adjacent.
[0085] S1-2-3: uniquely marking a number of initial real-time motion regions to obtain a number of final real-time motion regions of the real-time video stream;
[0086] S1-3: Obtaining the real-time motion region density of the real-time video stream according to the real-time motion region, and dynamically adjusting the preset frame extraction frequency according to the real-time motion region density to obtain the real-time frame extraction frequency, including the following steps:
[0087] S1-3-1: Obtain the real-time motion area density of each real-time motion area;
[0088] The formula is:
[0089]
[0090] Where, ρ(R t,i ) is the real-time motion region R of the t-th frame t,i Real-time motion area density; R t,i is the i-th real-time motion region of the t-th frame; i is the real-time motion region indicator; N t,i is the real-time motion region R of the tth frame t,i The number of feature points; A t,i is the real-time motion region R of the tth frame t,i The area (i.e. the number of pixels it contains);
[0091] S1-3-2: Statistically analyze the real-time motion region density of all real-time motion regions to obtain the real-time motion region density distribution;
[0092] The real-time motion area density distribution includes the real-time motion area density mean and the real-time motion area density standard deviation;
[0093] The formula is:
[0094]
[0095] Where μ ρ,t is the mean density of the real-time motion region in the tth frame; N is the total number of real-time motion regions;
[0096]
[0097] Where, σ ρ,t is the standard deviation of the real-time motion area density of the t-th frame;
[0098] S1-3-3: Dynamically adjust the preset frame rate according to the density distribution of the real-time motion area to obtain the real-time frame rate;
[0099] The formula is:
[0100]
[0101] Where, f t is the real-time frame extraction frequency; f base is the preset frame frequency; ε is a small positive number;
[0102] S1-4: Dynamically and adaptively extract frames from the real-time video stream according to the real-time frame extraction frequency to obtain several real-time key frames;
[0103] S2: Use the spatiotemporal attention feature extraction model to extract real-time spatiotemporal attention fusion features of several real-time key frames;
[0104] The spatiotemporal attention feature extraction model is constructed based on the Spatial-TemporalAttention Convolutional Neural Network (ST-Attention-CNN) algorithm, and the spatiotemporal attention feature extraction model includes a spatial attention feature extraction channel constructed based on the Convolutional Block Attention Module (CBAM) algorithm, a temporal attention feature extraction channel constructed based on the 3D-CNN-Attention algorithm, and a feature fusion layer constructed based on the CNN algorithm. The spatial attention feature extraction channel and the temporal attention feature extraction channel are both connected to the feature fusion layer, and the spatial attention feature extraction channel includes a channel attention module and a spatial attention module connected in sequence, and the temporal attention feature extraction channel includes a 3D convolution module and a temporal self-attention module connected in sequence.
[0105] The channel attention module of the spatial attention feature extraction channel is used to identify and strengthen important channel information in the input feature map, while suppressing unimportant channels. By learning the importance of different channels, the module can highlight the features that are most helpful for target recognition. The spatial attention module is used to focus on the spatial position information in the feature map, strengthen the spatial features of the target area, and suppress the interference of the background area. Through the dual attention mechanism of channels and space, the model can more accurately capture and strengthen the key features of the target, improve the expressiveness and discrimination of the features, suppress the interference of irrelevant channels and background areas, make the model more focused on the target area, and enhance the anti-interference ability in complex backgrounds. The spatial attention module can highlight the spatial features of small targets, making them more prominent in the feature map, thereby improving the detection performance of small targets;
[0106] The 3D convolution module of the temporal attention feature extraction channel is used to extract spatiotemporal features from video sequences. The 3D convolution operation simultaneously captures spatial and temporal information, enhancing the model's ability to perceive dynamic changes. The temporal self-attention module is used to learn the relationship between different frames in the video sequence, strengthen the features of key frames, and suppress the interference of redundant frames. The 3D convolution module can capture dynamic changes in the video, enabling the model to better adapt to changes in target posture, speed, etc., and improve recognition accuracy. Through the temporal self-attention module, the model can focus on key frames, reduce the processing of redundant frames, improve data processing efficiency, and enhance real-time performance. The temporal self-attention module can learn the relationship between frames, enhance the model's ability to model video sequences, and improve recognition performance in complex sequence scenarios;
[0107] The feature fusion layer fuses the output features of the spatial attention feature extraction channel and the temporal attention feature extraction channel to form the final feature representation. By fusing spatial and temporal features, the model can integrate spatiotemporal information to form a more comprehensive and accurate feature representation, thereby improving recognition performance. The fused features contain richer information, making the model more robust to interference such as noise and illumination changes. The integration of spatiotemporal features helps the model learn more generalizable feature representations, improving its generalization capabilities across different scenarios and data.
[0108] Use the spatiotemporal attention feature extraction model to extract real-time spatiotemporal attention fusion features of several real-time keyframes, including the following steps:
[0109] S2-1: Use the spatial attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time spatial attention features for each real-time keyframe;
[0110] S2-2: Use the temporal attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time temporal attention features of several consecutive real-time keyframes;
[0111] S2-3: Use the feature fusion layer of the spatiotemporal attention feature extraction model to fuse the real-time spatial attention features and the real-time temporal attention features to obtain the real-time spatiotemporal attention fusion features of each real-time keyframe;
[0112] S3: Based on the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results;
[0113] The target recognition model is built based on the YOLOv8-Domain-Adversarial Neural Network (DANN) algorithm, and the target recognition model includes a target recognition module built based on the YOLOv8 algorithm and a domain adversarial training module built based on the DANN algorithm, which are connected in sequence.
[0114] The YOLOv8 object recognition module is an efficient object detection algorithm that directly predicts the category and location of the object from the input image through a single-stage detection mechanism. YOLOv8 uses an advanced backbone network and neck architecture to quickly process images and achieve high-precision object detection.
[0115] The DANN in the domain adversarial training module is a domain adaptation algorithm that can align features between the source domain (labeled data) and the target domain (unlabeled data), reducing the impact of domain differences on model performance. Through the Gradient Reversal Layer (GRL) and adversarial training mechanism, DANN can learn universal features that are independent of the domain. Through domain adaptation, the model can maintain high recognition accuracy in the target domain. Even if the target domain data distribution is different from the source domain, DANN can effectively reduce domain differences caused by changes in lighting, viewing angle, etc., improving the model's robustness in complex environments. In the absence of labeled data in the target domain, DANN can still achieve effective target recognition through the adversarial training mechanism, reducing data labeling costs.
[0116] Based on the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results, including the following steps:
[0117] S3-1: The real-time spatiotemporal attention fusion features of each real-time keyframe are sequentially input into the target recognition model;
[0118] S3-2: Based on the real-time spatiotemporal attention fusion features, the target recognition module of the target recognition model is used to perform target recognition and obtain the real-time recognition target;
[0119] S3-3: Traverse the real-time spatiotemporal attention fusion features of all real-time keyframes and use the Simple Online and Realtime Tracking (SORT) algorithm to track the same real-time recognition target to obtain the real-time target recognition result;
[0120] S4: Based on the real-time target recognition results, use the early warning strategy generation model to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy;
[0121] The early warning strategy generation model is constructed based on the Meta-Policy Optimization (MPO)-Multi-Objective Group Relative Policy Optimization (MOGRPO) algorithm. The early warning strategy generation model includes a meta-policy optimization module constructed based on the MPO algorithm and an early warning strategy generation module constructed based on the MOGRPO algorithm. The early warning strategy generation module includes an objective function set, a policy network, an experience replay pool, and an intelligent agent. The intelligent agent is connected to the objective function set and the policy network respectively, and the meta-policy optimization module is connected to the policy network.
[0122] The meta-strategy optimization module is used to adjust the network parameters of the policy network in the early warning strategy generation module so that these parameters can quickly adapt to new and unseen target recognition results, thereby improving the generalization ability of the model. Even under unseen target recognition results, the policy network can be updated based on previous learning experience, thereby improving the adaptability of the early warning strategy generation model. The objective function set of the early warning strategy generation module can handle multiple conflicting objectives, such as minimizing early warning costs, maximizing early warning accuracy, maximizing early warning efficiency, etc., and generate early warning strategies that balance these objectives. The intelligent agent learns historical early warning strategies through the experience replay pool and continuously optimizes its own strategy generation capabilities. The intelligent agent controls the policy network based on the learned experience to generate more effective early warning strategies. The design of the experience replay pool and the intelligent agent enables the model to continuously learn and optimize, thereby improving the quality of strategy generation. The early warning strategy generation module can avoid falling into local optimal solutions to a certain extent due to the use of a group exploration method. The policy network outputs the distribution probability of actions under a given state. The early warning strategy generation module directly updates the policy network through gradients, eliminating the Critic network in traditional reinforcement learning, making the algorithm structure more concise.
[0123] Based on the real-time target recognition results, the early warning strategy generation model is used to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy, including the following steps:
[0124] S4-1: Based on the real-time target recognition results, the meta-strategy optimization module of the early warning strategy generation model is used to update the policy network of the early warning strategy generation module to obtain an updated policy network;
[0125] S4-2: Based on the real-time target recognition results, the state space of the intelligent agent of the early warning strategy generation module is updated to obtain an updated state space;
[0126] S4-3: Randomly extract several historical early warning strategy generation experiences from the experience replay pool, and update the action space of the agent of the early warning strategy generation module based on the several historical early warning strategy generation experiences to obtain an updated action space;
[0127] S4-4: Select a real-time objective function from the objective function set. Based on the real-time objective function, use the agent to control the updated policy network according to the updated state space and updated action space, generate a warning strategy, obtain a real-time warning strategy, and execute the real-time warning strategy. The steps include:
[0128] S4-4-1: Select a real-time objective function from the objective function set, and based on the real-time objective function, use the agent of the warning strategy generation module to control the updated policy network to generate the probability distribution of all possible warning actions in the updated action space corresponding to each real-time target recognition state in the updated state space;
[0129] S4-4-2: The possible warning action with the highest probability distribution in the updated action space is used as the execution warning action of the corresponding real-time target recognition state;
[0130] S4-4-3: Integrate the execution warning actions of all real-time target recognition states in the updated state space to obtain the real-time warning strategy, and execute the real-time warning strategy.
[0131] Example 2:
[0132] like Figure 2 As shown, this embodiment provides a target recognition system based on video computing, which is used to implement a target recognition method. The system includes a key frame extraction unit, a fusion feature extraction unit, a target recognition unit, and a warning strategy generation unit connected in sequence;
[0133] A key frame extraction unit is used to collect real-time video data and use a dynamic adaptive frame extraction algorithm to extract several real-time key frames of the real-time video stream;
[0134] A fusion feature extraction unit, configured to extract real-time spatiotemporal attention fusion features of a plurality of real-time key frames using a spatiotemporal attention feature extraction model;
[0135] The target recognition unit is used to perform target recognition based on the real-time spatiotemporal attention fusion features and the target recognition model to obtain real-time target recognition results;
[0136] The early warning strategy generation unit is used to generate early warning strategies based on the real-time target recognition results using the early warning strategy generation model to obtain the real-time early warning strategies and execute the real-time early warning strategies.
[0137] The present invention discloses a target recognition method and system based on video computing. Through the spatiotemporal attention feature extraction model, it can better capture the spatiotemporal context information of the target, enhance the model's adaptability to complex backgrounds, lighting changes, occlusions and the like, and thus improve recognition accuracy; the spatiotemporal attention mechanism can automatically focus on the key areas of the target and extract more discriminative features. Compared with traditional manually designed features, it can better reflect the essential attributes of the target and effectively improve the recognition effect under conditions of posture changes and deformations; combined with the dynamic adaptive frame extraction algorithm and the improved DBSCAN clustering algorithm, it can more effectively detect and track small targets and reduce the missed detection rate; the dynamic adaptive frame extraction algorithm dynamically adjusts the frame extraction frequency according to the video content, reduces the amount of data to be processed, reduces the computational complexity, and improves the processing speed; by identifying and tracking real-time key frames, the inefficiency of frame-by-frame processing is avoided, and the data processing is improved The algorithm can improve the processing efficiency and better meet the real-time requirements; the spatiotemporal attention mechanism can help the model focus on the target area, suppress noise interference, and improve the robustness of recognition; by detecting and tracking the real-time motion area, it can better adapt to environmental changes such as lighting, weather, scene switching, etc., to ensure the stability of recognition performance; the warning strategy generation model based on reinforcement learning can actively predict potential risks based on target recognition results and historical data, and realize advance warning rather than passive recognition; through reinforcement learning, the warning strategy generation model can adaptively adjust the warning strategy according to different scenarios and targets, thereby improving the accuracy and effectiveness of the warning; by learning more generalized spatiotemporal features, it can better adapt to different scenarios and data and improve the generalization ability of the model; the dynamic adaptive frame extraction algorithm and efficient recognition and tracking algorithm reduce the algorithm's demand for computing resources, enabling it to be applied in resource-constrained environments.
[0138] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.
Claims
1. A target recognition method based on video computing, characterized by: The steps include: Collect real-time video data and use a dynamic adaptive frame extraction algorithm to extract several real-time key frames of the real-time video stream; Use the spatiotemporal attention feature extraction model to extract real-time spatiotemporal attention fusion features of several real-time key frames; According to the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results; According to the real-time target recognition results, the early warning strategy generation model is used to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy.
2. The target recognition method based on video computing according to claim 1, characterized in that: Collect real-time video data and use a dynamic adaptive frame extraction algorithm to extract several real-time key frames of the real-time video stream, including the following steps: Collect real-time video data, convert the real-time video data into a real-time video stream, and obtain a real-time motion energy distribution map of the real-time video stream; According to the real-time motion energy distribution map, the improved DBSCAN clustering algorithm is used to divide the real-time video stream into several real-time motion areas; According to the real-time motion area, the real-time motion area density of the real-time video stream is obtained, and according to the real-time motion area density, the preset frame extraction frequency is dynamically adjusted to obtain the real-time frame extraction frequency; According to the real-time frame extraction frequency, the real-time video stream is dynamically and adaptively extracted to obtain several real-time key frames.
3. The target recognition method based on video computing according to claim 2, characterized in that: Collecting real-time video data, converting the real-time video data into a real-time video stream, and obtaining a real-time motion energy distribution map of the real-time video stream includes the following steps: Collect real-time video data, convert the real-time video data into a real-time video stream, and perform frame interception on the real-time video stream to obtain a number of real-time frame images; Using the optical flow method, the real-time motion vector field of two consecutive frames of real-time frame images in the real-time video stream is obtained; According to the real-time motion vector field, the real-time motion amplitude of each real-time pixel point in the real-time frame image is obtained; Performing statistical analysis on the real-time motion amplitudes of all real-time pixel points in the real-time frame image to obtain an initial real-time motion energy distribution map of the real-time video stream; The initial real-time motion energy distribution map is normalized to obtain a final real-time motion energy distribution map of the real-time video stream.
4. The method for target recognition based on video computing according to claim 3, characterized in that: According to the real-time motion energy distribution map, the improved DBSCAN clustering algorithm is used to divide the real-time video stream into several real-time motion areas, including the following steps: Use the NMS algorithm to extract several real-time feature points of the real-time motion energy distribution map, and extract the local real-time motion energy features around the real-time feature points as the corresponding real-time descriptors; Using the improved DBSCAN algorithm based on the spatiotemporal constraint mechanism, several real-time feature points are clustered and several real-time feature points with spatiotemporal proximity are divided into the same motion region to obtain several initial real-time motion regions. Several initial real-time motion regions are uniquely marked to obtain several final real-time motion regions of the real-time video stream.
5. The method for target recognition based on video computing according to claim 4, characterized in that: According to the real-time motion area, the real-time motion area density of the real-time video stream is obtained, and according to the real-time motion area density, the preset frame extraction frequency is dynamically adjusted to obtain the real-time frame extraction frequency, including the following steps: Obtaining the real-time motion region density of each real-time motion region; Performing statistical analysis on the real-time motion region density of all real-time motion regions to obtain the real-time motion region density distribution; According to the density distribution of the real-time motion area, the preset frame frequency is dynamically adjusted to obtain the real-time frame frequency.
6. The method for target recognition based on video computing according to claim 5, characterized in that: The spatiotemporal attention feature extraction model is constructed based on the ST-Attention-CNN algorithm, and the spatiotemporal attention feature extraction model includes a spatial attention feature extraction channel constructed based on the CBAM algorithm, a temporal attention feature extraction channel constructed based on the 3D-CNN-Attention algorithm, and a feature fusion layer constructed based on the CNN algorithm. The spatial attention feature extraction channel and the temporal attention feature extraction channel are both connected to the feature fusion layer, and the spatial attention feature extraction channel includes a channel attention module and a spatial attention module connected in sequence, and the temporal attention feature extraction channel includes a 3D convolution module and a temporal self-attention module connected in sequence; The target recognition model is constructed based on the YOLOv8-DANN algorithm, and the target recognition model includes a target recognition module constructed based on the YOLOv8 algorithm and a domain adversarial training module constructed based on the DANN algorithm, which are connected in sequence; The early warning strategy generation model is constructed based on the MPO-MOGRPO algorithm, and the early warning strategy generation model includes a meta-strategy optimization module constructed based on the MPO algorithm and a early warning strategy generation module constructed based on the MOGRPO algorithm. The early warning strategy generation module includes an objective function set, a strategy network, an experience replay pool and an intelligent agent. The intelligent agent is connected to the objective function set and the strategy network respectively, and the meta-strategy optimization module is connected to the strategy network.
7. The method for target recognition based on video computing according to claim 6, characterized in that: Use the spatiotemporal attention feature extraction model to extract real-time spatiotemporal attention fusion features of several real-time keyframes, including the following steps: Use the spatial attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time spatial attention features for each real-time keyframe; Use the temporal attention feature extraction channel of the spatiotemporal attention feature extraction model to extract real-time temporal attention features of several consecutive real-time key frames; The feature fusion layer of the spatiotemporal attention feature extraction model is used to fuse the real-time spatial attention features and the real-time temporal attention features to obtain the real-time spatiotemporal attention fusion features of each real-time key frame.
8. The method for target recognition based on video computing according to claim 7, characterized in that: Based on the real-time spatiotemporal attention fusion features, the target recognition model is used to perform target recognition and obtain real-time target recognition results, including the following steps: The real-time spatiotemporal attention fusion features of each real-time keyframe are sequentially input into the target recognition model; According to the real-time spatiotemporal attention fusion features, the target recognition module of the target recognition model is used to perform target recognition and obtain the real-time recognition target; The real-time spatiotemporal attention fusion features of all real-time key frames are traversed, and the SORT algorithm is used to perform single target tracking on the same real-time recognition target to obtain the real-time target recognition result.
9. The method for target recognition based on video computing according to claim 8, characterized in that: Based on the real-time target recognition results, the early warning strategy generation model is used to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy, including the following steps: Based on the real-time target recognition results, the meta-strategy optimization module of the early warning strategy generation model is used to update the strategy network of the early warning strategy generation module to obtain an updated strategy network; According to the real-time target recognition results, the state space of the intelligent agent of the early warning strategy generation module is updated to obtain an updated state space; Randomly extract several historical early warning strategy generation experiences from the experience replay pool, and update the action space of the intelligent agent of the early warning strategy generation module based on the several historical early warning strategy generation experiences to obtain an updated action space; A real-time objective function is selected from the objective function set. Based on the real-time objective function, according to the updated state space and the updated action space, the intelligent agent is used to control the updated policy network to generate the early warning strategy, obtain the real-time early warning strategy, and execute the real-time early warning strategy.
10. A target recognition system based on video computing, used to implement the target recognition method according to any one of claims 1 to 9, characterized in that: The system comprises a key frame extraction unit, a fusion feature extraction unit, a target recognition unit and an early warning strategy generation unit which are connected in sequence.