Visual multi-target tracking method, system and electronic device based on deep learning

By deploying cameras at the entrances and exits of the mall, using deep learning technology to extract and integrate monitoring feature maps, we can determine whether there are abnormal personnel in the mall, and solve the problem that traditional artificial visual monitoring is difficult to meet the security needs of large areas, and an efficient and intelligent monitoring system is realized.

CN118644804BActive Publication Date: 2025-05-16SHENZHEN ENGINEER-LINK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410922436.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-05-16
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

Traditional artificial vision is difficult to meet the monitoring needs of large areas, there are security bottlenecks, and it is difficult to track and identify abnormal people in the mall in real time.

Method used

Using a visual multi-objective tracking method based on deep learning, the camera is deployed to collect surveillance videos at the entrance and exit of the mall, the entrance and exit monitoring feature maps are extracted, and the entrance and exit monitoring feature maps are fused into a multi-objective monitoring feature map, and whether there are abnormal people in the mall are judged through the classifier.

Benefits of technology

It realizes high-quality robust tracking, improves the intelligence level of the monitoring system, can identify and judge abnormal personnel in the mall in real time, and improves the efficiency and accuracy of security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118644804B_ABST
    Figure CN118644804B_ABST
Patent Text Reader

Abstract

The present application relates to the field of visual multi-target technology, and more specifically discloses a visual multi-target tracking method, system and electronic device based on deep learning, which collects monitoring videos by deploying cameras at the entrance and exit of a shopping mall, extracts entrance monitoring feature maps and exit monitoring feature maps from them, and then fuses these feature maps into multi-target monitoring feature maps, and classifies them through a classifier to determine whether there are abnormal people in the shopping mall. Taking advantage of deep learning, by automatically learning the appearance characteristics and motion characteristics of the target, high-quality robust tracking is achieved, and the intelligence level of the monitoring system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present application relates to the field of visual multi-target technology, and more specifically, to a visual multi-target tracking method, system and electronic device based on deep learning. Background Art

[0002] With the rapid development of science and technology and the economy, large areas (such as shopping malls) are becoming increasingly numerous and expanding, and the requirements for surveillance in these areas are also becoming increasingly demanding. For example, current surveillance systems in shopping malls are relatively simple, relying primarily on manual visual monitoring to provide security information on screens. Due to the increasing number of shoppers in shopping malls, traditional manual visual monitoring and management methods present security bottlenecks, and inadequate supervision can easily lead to safety incidents. Therefore, traditional manual visual monitoring is no longer sufficient to meet the monitoring needs of large areas, and there is an urgent need to improve the intelligence level of video surveillance systems for large areas.

[0003] Visual object tracking is a hot topic in computer vision research. With the rapid development of computer technology, object tracking technology has also made significant progress. With the rapid rise of artificial intelligence in recent years, research on object tracking technology has received increasing attention. Deep learning technology has powerful feature representation capabilities and has achieved better results than traditional methods in applications such as image classification, object recognition, and natural language processing. Therefore, it has gradually become a mainstream technology in image and video research. Deep learning-based tracking methods are an important branch of object tracking methods. They leverage the advantages of end-to-end training of deep convolutional networks to enable the model to automatically learn the appearance and motion characteristics of the target to be tracked, achieving high-quality robust tracking.

[0004] Therefore, a deep learning-based visual multi-target tracking method, system, and electronic device are desired. Summary of the Invention

[0005] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide a deep learning-based visual multi-target tracking method, system, and electronic device, which fuses entrance and exit monitoring feature maps and uses a classifier to determine whether there are abnormal people in the mall.

[0006] Accordingly, according to one aspect of the present application, a deep learning-based visual multi-target tracking method is provided, which includes:

[0007] Obtain multiple entrance surveillance videos and multiple exit surveillance videos collected by cameras deployed at the entrances and exits of the shopping mall;

[0008] Extracting an entrance monitoring feature map and an exit monitoring feature map from the plurality of entrance monitoring videos and the plurality of exit monitoring videos respectively;

[0009] fusing the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map;

[0010] Performing a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector;

[0011] Performing reverse feature consistency enhancement based on the class regression domain on the multi-target monitoring feature vector to obtain an optimized multi-target monitoring feature vector;

[0012] The optimized multi-target monitoring feature vector is passed through a classifier to obtain a classification result, and the classification result is used to indicate whether there are abnormal people in the shopping mall.

[0013] According to another aspect of the present application, a deep learning-based visual multi-target tracking system is provided, comprising:

[0014] The camera acquisition module is used to obtain multiple entrance surveillance videos and multiple exit surveillance videos collected by cameras deployed at the entrances and exits of the shopping mall;

[0015] A feature map extraction module is used to extract an entry monitoring feature map and an exit monitoring feature map from the multiple entry monitoring videos and the multiple exit monitoring videos respectively;

[0016] A fusion multi-target module, used for fusing the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map;

[0017] A feature map pooling module, configured to perform a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector;

[0018] A multi-objective feature optimization module, configured to perform reverse feature consistency enhancement based on a class regression domain on the multi-objective monitoring feature vector to obtain an optimized multi-objective monitoring feature vector;

[0019] The multi-objective result analysis module is used to pass the optimized multi-objective monitoring feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether there are abnormal people in the shopping mall.

[0020] According to another aspect of the present application, an electronic device is provided, comprising: a processor; a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the deep learning-based visual multi-target tracking method as described in any one of the above.

[0021] Compared to existing technologies, this application provides a deep learning-based visual multi-target tracking method, system, and electronic device. This method uses cameras deployed at the entrance and exit of a shopping mall to capture surveillance video, extracting entrance and exit monitoring feature maps from these videos, and then fusing these feature maps into a multi-target monitoring feature map. This map is then classified using a classifier to determine whether there are any unusual individuals in the mall. Leveraging the advantages of deep learning, this method automatically learns the appearance and motion characteristics of targets, achieving high-quality robust tracking and enhancing the intelligence of the surveillance system. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0023] Figure 1 Flowchart of a deep learning-based visual multi-target tracking method according to an embodiment of the present application.

[0024] Figure 2 This is a flowchart of extracting entrance monitoring feature maps and exit monitoring feature maps from the multiple entrance monitoring videos and the multiple exit monitoring videos respectively in the deep learning-based visual multi-target tracking method according to an embodiment of the present application.

[0025] Figure 3 The present invention provides a flowchart of a method for visual multi-target tracking based on deep learning according to an embodiment of the present application, in which the multiple entry monitoring key frames are respectively subjected to entry target detection to obtain multiple entry target object area of ​​interest maps, and convolution encoding is performed to obtain the entry monitoring feature map.

[0026] Figure 4 Schematic diagram of a deep learning-based visual multi-target tracking system according to an embodiment of the present application.

[0027] Figure 5 Schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0029] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0030] In addition, numerous specific details are provided in the following detailed description to better illustrate the present application. Those skilled in the art will appreciate that the present application can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.

[0031] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0032] Figure 1 FIG is a flowchart of a method for visual multi-target tracking based on deep learning according to an embodiment of the present application. Figure 1 As shown, the visual multi-target tracking method based on deep learning according to the embodiment of the present application includes the steps of: S110, obtaining multiple entrance monitoring videos and multiple exit monitoring videos respectively collected by cameras deployed at the entrances and exits of a shopping mall; S120, extracting entrance monitoring feature maps and exit monitoring feature maps from the multiple entrance monitoring videos and the multiple exit monitoring videos respectively; S130, fusing the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map; S140, performing a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector; S150, performing reverse feature consistency enhancement based on the class regression domain on the multi-target monitoring feature vector to obtain an optimized multi-target monitoring feature vector; S160, passing the optimized multi-target monitoring feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether there are abnormal people in the shopping mall.

[0033] In step S110 of the present embodiment, multiple entrance and exit surveillance videos are captured by cameras deployed at the mall's entrances and exits. It should be understood that in order to comprehensively monitor the mall's entrances and exits, capturing multiple entrance and exit surveillance videos provides a more comprehensive perspective and more information, enabling more accurate tracking and monitoring of personnel activities within the mall. Furthermore, different entrances and exits may present different security risks and anomalies. By capturing separate surveillance videos, the security status of each area of ​​the mall can be better understood, allowing for timely detection and response to potential security issues. Therefore, capturing multiple entrance and exit surveillance videos helps improve the comprehensiveness and accuracy of the monitoring system. Deep learning-based object detection algorithms can automatically identify different targets in surveillance videos, such as people and vehicles. Object detection can extract target information in the entrance and exit areas. Deep learning-based object tracking algorithms can continuously track targets in surveillance videos, even if the targets appear obscured, change their posture, or move at varying speeds. Object tracking allows real-time acquisition of target movement trajectories and behavioral information within the mall. Mall entrances and exits typically have high traffic volumes, and multiple targets may be present simultaneously. Deep learning-based multi-target tracking algorithms can simultaneously process, differentiate, and track multiple targets. This allows for accurate counting of entrance and exit traffic, analysis of personnel flow and congestion, and more. Deep learning algorithms have made significant progress in target detection and tracking, achieving high real-time performance and accuracy. This makes deep learning-based visual multi-target tracking an effective tool for obtaining entrance and exit surveillance video within shopping mall surveillance systems. Specifically, cameras are installed at each entrance and exit of the mall. Ensure that the cameras are positioned appropriately to fully cover the entrance and exit areas and capture sufficient surveillance footage.

[0034] In step S120 of the embodiment of the present application, entrance monitoring feature maps and exit monitoring feature maps are extracted from the multiple entrance surveillance videos and the multiple exit surveillance videos, respectively. It should be understood that by extracting the entrance and exit monitoring feature maps, human behavior can be analyzed and modeled. Deep learning algorithms can learn normal human behavior patterns and identify abnormal behaviors, such as running, walking against traffic, and lingering. By comparing these with normal behavior patterns, it is possible to determine whether any abnormal individuals are present. The entrance and exit monitoring feature maps can be used to define the boundaries of a shopping mall. By detecting the presence of individuals in the boundary areas, it is possible to determine whether any individuals have entered or left the mall without authorization, thereby identifying potential abnormal situations. Extracting the entrance and exit monitoring feature maps enables real-time monitoring of the shopping mall. The deep learning algorithm can analyze the feature maps in real time, quickly detecting the presence of any abnormal individuals and taking timely appropriate measures to ensure safety and order in the shopping mall. Shopping malls typically have multiple entrances and exits, and observing from different angles can provide more comprehensive information. By extracting multiple entrance and exit monitoring feature maps separately, it is possible to simultaneously observe and analyze the presence of individuals in different areas, more accurately determining whether any abnormal individuals are present.

[0035] Specifically, in one embodiment of the present application, Figure 2 The figure shows a flow chart of extracting entrance monitoring feature graphs and exit monitoring feature graphs from the multiple entrance monitoring videos and the multiple exit monitoring videos in the deep learning-based visual multi-target tracking method according to an embodiment of the present application. Figure 2 As shown, in Figure 1 Based on the embodiment shown, the step S120 includes: S210, extracting multiple entry monitoring key frames and multiple exit monitoring key frames from the multiple entry monitoring videos and the multiple exit monitoring videos respectively; S220, subjecting the multiple entry monitoring key frames to entry target detection to obtain multiple entry target object area of ​​interest maps, and performing convolution coding to obtain the entry monitoring feature map; S230, subjecting the multiple exit monitoring key frames to exit target detection to obtain multiple exit target object area of ​​interest maps, and performing convolution coding to obtain the exit monitoring feature map.

[0036] Specifically, in a specific example of this application, step S210 extracts multiple entry monitoring keyframes and multiple exit monitoring keyframes from the multiple entry surveillance videos and the multiple exit surveillance videos, respectively. It should be understood that by extracting multiple entry monitoring keyframes and multiple exit monitoring keyframes, more visual information from different time points and perspectives can be obtained. These keyframes can capture the changes in different people, different behaviors, and different backgrounds, thereby providing a richer and more comprehensive visual representation. Entrances and exits are key locations for people entering and exiting a shopping mall, and people's behavior often changes at these locations. Extracting multiple entry monitoring keyframes and multiple exit monitoring keyframes can capture the dynamic behavior of people entering and leaving the mall, such as opening doors, passing through doors, and stopping, which helps to more accurately model and identify abnormal behavior. Mall entrances and exits are often busy and complex areas, potentially subject to occlusion, lighting changes, and noise. By extracting multiple entry monitoring keyframes and multiple exit monitoring keyframes, the robustness and reliability of the system can be increased, and information from multiple time points and perspectives can be used to reduce interference and misjudgment. Multiple entry monitoring keyframes and multiple exit monitoring keyframes can be used for multi-target tracking tasks. By performing target detection and tracking on key frames, the position and movement trajectory of people can be tracked in real time, and further analysis and judgment can be performed, such as calculating personnel flow and detecting abnormal behavior.

[0037] Furthermore, in a specific example of the present application, step S220 involves subjecting the multiple entry monitoring keyframes to entry target detection to obtain multiple entry target object region of interest maps, which are then subjected to convolutional coding to obtain the entry monitoring feature map. It should be understood that entry target detection can accurately locate and extract target objects within the entry area. By detecting target objects in entry keyframes, attention can be focused on areas related to the entry, eliminating interference from other irrelevant areas. This helps improve the accuracy and efficiency of subsequent analysis. The entry target object region of interest maps capture important information within the entry area. These region maps may include key entry and exit movements, occupant density, occupant flow, and other important features related to the entry. By extracting this key information, the state and changes of the entry area can be better understood. By subjecting the entry target object region of interest maps to convolutional coding, complex image data can be converted into a higher-level feature representation. Convolutional coding can extract features with semantic information by learning local patterns and structures in the image. These entry monitoring feature maps can better represent the characteristics of the entry area, providing more meaningful input for subsequent anomaly detection and behavior analysis. The combination of object detection and convolutional coding can reduce the amount of data required for processing. Entry point detection can narrow the region of interest and reduce the computational effort required for subsequent processing. Convolutional coding can also transform image data into low-dimensional feature representations, reducing storage and computational requirements and improving algorithm efficiency and real-time performance.

[0038] Figure 3 The figure shows a flow chart of the method for visual multi-target tracking based on deep learning according to an embodiment of the present application, wherein the multiple entrance monitoring key frames are respectively subjected to entrance target detection to obtain multiple entrance target object interest region maps, and convolution coding is performed to obtain the entrance monitoring feature map. Figure 3 As shown, in Figure 2 Based on the embodiment shown, the step S220 includes: S2201, respectively passing the multiple entry monitoring key frames through the entry target detection network based on the anchor-free window to obtain the multiple entry target object region of interest maps; S2202, arranging the multiple entry target object region of interest maps into an entry input tensor and then passing them through a three-dimensional convolutional neural network to obtain the entry monitoring feature map.

[0039] Specifically, step S2201 involves passing each of the multiple entry monitoring keyframes through an anchor-window-free entry target detection network to obtain region-of-interest maps for the multiple entry target objects. It should be understood that traditional target detection methods typically use anchor boxes to represent and match the location and scale of targets. However, in surveillance scenarios with multiple entry points, the scale and location of targets vary significantly, making it difficult to cover all possible scenarios using anchor boxes. An anchor-window-free object detection networks can adaptively detect and localize targets without the need for predefined anchor boxes, thus better adapting to targets of varying scales and locations. Multiple entry monitoring keyframes may contain multiple entry target objects, such as a person entering or leaving. By inputting each keyframe into the anchor-window-free entry target detection network, multiple entry target objects can be detected simultaneously and corresponding region-of-interest maps generated. An anchor-window-free object detection networks generally have high detection efficiency and accuracy. Through internal adaptive mechanisms and multi-scale feature fusion, they effectively capture target details and contextual information, thereby improving target detection accuracy. Furthermore, the anchor-window-free design reduces computational and memory consumption, improving the speed and efficiency of target detection. In entrance monitoring, in addition to normal entry and exit behavior, it is also necessary to detect abnormal targets, such as potential intruders or suspicious individuals. An anchor-free object detection network can identify abnormal targets that do not conform to normal targets by learning their characteristics and behavior patterns. This helps improve the security and reliability of entrance monitoring systems.

[0040] Accordingly, in a specific example of the present application, the multiple entry monitoring key frames are respectively passed through the entry target detection network based on the anchor-free window to obtain the multiple entry target object area of ​​interest maps, including: passing each entry monitoring key frame in the multiple entry monitoring key frames through multiple convolutional layers to obtain multiple shallow feature maps; passing the multiple shallow feature maps through the entry target detection network based on the anchor-free window to obtain the multiple entry target object area of ​​interest maps.

[0041] Furthermore, each of the multiple entry monitoring key frames is passed through multiple convolutional layers to obtain multiple shallow feature maps. It should be understood that multiple convolutional layers can achieve multi-scale feature extraction. Convolutional layers at different levels have different receptive fields. Shallower convolutional layers can capture more localized details, while deeper convolutional layers can capture more global semantic information. By extracting features at multiple levels, a more comprehensive and richer feature representation can be obtained, helping to more accurately describe the characteristics of the entry area. Multiple convolutional layers can also achieve fusion of features at different levels. Shallower features contain more detailed information, while deeper features contain more semantic information and contextual relationships. By fusing features from different levels, information from different levels can be comprehensively utilized, improving the expressiveness and discriminability of features. Multiple shallow feature maps can be used to visualize and analyze the characteristics of the entry area. Shallow feature maps have higher spatial resolution and can more clearly display the details and structure of the entry area. By visualizing and analyzing the feature maps, we can further understand the characteristics and changes of the entry area, providing more clues for subsequent behavior analysis and anomaly detection. By processing multiple convolutional layers, the amount of data that needs to be processed can be reduced. Deeper convolutional layers usually have lower resolution and higher number of channels, which can transform the input data into a more compact feature representation, reducing storage and computing requirements and improving the efficiency and real-time performance of the algorithm.

[0042] Specifically, each of the multiple entry monitoring key frames is passed through a multi-layer convolution layer to obtain a plurality of shallow feature maps, including: the multi-layer convolution layer includes N convolution layers, and N is greater than or equal to 1 and less than or equal to 6; the multiple layers of the multi-layer convolution layer perform convolution processing, pooling processing and nonlinear activation processing based on a two-dimensional convolution kernel on the input data in the forward pass of the layer to output the shallow feature map by the last convolution layer of the multi-layer convolution layer; the nonlinear activation function used in each layer of the multi-layer convolution layer is a Mish activation function.

[0043] Furthermore, the multiple shallow feature maps are separately passed through the anchor-window-free entry target detection network to obtain the multiple entry target object region-of-interest maps. It should be understood that the anchor-window-free object detection network can perform object detection and localization on the input feature maps, namely, determine the location and bounding box of the target. By separately inputting multiple shallow feature maps into the target detection network, the target's different scales and details can be captured from features at different levels, improving the accuracy and robustness of target detection. The anchor-window-free object detection network does not rely on predefined anchor boxes, but instead detects and localizes targets through an adaptive mechanism within the network. This allows the network to adapt to targets of varying scales, shapes, and poses, and achieves strong robustness. By separately inputting multiple shallow feature maps into the target detection network, the network can better adapt to the diversity and variability of entry target objects. By obtaining the region-of-interest maps of entry target objects, the target's location and bounding box can be visualized and analyzed. This helps to better understand target behavior and changes in entry monitoring scenarios and provides more information for subsequent behavior analysis and security monitoring.

[0044] Specifically, step S2202 arranges the multiple entry target object region of interest maps into an entry input tensor and then passes it through a three-dimensional convolutional neural network to obtain the entry monitoring feature map. It should be understood that in multiple entry monitoring scenarios, there may be multiple target objects, such as different people entering, exiting, or leaving. By arranging multiple entry target object region of interest maps as an input tensor, feature extraction for multiple targets can be processed simultaneously. A three-dimensional convolutional neural network can perform convolution operations on the input tensor in time and space to extract spatiotemporal features of multiple targets, including their appearance, shape, and motion information. Target behavior in entry monitoring scenarios is often sequential and dynamic. Using a three-dimensional convolutional neural network, spatiotemporal feature modeling can be performed on the input tensor to capture the target's motion trajectory and temporal changes. The three-dimensional convolution operation can learn features in both time and space, better modeling the spatiotemporal relationships of targets, and improving understanding and analysis of target behavior. By arranging multiple entry target object region of interest maps as an input tensor and performing feature extraction through a three-dimensional convolutional neural network, multiple target features can be fused. Convolutional neural networks have a hierarchical structure that can fuse feature representations of multiple targets at different levels, resulting in a richer and more comprehensive portal monitoring feature map. This helps improve the accuracy and robustness of target recognition, classification, and behavior analysis. The portal monitoring feature map can be visualized, analyzed, and subsequently processed. The feature map reflects the key characteristics and spatial distribution of targets, helping to understand target behavior and changes in portal monitoring scenarios. Furthermore, the feature map can serve as input for subsequent processing, such as behavior recognition and anomaly detection, to further extract and analyze high-level semantic information about the target.

[0045] Specifically, the multiple entry target object area of ​​interest maps are arranged as an entry input tensor and then passed through a three-dimensional convolutional neural network to obtain the entry monitoring feature map, which is used to: use each layer of the three-dimensional convolutional neural network to perform convolution processing, pooling processing and nonlinear activation processing based on a three-dimensional convolution kernel on the input data in the forward pass of the layer to output the entry monitoring feature map from the last layer of the three-dimensional convolution kernel.

[0046] Furthermore, in one embodiment of the present application, step S230 involves subjecting the multiple exit monitoring keyframes to exit target detection to obtain multiple exit target object region of interest maps, which are then convolutionally encoded to obtain the exit monitoring feature map. It should be understood that, using an anchor-free exit target detection network, target detection and localization can be performed on multiple exit monitoring keyframes, determining the location and bounding box of the exit target. This facilitates accurate identification of target objects in the exit monitoring scene and acquisition of their region of interest maps. After arranging the multiple exit target object region of interest maps into an exit input tensor, feature extraction is performed using a three-dimensional convolutional neural network. The three-dimensional convolutional neural network can perform convolution operations on the input tensor in both temporal and spatial dimensions to extract spatiotemporal features of the exit targets. This helps capture the appearance, shape, and motion information of the target, providing more comprehensive and rich exit monitoring features. By passing multiple exit monitoring keyframes through the exit target detection network to obtain exit target object region of interest maps, arranging these region of interest maps into an exit input tensor, and then applying them to the three-dimensional convolutional neural network to obtain the exit monitoring feature map, accurate exit target detection, multi-target feature extraction, and spatiotemporal feature modeling can be achieved. Such an approach can improve the ability to understand and analyze target behaviors in exit monitoring scenarios and provide richer and more accurate input data for subsequent behavior recognition and anomaly detection.

[0047] Specifically, in a specific example of the present application, the multiple exit monitoring key frames are respectively subjected to exit target detection to obtain multiple exit target object area of ​​interest maps, and convolution coding is performed to obtain the exit monitoring feature map, including: respectively passing the multiple exit monitoring key frames through an anchor-free window-based exit target detection network to obtain the multiple exit target object area of ​​interest maps; arranging the multiple exit target object area of ​​interest maps into exit input tensors and then passing them through a three-dimensional convolutional neural network to obtain the exit monitoring feature map.

[0048] In step S130 of the embodiment of the present application, the entrance monitoring feature map and the exit monitoring feature map are fused into a multi-target monitoring feature map. It should be understood that the entrance monitoring feature map and the exit monitoring feature map extract the feature information of the target from the perspective of the entrance and exit respectively. Fusion of them into a multi-target monitoring feature map can provide a more global target feature representation. By integrating the entrance and exit information, the overall behavior and path of the target can be better understood, and the movement and interaction pattern of the target in the entire monitoring area can be captured. The entrance and exit are important key points in the monitoring scene, and the behavior and movement path of the target between these two locations are correlated. After the entrance monitoring feature map and the exit monitoring feature map are fused, target association and trajectory analysis can be better performed. By analyzing the feature changes and movement patterns of the target between the entrance and exit, the target's entry and exit behavior can be identified, the residence time can be calculated, abnormal behavior can be detected, etc. The entrance monitoring feature map and the exit monitoring feature map both contain feature representations of multiple targets. Fusion of them into a multi-target monitoring feature map can achieve the fusion and integration of multiple target features. By fusing entry and exit features at different levels through methods such as convolutional neural networks, we can obtain richer and more accurate multi-target feature representations, improving target recognition, classification, and behavior analysis. By fusing entry and exit monitoring feature maps, we can obtain more comprehensive monitoring information for anomaly detection and early warning. By analyzing multi-target monitoring feature maps, we can detect abnormal behavior, unusual trajectories, and crowd congestion, issuing timely warnings and taking appropriate measures.

[0049] Specifically, in one embodiment of the present application, the entrance monitoring feature map and the exit monitoring feature map are fused into a multi-target monitoring feature map, and the multi-target monitoring feature map is obtained by fusing the entrance monitoring feature map and the exit monitoring feature map using the following formula, wherein the fusion formula is: ;in, is the multi-target monitoring feature map, is the entrance monitoring feature map, is the export monitoring characteristic diagram, represents the addition of the elements at corresponding positions of the entrance monitoring feature map and the exit monitoring feature map, and is a weighting parameter used to control the balance between the entry monitoring feature map and the exit monitoring feature map in the multi-target monitoring feature map.

[0050] In step S140 of the embodiment of the present application, the multi-target monitoring feature map is pooled to obtain a multi-target monitoring feature vector. It should be understood that the multi-target monitoring feature map is high-dimensional, and feature processing thereof may require a large amount of storage space and has a high computational complexity. By reducing the size of the feature map, the computational complexity and memory requirements of subsequent network layers can be significantly reduced. Therefore, the multi-target monitoring feature map is further subjected to a mean pooling operation based on a feature matrix to obtain a multi-target monitoring feature vector.

[0051] In particular, the technical solution of this application takes into account that the quality of the data collected from entrance and exit surveillance videos may vary due to different camera positions, angles, resolutions, or lighting conditions. The facing target detection network and the back-facing target detection network based on anchor-free windows may capture human features in different directions, and these features may be inconsistent when fused. The key frames extracted from the surveillance video may not fully represent the dynamic changes and complexity of the entire video sequence. The behavioral patterns of people in a shopping mall can be very complex and changeable. A single monitoring feature vector may not be able to capture all possible behavioral features, resulting in poor manifold geometric consistency of the overall feature distribution of the multi-target monitoring feature vector. When the distribution of multi-target monitoring feature vectors in the feature space is inconsistent, the classifier may have difficulty accurately learning the boundary between normal and abnormal people, especially in areas with sparse distribution in the feature space. This may cause the classifier to have biased predictions in these areas, affecting the accuracy of the classification results. If the distribution of multi-target monitoring feature vectors in the feature space is inconsistent, the classifier may perform well on the training set, but its generalization ability will be reduced when faced with new, unseen abnormal behavior patterns. In order to improve the accuracy of abnormal personnel detection in shopping malls, the multi-target monitoring feature vector is enhanced based on the reverse feature consistency of the class regression domain to obtain the optimized multi-target monitoring feature vector.

[0052] Specifically, in step S150 of the embodiment of the present application, the multi-target monitoring feature vector is subjected to reverse feature consistency enhancement based on the class regression domain to obtain an optimized multi-target monitoring feature vector, including: multiplying the multi-target monitoring feature vector with the classification weight matrix of the geological state classifier to obtain a first intermediate multi-target monitoring feature vector; using the concat function to process the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector to obtain a second intermediate multi-target monitoring feature vector; adding the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector by position to obtain a third intermediate multi-target monitoring feature vector; multiplying the first weight matrix by the second intermediate multi-target monitoring feature vector and adding the first bias vector and then passing it through the sigmoid function to obtain a first activation value ; Multiply the second weight matrix by the third intermediate multi-target monitoring feature vector and add the second bias vector, and then pass it through the sigmoid function to obtain a second activation value; calculate the mean of the first activation value and the second activation value to obtain the activation mean; calculate the difference between one and the activation mean, and use the difference as the first weighting coefficient and the activation mean as the second weighting coefficient, weight the second intermediate multi-target monitoring feature vector and the multi-target monitoring feature vector by position to obtain a fourth intermediate multi-target monitoring feature vector; perform an exponential operation with a natural constant as the base on each eigenvalue of the backward anchor reference vector to obtain a fifth intermediate multi-target monitoring feature vector; subtract the fifth intermediate multi-target monitoring feature vector from the fourth intermediate multi-target monitoring feature vector by position, and then pass it through the ReLU function to obtain an optimized multi-target monitoring feature vector.

[0053] In addition, in step S150 of the embodiment of the present application, the multi-target monitoring feature vector is subjected to reverse feature consistency enhancement based on the class regression domain to obtain an optimized multi-target monitoring feature vector, further comprising: performing reverse feature consistency enhancement based on the class regression domain on the multi-target monitoring feature vector using the following formula, wherein the formula is: ;in, represents the multi-target monitoring feature vector, represents the classification weight matrix of the geological state classifier, represents matrix multiplication, Represents a cascade function, is the backward anchor reference vector, represents the first weight matrix, represents the first bias vector, represents the first activation value, represents the second weight matrix, represents the second bias vector, represents the second activation value, represents the logistic function, Indicates subtraction by position, represents the linear rectification function, represents the natural exponential function, Represents the optimized multi-target monitoring feature vector.

[0054] That is, in the technical solution of the present application, the manifold geometric consistency of the overall feature distribution of the multi-target monitoring feature vector is poor, resulting in a long-range distribution regression deviation across the geological state classifier when it is classified by the classifier, affecting the accuracy of the classification result. Therefore, in the technical solution of the present application, the multi-target monitoring feature vector is subjected to reverse feature consistency enhancement based on the class regression domain, which uses the classification weight matrix of the geological state classifier to perform an auxiliary description of the attribute distribution of the multi-target monitoring feature vector to support the descriptiveness of the improved multi-target monitoring feature vector for the different distance feature descriptions of the classification weight matrix of the geological state classifier to the category probability of the preset classification, wherein the backward anchor reference vector is used as an offset and activated by an activation operation to maintain the reinforcement of the distribution description dependency with a positive effect, so that the manifold geometric consistency of the overall attribute distribution of the multi-target monitoring feature vector can be significantly improved, thereby reducing the class probability distribution regression deviation of the multi-target monitoring feature vector when passing through the classifier, so as to improve the accuracy of the classification result.

[0055] In step S160 of the embodiment of the present application, the optimized multi-target monitoring feature vector is passed through a classifier to obtain a classification result. This classification result is used to indicate whether the mall contains any unusual individuals. It will be appreciated that mall surveillance systems typically require real-time monitoring and identification of unusual individuals, such as potential threats, suspicious behavior, or illegal activities. By inputting the multi-target monitoring feature map into the classifier, targets in the monitoring scene can be classified and determined to be unusual individuals. The classifier can learn and identify features that deviate from normal behavior and appearance patterns, thereby assisting in identifying unusual individuals. The multi-target monitoring feature map incorporates entry and exit monitoring features, providing a more comprehensive and accurate representation of target behavior. By inputting these features into the classifier, the target's appearance, shape, motion, and other characteristics can be comprehensively considered, thereby improving the accuracy of unusual individual identification. The classifier can learn and model the characteristic distribution of normal behavior, classify targets that deviate from this distribution, and then determine whether an unusual individual is present. By inputting the multi-target monitoring feature map into the classifier, real-time monitoring and early warning of the monitoring scene can be achieved. The classifier can quickly analyze the target's features and provide a classification result. If the classification results indicate the presence of abnormal personnel, the early warning mechanism can be triggered in time to notify relevant personnel to take necessary measures to ensure the safety and order of the mall.

[0056] Specifically, in a specific example of the present application, the optimized multi-target monitoring feature vector is passed through a classifier to obtain a classification result, and the classification result is used to indicate whether there are abnormal personnel in the shopping mall, and is used to: use the classifier to process the optimized multi-target monitoring feature vector using the following formula to obtain the classification result; wherein, the formula is: ;in, arrive is the weight matrix, arrive is the bias vector, To optimize the multi-target monitoring feature vector, represents the normalized exponential function, Indicates the classification result.

[0057] In summary, the deep learning-based visual multi-target tracking method, system, and electronic device described in the embodiments of this application collect surveillance video by deploying cameras at the entrance and exit of a shopping mall, extracting entrance and exit monitoring feature maps from them, then fusing these feature maps into multi-target monitoring feature maps, and classifying them using a classifier to determine whether there are any abnormal people in the mall. Leveraging the advantages of deep learning, by automatically learning the appearance and motion characteristics of the target, high-quality robust tracking is achieved, improving the intelligence level of the monitoring system.

[0058] Figure 4 FIG. 1 is a block diagram of a deep learning-based visual multi-target tracking system according to an embodiment of the present application. Figure 4 As shown, the deep learning-based visual multi-target tracking system 100 according to the embodiment of the present application includes: a camera acquisition module 110, used to obtain multiple entrance surveillance videos and multiple exit surveillance videos respectively collected by cameras deployed at the entrances and exits of a shopping mall; a feature map extraction module 120, used to extract entrance monitoring feature maps and exit monitoring feature maps from the multiple entrance surveillance videos and the multiple exit surveillance videos respectively; a multi-target fusion module 130, used to fuse the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map; a feature map pooling module 140, used to perform a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector; a multi-target feature optimization module 150, used to perform reverse feature consistency enhancement based on the class regression domain on the multi-target monitoring feature vector to obtain an optimized multi-target monitoring feature vector; a multi-target result analysis module 160, used to pass the optimized multi-target monitoring feature vector through a classifier to obtain a classification result, and the classification result is used to indicate whether there are abnormal people in the shopping mall.

[0059] Here, those skilled in the art will appreciate that the specific functions and operations of the various units and modules in the above-mentioned deep learning-based visual multi-target tracking system have been described in detail above. Figures 1 to 3 It has been introduced in detail in the description of the deep learning based visual multi-target tracking method, and therefore, its repeated description will be omitted.

[0060] As described above, the deep learning-based visual multi-target tracking system 100 according to the embodiment of the present application can be implemented in various terminal devices, such as a server deployed with a deep learning-based visual multi-target tracking algorithm. In one example, the deep learning-based visual multi-target tracking system 100 can be integrated into the terminal device as a software module and / or a hardware module. For example, the deep learning-based visual multi-target tracking system 100 can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the deep learning-based visual multi-target tracking system 100 can also be one of the many hardware modules of the terminal device.

[0061] Alternatively, in another example, the deep learning-based visual multi-target tracking system 100 and the terminal device may also be separate devices, and the deep learning-based visual multi-target tracking system 100 may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0062] Specifically, this application also provides another embodiment, below, refer to Figure 5 To describe the electronic device according to the embodiment of the present application.

[0063] Figure 5 FIG2 shows a block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device 50 according to an embodiment of the present disclosure includes a memory 501 and a processor 502. The components in the electronic device 50 are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0064] The memory 501 is used to store computer-readable instructions. Specifically, the memory 501 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.

[0065] The processor 502 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 50 to perform desired functions. In one embodiment of the present disclosure, the processor 502 is configured to execute the computer-readable instructions stored in the memory 501, so that the electronic device 50 executes the reference Figure 1 and Figure 4 Describe the deep learning-based visual multi-target tracking method or reference Figure 4 Described deep learning based visual multi-object tracking system.

[0066] Furthermore, it is important to understand that Figure 5 The components and structures of the electronic device 50 shown are only exemplary and non-restrictive. The electronic device 50 may also have other components and structures as needed. For example, a surveillance video acquisition device and an output device (not shown). The surveillance video acquisition device can be used to acquire multiple entrance surveillance videos and multiple exit surveillance videos respectively acquired by cameras deployed at the entrances and exits of a shopping mall, and store the acquired surveillance videos in a memory 501 for use by other components. Of course, other surveillance video acquisition devices can also be used to acquire the multiple entrance surveillance videos and multiple exit surveillance videos respectively acquired by cameras deployed at the entrances and exits of a shopping mall, and send the acquired surveillance videos to the electronic device 50, which can store the received surveillance videos in the memory 501. The output device can output various information to the outside (e.g., a user), such as whether there are any abnormal people in the shopping mall. The output device may include one or more of a display, a speaker, a projector, a network card, etc.

[0067] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0068] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.

[0069] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.

[0070] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in hardware, or, as described above, those skilled in the art will readily appreciate that the present invention may be implemented in hardware, firmware, or a combination thereof. When implemented using software, the aforementioned functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media includes any medium that facilitates the transfer of computer programs from one location to another. Storage media may be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection may appropriately constitute a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of protection of computer-readable media.

[0071] In short, the above description is only a preferred embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A visual multi-target tracking method based on deep learning, characterized in that: include: Obtain multiple entrance surveillance videos and multiple exit surveillance videos collected by cameras deployed at the entrances and exits of the shopping mall; Extracting an entrance monitoring feature graph and an exit monitoring feature graph from the plurality of entrance monitoring videos and the plurality of exit monitoring videos respectively; Merging the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map; Performing a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector; Performing reverse feature consistency enhancement based on the class regression domain on the multi-target monitoring feature vector to obtain an optimized multi-target monitoring feature vector; Passing the optimized multi-target monitoring feature vector through a classifier to obtain a classification result, wherein the classification result is used to indicate whether there are abnormal persons in the shopping mall; Among them, the multi-target monitoring feature vector is enhanced by reverse feature consistency based on the class regression domain to obtain an optimized multi-target monitoring feature vector, including: Multiplying the multi-objective monitoring feature vector with a classification weight matrix of a geological state classifier to obtain a first intermediate multi-objective monitoring feature vector; Using a concat function to process the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector to obtain a second intermediate multi-target monitoring feature vector; Adding the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector by position to obtain a third intermediate multi-target monitoring feature vector; The first weight matrix is ​​multiplied by the second intermediate multi-target monitoring feature vector and then added to the first bias vector and then passed through a sigmoid function to obtain a first activation value; The second weight matrix is ​​multiplied by the third intermediate multi-target monitoring feature vector and then added to the second bias vector and then passed through the sigmoid function to obtain a second activation value; Calculating a mean of the first activation value and the second activation value to obtain an activation mean; Calculate a difference value minus the activation mean, and use the difference value as a first weighting coefficient and the activation mean value as a second weighting coefficient to weight the first intermediate multi-target monitoring feature vector and the multi-target monitoring feature vector by position to obtain a fourth intermediate multi-target monitoring feature vector; Performing an exponential operation with a natural constant as the base on each eigenvalue of the backward anchor reference vector to obtain a fifth intermediate multi-target monitoring eigenvector; The fourth intermediate multi-target monitoring feature vector is subtracted from the fifth intermediate multi-target monitoring feature vector by position and then passed through the ReLU function to obtain an optimized multi-target monitoring feature vector.

2. The method for visual multi-target tracking based on deep learning according to claim 1, characterized in that: Extracting an entrance monitoring feature graph and an exit monitoring feature graph from the plurality of entrance monitoring videos and the plurality of exit monitoring videos respectively includes: Extracting a plurality of entry monitoring key frames and a plurality of exit monitoring key frames from the plurality of entry monitoring videos and the plurality of exit monitoring videos respectively; The plurality of entrance monitoring key frames are respectively subjected to entrance target detection to obtain a plurality of entrance target object region of interest maps, and are subjected to convolution encoding to obtain the entrance monitoring feature map; The multiple exit monitoring key frames are respectively subjected to exit target detection to obtain multiple exit target object region of interest maps, and are then subjected to convolution coding to obtain the exit monitoring feature map.

3. The method for visual multi-target tracking based on deep learning according to claim 2, characterized in that: The multiple entrance monitoring key frames are respectively subjected to entrance target detection to obtain multiple entrance target object interest area maps, and the entrance monitoring feature map is obtained by convolution coding, including: Passing the multiple entrance monitoring key frames through the entrance target detection network based on anchor-free window respectively to obtain the multiple entrance target object region of interest maps; The plurality of entrance target object region of interest maps are arranged as an entrance input tensor and then passed through a three-dimensional convolutional neural network to obtain the entrance monitoring feature map.

4. The method for visual multi-target tracking based on deep learning according to claim 3, characterized in that: The multiple entrance monitoring key frames are respectively passed through an entrance target detection network based on a non-anchor window to obtain the multiple entrance target object region of interest maps, including: Passing each of the plurality of entry monitoring key frames through multiple convolutional layers to obtain a plurality of shallow feature maps; The multiple shallow feature maps are respectively passed through the anchor-free window-based entry target detection network to obtain the multiple entry target object region of interest maps.

5. The method for visual multi-target tracking based on deep learning according to claim 4, characterized in that: Each of the plurality of entry monitoring key frames is passed through multiple convolutional layers to obtain a plurality of shallow feature maps, including: The multi-layer convolutional layer includes N convolutional layers, where N is greater than or equal to 1 and less than or equal to 6; The multiple layers of the multi-layer convolution layer respectively perform convolution processing based on a two-dimensional convolution kernel, pooling processing and non-linear activation processing on the input data in the forward transmission of the layer, so that the last convolution layer of the multi-layer convolution layer outputs the shallow feature map; The nonlinear activation function used in each layer of the multi-layer convolutional layer is the Mish activation function.

6. The method for visual multi-target tracking based on deep learning according to claim 5, characterized in that: The multiple entrance target object area of ​​interest maps are arranged as an entrance input tensor and then passed through a three-dimensional convolutional neural network to obtain the entrance monitoring feature map, including: using each layer of the three-dimensional convolutional neural network to perform convolution processing, pooling processing and non-linear activation processing based on a three-dimensional convolution kernel on the input data in the forward pass of the layer to output the entrance monitoring feature map from the last layer of the three-dimensional convolution kernel.

7. The deep learning-based visual multi-target tracking method according to claim 6, characterized in that: The multiple exit monitoring key frames are respectively subjected to exit target detection to obtain multiple exit target object region of interest maps, and are subjected to convolution coding to obtain the exit monitoring feature map, including: Passing the multiple exit monitoring key frames through an exit target detection network based on no anchor window respectively to obtain the multiple exit target object region of interest maps; The multiple exit target object region of interest maps are arranged as an exit input tensor and then passed through a three-dimensional convolutional neural network to obtain the exit monitoring feature map.

8. A visual multi-target tracking system based on deep learning, characterized in that: include: The camera acquisition module is used to obtain multiple entrance surveillance videos and multiple exit surveillance videos respectively collected by cameras deployed at the entrances and exits of the shopping mall; A feature map extraction module, used to extract an entry monitoring feature map and an exit monitoring feature map from the multiple entry monitoring videos and the multiple exit monitoring videos respectively; A fusion multi-target module, used for fusing the entrance monitoring feature map and the exit monitoring feature map into a multi-target monitoring feature map; A feature map pooling module, used for performing a pooling operation on the multi-target monitoring feature map to obtain a multi-target monitoring feature vector; A multi-objective feature optimization module, used for performing reverse feature consistency enhancement based on a class regression domain on the multi-objective monitoring feature vector to obtain an optimized multi-objective monitoring feature vector; A multi-objective result analysis module, used for passing the optimized multi-objective monitoring feature vector through a classifier to obtain a classification result, wherein the classification result is used to indicate whether there are abnormal persons in the shopping mall; Among them, the multi-target monitoring feature vector is enhanced by reverse feature consistency based on the class regression domain to obtain an optimized multi-target monitoring feature vector, including: Multiplying the multi-objective monitoring feature vector with a classification weight matrix of a geological state classifier to obtain a first intermediate multi-objective monitoring feature vector; Using a concat function to process the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector to obtain a second intermediate multi-target monitoring feature vector; Adding the multi-target monitoring feature vector and the first intermediate multi-target monitoring feature vector by position to obtain a third intermediate multi-target monitoring feature vector; The first weight matrix is ​​multiplied by the second intermediate multi-target monitoring feature vector and then added to the first bias vector and then passed through a sigmoid function to obtain a first activation value; The second weight matrix is ​​multiplied by the third intermediate multi-target monitoring feature vector and then added to the second bias vector and then passed through the sigmoid function to obtain a second activation value; Calculating a mean of the first activation value and the second activation value to obtain an activation mean; Calculate a difference value minus the activation mean, and use the difference value as a first weighting coefficient and the activation mean value as a second weighting coefficient to weight the first intermediate multi-target monitoring feature vector and the multi-target monitoring feature vector by position to obtain a fourth intermediate multi-target monitoring feature vector; Performing an exponential operation with a natural constant as the base on each eigenvalue of the backward anchor reference vector to obtain a fifth intermediate multi-target monitoring eigenvector; The fourth intermediate multi-target monitoring feature vector is subtracted from the fifth intermediate multi-target monitoring feature vector by position and then passed through the ReLU function to obtain an optimized multi-target monitoring feature vector.

9. An electronic device, comprising: processor; A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the deep learning-based visual multi-target tracking method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Divulging behavior identification method based on deep learning

    CN115861877A

  • Fish behavior detection method based on machine vision

    CN117975572A