Target tracking method and device based on dynamic data stream, electronic equipment and medium

By adopting a dynamic data flow-based target tracking method in the substation operation scenario, using RGB images of multiple terminal devices for feature extraction and merging, and combining Kalman filtering algorithm for prediction and update, the problems of limited view of single cameras and multi-camera data processing computing resources are solved, and efficient and flexible target tracking is achieved.

CN119941783APending Publication Date: 2025-05-06SHENZHEN GUOCHUANG EMBODIED INTELLIGENT ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411775362.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the substation operation scenario, the prior art causes the target trajectory to be lost and fragmented due to the limited view of a single camera in substations, and the computing resources for multi-camera data processing are large, and the tracking accuracy depends on accurate prior parameters.

Method used

The target tracking method based on dynamic data flow is adopted, and the RGB pictures of multiple terminal devices are obtained, feature extraction and projection into BEV features, and probability occupancy maps are generated after merging and decoding. Combining the Kalman filtering algorithm for prediction and update, to achieve flexible tracking of the target.

Benefits of technology

It improves the flexibility of automatic target tracking, adapts to dynamically changing data flow input, reduces the use of computing resources, reduces the dependence on prior parameters, and improves tracking accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941783A_ABST
    Figure CN119941783A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a target tracking method and device based on a dynamic data stream, electronic equipment and a medium, and the method comprises the steps: obtaining S RGB pictures of a target object collected by S terminal equipment at the current moment, carrying out the feature extraction of each RGB picture, obtaining a corresponding picture feature, projecting the picture feature into a BEV feature, and carrying out the feature extraction of the BEV feature; and combining the S BEV features into a first BEV feature map, decoding the first BEV feature map to obtain a second BEV feature map, processing the second BEV feature map to obtain a probability occupancy map, and obtaining an initial tracking result of the target object based on the second BEV feature map and the probability occupancy map. The method and the device can adapt to multi-target detection tracking of dynamic change effective data stream input, and the flexibility of target automatic tracking is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a target tracking method, device, electronic device and medium based on dynamic data stream. Background Art

[0002] Substations are important hubs that connect power systems and power users, and their operational safety is directly related to the stable operation of the power system. Substation operation scenarios have the characteristics of large working scope, long working hours, and high risk factors. Although existing technologies can achieve visual monitoring of the operation site by deploying a large number of video terminals, the monitoring range of a single video terminal device is limited, the continuous tracking capability is insufficient, and the massive generation of multi-view data requires a large workload for manual screening and analysis, which makes it difficult to fully cover safety supervision and inefficient.

[0003] In order to overcome the problem of trajectory loss and trajectory fragmentation caused by the limited field of view of a single camera, it is necessary to establish a monitoring network of multiple cameras and jointly track the target through camera information from multiple perspectives. Although the existing front-fusion end-to-end multi-camera multi-target tracking algorithm can solve the above problems to a certain extent, it needs to process the data streams of all camera perspectives at the same time, which takes up a lot of computing resources and relies on accurate prior parameters. When the parameters are inaccurate, the tracking accuracy drops significantly. Summary of the invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides a target tracking method, device, electronic device and medium based on dynamic data stream, which can adapt to multi-target detection and tracking of dynamically changing valid data stream input and improve the flexibility of automatic target tracking.

[0005] A first aspect of the present application provides a target tracking method based on dynamic data stream, the method comprising: Obtain an RGB image of the target object collected by each of the S terminal devices at the current moment, and obtain S RGB images, where S is an integer greater than 1; Perform feature extraction on a target RGB image to obtain image features, wherein the target RGB image is any one of the S RGB images; Projecting the image features into BEV features; Merging the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map; Decoding the first BEV characteristic graph to obtain a second BEV characteristic graph; Processing the second BEV characteristic map to obtain a probability occupancy map; An initial tracking result of the target object is obtained based on the second BEV feature map and the probability occupancy map.

[0006] In an optional embodiment, the method further comprises: Predicting a first prediction center of the target object at a next moment based on a Kalman filter algorithm, wherein the next moment and the current moment are two adjacent moments; The initial tracking result is updated based on the first prediction center to obtain a target tracking result corresponding to the target object.

[0007] The updating of the initial tracking result based on the first prediction center to obtain the target tracking result corresponding to the target object includes: Determine a first detection result in which the occupancy probability value in the probabilistic occupancy map is greater than a first preset threshold; Extracting a first Reid feature of a first detection center of the first detection result; Calculating a first Mahalanobis distance between the first prediction center and the first detection center; Calculate a first cosine distance between the first prediction center and the first detection center based on the first Reid feature; Matching the first predicted center with the initial tracking result based on the first Mahalanobis distance and the first cosine distance to obtain a first matching result; The initial tracking result is updated based on the first matching result to obtain the target tracking result.

[0008] In an optional implementation, before updating the initial tracking result based on the first matching result, the method further includes: Determine a second detection result in which the occupancy probability value in the probabilistic occupancy map is greater than a second preset threshold; Extracting a second ReID feature of a second detection center of the second detection result; Calculate a second Mahalanobis distance between the first prediction center and the second detection center; Calculate a second cosine distance between the first prediction center and the second detection center based on the second Reid feature; Matching the first predicted center with the initial trajectory based on the second Mahalanobis distance and the second cosine distance to obtain a second matching result; The initial tracking result is updated based on the first matching result and the second matching result to obtain the target tracking result.

[0009] In an optional implementation, extracting features from the target RGB image to obtain image features includes: Input the target RGB image into a first convolution block of a ResNet-18 architecture, so that the first convolution block downsamples the target RGB image to obtain a first feature map; Inputting the target RGB image into the second convolution block of the ResNet-18 architecture, so that the second convolution block downsamples the target RGB image to obtain a second feature map; Inputting the target RGB image into the third convolution block of the ResNet-18 architecture, so that the third convolution block downsamples the target RGB image to obtain a third feature map; Performing upsampling processing on the first feature map, the second feature map, and the third feature map; The first feature map, the second feature map, and the third feature map after upsampling are concatenated to obtain the image features.

[0010] In an optional implementation, decoding the first BEV characteristic graph to obtain the second BEV characteristic graph includes: Inputting the first BEV feature map into the first convolutional layer, so that the first convolutional layer performs downsampling processing on the first BEV feature map to obtain a first downsampling feature; Inputting the first BEV feature map into the second convolutional layer, so that the second convolutional layer performs downsampling processing on the first BEV feature map to obtain a second downsampling feature; inputting the first BEV feature map into the third convolutional layer, so that the third convolutional layer performs downsampling processing on the first BEV feature map to obtain a third downsampling feature; Performing upsampling processing on the first down-sampling feature, the second down-sampling feature and the third down-sampling feature; The first down-sampling feature, the second down-sampling feature and the third down-sampling feature after the up-sampling process are connected in the channel dimension to obtain the second BEV feature map.

[0011] In an optional implementation, before extracting features from the target RGB image, the method further includes: Performing centering and normalization processing on the target RGB image; Image enhancement is performed on the target RGB image after the centering and normalization processing is completed.

[0012] A second aspect of the present application provides a target tracking device based on dynamic data stream, the device comprising: An image acquisition module is used to acquire an RGB image of the target object collected by each of the S terminal devices at the current moment, and obtain S RGB images, where S is an integer greater than 1; A feature extraction module is used to extract features of a target RGB image to obtain image features, wherein the target RGB image is any one of the S RGB images; A feature projection module, used for projecting the image features into BEV features; A feature merging module, used for merging the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map; A feature decoding module, used for decoding the first BEV feature map to obtain a second BEV feature map; A feature processing module, used for processing the second BEV feature map to obtain a probability occupancy map; An identification and tracking module is used to obtain an initial tracking result of the target object based on the second BEV feature map and the probability occupancy map.

[0013] A third aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the target tracking method based on dynamic data stream when executing the computer program.

[0014] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned target tracking method based on dynamic data stream.

[0015] In summary, the target tracking method, device, electronic device and medium based on dynamic data stream provided by the present application collects RGB images of the target object in real time through S terminal devices to obtain S RGB images. Where S is an integer greater than 1, indicating that multiple data sources provide data at the same time, and the data is collected in real time to generate multiple dynamic data streams that are constantly changing. Feature extraction is performed on each RGB image to obtain image features, and the image features are projected as BEV features, which helps to better understand and process targets on the ground. Then, merging the BEV features of multiple dynamic data sources can further improve the redundancy and complementarity of information, which helps to track targets more accurately in a dynamically changing environment. Then, the first BEV feature map is decoded to obtain a second BEV feature map, and the second BEV feature map is processed to obtain a probability occupancy map, which can dynamically reflect the position and state changes of the target in space. Finally, based on the second BEV feature map and the probability occupancy map, the initial tracking result of the target object is obtained, which can accurately reflect the position and state of the target in a dynamic environment, thereby effectively adapting to the dynamically changing data stream input and improving the flexibility of multi-target detection and tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a target tracking method based on dynamic data stream shown in an embodiment of the present application; Figure 2 is another flow chart of a target tracking method based on dynamic data stream shown in an embodiment of the present application; Figure 3 This is a schematic diagram of a scene map divided into grids as shown in an embodiment of the present application; Figure 4 It is a functional module diagram of a target tracking device based on dynamic data stream shown in an embodiment of the present application; Figure 5 It is a structural schematic diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0018] The following will clearly and completely describe the concept, specific structure and technical effects of the present invention in combination with the embodiments and drawings, so as to fully understand the purpose, characteristics and effects of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by technicians in this field without creative work are all within the scope of protection of the present invention. In addition, all the connection / connection relationships involved in the patent do not refer to the direct connection of components, but refer to the formation of a better connection structure by adding or reducing connection accessories according to the specific implementation situation. The various technical features in the invention can be combined interchangeably without conflicting with each other.

[0019] Reference Figure 1 As shown, it is a flow chart of a target tracking method based on dynamic data stream shown in an embodiment of the present application, and the target tracking method based on dynamic data stream includes the following steps.

[0020] S11, obtaining an RGB image of the target object collected by each of the S terminal devices at the current moment, and obtaining S RGB images.

[0021] Wherein, S is an integer greater than 1, and the terminal device refers to a collection device for collecting RGB images including the target of interest (i.e., the target object), such as a camera, etc. The RGB image contains the temporal and spatial information of the target object, and the image size of the RGB image is , is the height of the RGB image, is the width of the RGB image. The target object refers to the subject that needs to be identified and tracked, which may include, but is not limited to: pedestrians, vehicles, animals, etc.

[0022] In the fields of video surveillance, autonomous driving, and intelligent security, pedestrians are important target objects. By identifying and tracking pedestrians, functions such as crowd monitoring, behavior analysis, and anomaly detection can be achieved. In order to more clearly understand the inventive concept of this application, the target object in the embodiments of this application is described using pedestrians as an example.

[0023] In the embodiment of the present application, the substation operation scene usually has a wide working range and is fully monitored by deploying a large number of cameras. However, only the camera field of view that includes pedestrians can provide effective tracking information. Therefore, the electronic device can ensure that only the camera data stream containing pedestrians is transmitted to the end-to-end network of the electronic device by dynamically enabling and disabling the camera data stream strategy. Specifically, the electronic device can first calibrate the substation operation scene, that is, divide the scene map corresponding to the substation operation scene into grids, and refer to Figure 3As shown. The corresponding monitoring range is calculated by the installation position and field of view (FOV) of the terminal device, and the terminal device number that can be covered in each square is recorded. When the target is currently located in square A, the terminal device numbers that can be covered in the adjacent squares adjacent to square A are determined. Combined with the terminal device numbers that can be covered in the current square, the camera data stream of the terminal device that needs to be activated to monitor the target within a certain period of time is obtained, and the collected RGB images are transmitted to the electronic device for subsequent target tracking processing. Other irrelevant data streams are closed and not input into the electronic device for target tracking processing.

[0024] When monitoring the situation of the target within a certain time period requires the activation of camera data streams of S terminal devices, the electronic device only processes the camera data stream, obtains the RGB picture of each pedestrian at the current moment in the camera data stream, and obtains S RGB images.

[0025] Through the above optional implementation, by dynamically adjusting the camera data stream, not only the amount of algorithm calculation is significantly reduced, but also the overall efficiency of the system is improved. In addition, the monitoring system can focus on the target object more accurately, achieve efficient real-time tracking, and reduce resource consumption, providing strong support for the safety management of the substation.

[0026] S12, extracting features from the target RGB image to obtain image features.

[0027] The target RGB picture is any one of the S RGB pictures.

[0028] Combined Figure 2 In some embodiments, when S RGB images including pedestrians at the current moment collected by S terminal devices are obtained, the electronic device can encode the S RGB images using the ResNet-18 architecture. In the encoding stage, three convolution blocks of the ResNet-18 architecture, namely the first convolution block, the second convolution block and the third convolution block, are used to extract features from the S input RGB images. Each convolution block gradually downsamples the image size of each RGB image to half of the original size to obtain three feature maps of different sizes. Among them, the feature map obtained by the downsampling process of the first convolution block is called the first feature map, the feature map obtained by the downsampling process of the second convolution block is called the second feature map, and the feature map obtained by the downsampling process of the third convolution block is called the third feature map. Then, the first feature map, the second feature map and the third feature map output by the three convolution blocks are upsampled in turn, and the upsampled first feature map, the second feature map and the third feature map are spliced ​​to finally generate a feature map with the target number of output channels and the target size, which is called the picture feature. Among them, the target output channel number is , the target size is ,in , Therefore, by performing feature extraction on S RGB images, the image features corresponding to the S RGB images can be obtained.

[0029] Through the above optional implementation, the ResNet-18 architecture is used as an encoder to encode the RGB image, so that the complex information in the original RGB image can be compressed into a recognizable feature representation, which improves the efficiency and accuracy of feature extraction and provides a solid foundation for subsequent processing steps.

[0030] In an optional implementation, before extracting features from the target RGB image, the method further includes: Performing centering and normalization processing on the target RGB image; Image enhancement is performed on the target RGB image after the centering and normalization processing is completed.

[0031] In some embodiments, after acquiring S RGB images, the electronic device may pre-process each RGB image, such as centering, normalization, and image enhancement, etc. Image enhancement refers to rotating, scaling, and translating the RGB image to obtain an adjusted RGB image, and performing random noise addition and Gaussian filtering on the adjusted RGB image to achieve data amplification of the adjusted RGB image.

[0032] S13, projecting the image features into BEV features.

[0033] Combined Figure 2 In some embodiments, after encoding the RGB image to obtain the image features, the electronic device may project the image features into Bird's Eye View (BEV) features through a projection matrix.

[0034] Specifically, the RGB image is transformed from the two-dimensional image pixel coordinates into Convert to three-dimensional space : ; Among them, s is the scaling factor, P is the perspective transformation matrix, , is the internal parameter matrix of the terminal device, is the external parameter matrix of the terminal device. For BEV features, , so the projection process can be simplified as: ; in, is the projection matrix, through the projection matrix Will The image features are projected to a predefined size The size of the ground plane grid depends on the size of the observation area and the annotation area. The projected feature maps are stacked to obtain a size of BEV characteristics.

[0035] Through the above optional implementation, by converting RGB image features into BEV features through a projection matrix, spatial alignment of multi-view image features is achieved, so that the perspective information of different cameras can be fused on the same plane, thereby generating a BEV feature map under a global perspective, which not only improves the accuracy and efficiency of feature fusion, but also provides more comprehensive and reliable environmental perception information for subsequent target tracking.

[0036] S14, merging the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map.

[0037] Combined Figure 2 In some embodiments, after projecting each of the S RGB images to obtain S BEV features, the electronic device may merge the S BEV features into a high-dimensional BEV feature map, that is, concatenate the S BEV features in the channel dimension to obtain a first BEV feature map. Specifically, if the dimension of each BEV feature is , then the dimension of the first BEV feature map after splicing is , and then adjust the number of channels of BEV features to .

[0038] It should be noted that adjusting the number of channels to 128 helps optimize computational efficiency and model performance while ensuring feature expression capabilities, providing high-quality feature input for subsequent decoding and task heads.

[0039] Through the above optional implementation, the BEV features corresponding to the S RGB images are merged in the channel dimension to obtain a high-dimensional first BEV feature map, and the number of channels is adjusted using large kernel convolution to optimize the computational efficiency and model performance, effectively integrate multi-view information, and improve feature expression capabilities. Secondly, while ensuring feature expression capabilities, the computational efficiency and model performance are also optimized to provide high-quality feature input for subsequent decoding and task heads.

[0040] S15, decoding the first BEV characteristic map to obtain a second BEV characteristic map.

[0041] Combined Figure 2 In some embodiments, the electronic device may input the aggregated first BEV feature map into a decoder based on the ResNet-18 architecture, decode the first BEV feature map, and obtain a second BEV feature map. Specifically, the input first BEV feature map is first convolved using the initial convolution layer of the ResNet-18 architecture, and processed with the subsequent batch normalization and ReLU activation function of the initial convolution layer. Then, the first BEV feature map after the above processing is passed through the first convolution layer, the second convolution layer, and the third convolution layer of the ResNet-18 architecture in turn. In each convolution layer, the first BEV feature map is downsampled to half of the original, and the first downsampled feature, the second downsampled feature, and the third downsampled feature are obtained respectively. During the downsampling process, the output of each convolution layer is processed using the pyramid network architecture, that is, the upsampled features of each convolution layer are connected with the downsampled features of the output of the previous convolution layer in the channel dimension. The output dimension of the feature pyramid decoding is , which has the same dimension as the first BEV feature map, but each grid position has a higher receptive field, thereby realizing multi-scale fusion of features. The first BEV feature map after multi-scale feature fusion is called the second BEV feature map.

[0042] Through the above optional implementation, the first BEV feature map is decoded by the decoder of the ResNet-18 architecture and the pyramid network architecture to achieve multi-scale fusion of features, improve the receptive field of each grid position, and make the final feature map have stronger target representation capabilities.

[0043] S16, processing the second BEV feature map to obtain a probability occupancy map.

[0044] Combined Figure 2 In some embodiments, the electronic device may perform output task head processing on the second BEV feature map. Specifically, a central prediction head is constructed based on a convolutional network, an instance normalization (IN) layer, and a Relu activation function, and the central prediction head is used to reduce the dimension of the second BEV feature map to a predetermined size, for example , and obtain the probability occupancy map on the ground. The probability occupancy map represents the occupancy probability value of each position on the ground occupied by pedestrians.

[0045] S17: Obtain an initial tracking result of the target object based on the second BEV feature map and the probability occupancy map.

[0046] Combined Figure 2In some embodiments, in order to improve the accuracy of the pedestrian position prediction at the next moment, the electronic device can also use an offset head with the same structure as the center prediction head to predict the offset between the pedestrian and the corresponding grid in the probability occupancy map. The offset head learns how to adjust the predicted position of the pedestrian according to the second BEV feature map through training, so that the predicted position more accurately matches the position of the actual target object. Among them, the center prediction head is trained based on the Focal loss function, and the offset head is trained based on the L1 loss function. At the same time, the electronic device also adds a detection head for image features. The detection head predicts the center of the 2D bounding box and the estimated foot position at the bottom center of the 2D bounding box, that is, at the bottom center position of the 2D bounding box, the detection head further predicts a point to represent the foot position of the pedestrian, so that the image feature has a higher activation value at each pedestrian position, which helps to improve the accuracy of target object detection. Among them, the 2D bounding box refers to a rectangular box surrounding the pedestrian in the image, which is used to represent the position and size of the pedestrian in the image. The center of the 2D bounding box is also the center of the pedestrian, which is called the detection center. Therefore, the electronic device can initialize a set of trajectories based on the detection center and use the trajectories as initial tracking results of pedestrians. Each trajectory represents the motion path of a potential target, that is, one trajectory represents the motion path of a pedestrian and represents an initial tracking result of a pedestrian.

[0047] In order to make the similarity between different pedestrians smaller than the affinity between the same pedestrians, the electronic device can further learn re-identification features through classification tasks and metric learning tasks. Specifically, the electronic device first builds a ReID head based on convolution, IN layer and ReLU layer, then determines the center position of the pedestrian through the center detection algorithm in two planes (ground plane and image plane), and uses the ReID head to extract the shape near the center position of the pedestrian. The ground Reid features and shapes are Image Reid features, channel dimension . Among them, the ground Reid features are extracted from the second BEV feature map (i.e., the ground plane), while the image Reid features are extracted from the image features (i.e., the image plane). Then, the ground Reid features and the image Reid features are concatenated and input into a linear layer, through which the concatenated features are mapped to a new space to form a class identity distribution, where the class identity distribution represents the probability distribution of different pedestrian identities. Finally, the electronic device can train the true class identity using cross entropy loss and supervised contrast (SupervisedContrastive, SupCon) loss. Specifically, the classification task is performed through the cross entropy loss, that is, the class identity distribution is compared with the true class identity, the loss value is calculated, and the Reid head is optimized through the back propagation algorithm so that it can more accurately classify the pedestrian identity; the metric learning task is performed through the SupCon loss, that is, by calculating the similarity between the features of different pedestrians, the Reid head is optimized so that the similarity between the features of the same pedestrian is greater than the similarity between the features of different pedestrians.

[0048] The Reid head is trained through cross entropy loss so that it can accurately distinguish the identities of different pedestrians, and the Reid head is trained through SupCon loss to optimize its extracted features so that the similarity between different pedestrians is less than the affinity between the same pedestrian. That is, by continuously training and optimizing the Reid head, the Reid head can learn how to extract the unique features of pedestrians and accurately re-identify the same pedestrian in scenarios across terminal devices or across time points.

[0049] Through the above optional implementation, the accuracy of pedestrian position prediction is improved by constructing a center prediction head and an offset head; at the same time, the detection head is added to predict the 2D bounding box and foot position, thereby improving the accuracy of pedestrian detection; in addition, by learning re-identification features through classification tasks and metric learning tasks, and using cross-entropy loss and SupCon loss to train the Reid head, accurate identification of pedestrians and feature optimization are achieved, thereby improving the ability to re-identify pedestrians across scenarios.

[0050] In an optional embodiment, the method further comprises: Predicting a first prediction center of the target object at a next moment based on a Kalman filter algorithm, wherein the next moment and the current moment are two adjacent moments; The initial tracking result is updated based on the first prediction center to obtain a target tracking result corresponding to the target object.

[0051] In some embodiments, after obtaining the initial tracking results corresponding to each pedestrian, the electronic device can use a Kalman filter algorithm (such as an adaptive Kalman filter) to predict the position of each pedestrian in each initial tracking result at the next moment, called the first prediction center, and update the initial tracking result of each new person based on the predicted first prediction center area to form a new tracking result, called a target tracking result.

[0052] Among them, the adaptive Kalman filter algorithm is used to adapt to different environmental noises. It models the process noise as a Student t distribution, and the target transfer density is expressed as: ; ; in, k is the time index, for k -1 moment target state, for k The state of the target at any moment, for k The target transfer density at the moment, is the mean vector, is the covariance matrix, For Students t The degrees of freedom of the distribution, is the state transfer equation, For Students t distributed, d is the dimension of the target state, is the measurement noise covariance matrix. Since the posterior probability density of the target state is assumed to be Gaussian, when it is multiplied by the transfer density of the Student t distribution, the posterior probability density will no longer obey the Gaussian distribution, that is, the filtering process does not have a closed analytical solution. Therefore, electronic devices can introduce auxiliary random variables , approximate the target transfer density as a Gaussian sum form: ; in, The degrees of freedom are Gamma distribution. In order to solve the target state and auxiliary random variables , that is, the joint posterior probability density , a variational Bayesian method is introduced to iteratively calculate the separable approximate solution of the joint posterior density.

[0053] Compared with the prior art of tracking the target through the Kalman filter algorithm as an end-to-end algorithm, in which the process noise obeys the Gaussian distribution, resulting in the tracking performance being more sensitive to the prior parameters, the present application models the process noise as a Student's t distribution, which is a heavy-tailed distribution and is more robust to inaccurate prior parameters, and then uses the variational Bayesian method to adjust the model parameters in real time to adapt to different process noises.

[0054] In an optional implementation, updating the initial tracking result based on the first prediction center to obtain a target tracking result corresponding to the target object includes: Determine a first detection result in which the occupancy probability value in the probabilistic occupancy map is greater than a first preset threshold; Extracting a first Reid feature of a first detection center of the first detection result; Calculating a first Mahalanobis distance between the first prediction center and the first detection center; Calculate a first cosine distance between the first prediction center and the first detection center based on the first Reid feature; Matching the first predicted center with the initial tracking result based on the first Mahalanobis distance and the first cosine distance to obtain a first matching result; The initial tracking result is updated based on the first matching result to obtain the target tracking result.

[0055] In some embodiments, as each subsequent time step is processed, a two-stage matching strategy is used to connect the predicted center to the existing motion trajectory. In the first stage, after the probability occupancy map is obtained based on the center prediction head, the electronic device can implement non-maximum suppression through a 3×3 maximum pooling operation, and then extract only the pedestrians whose occupancy probability values ​​in the probability occupancy map exceed a first preset threshold (for example, set to 0.4) as the first detection result based on the initial time step, and determine the center position of each pedestrian in the first detection result as the first detection center. At the same time, the electronic device can also extract the first Reid feature corresponding to the first detection center, and input the first detection result and the first Reid feature into the subsequent data association step.

[0056] Specifically, after the electronic device predicts the first prediction center using the adaptive Kalman filter, the Mahalanobis distance between the first prediction center and the first detection center (called the first Mahalanobis distance) is calculated to measure the proximity of the two in spatial position. At the same time, the cosine distance between the first prediction center and the first detection center (called the first pre-distance) can be calculated based on the first Reid feature to measure the similarity between the two in appearance features. Then, the first Mahalanobis distance Distance to first cosine Combined into a comprehensive distance metric ,in is a weight parameter predetermined in the experiment and can be set to 0.98. If the first Mahalanobis distance exceeds a preset distance threshold (for example, set to obey a chi-square distribution with 2 degrees of freedom), D Set to infinity to prevent tracking of unreasonable motion trajectories. Finally, the electronic device can use the Hungarian algorithm to successfully match the motion trajectory corresponding to the first preset threshold (the initial tracking result corresponding to any pedestrian) from the initial tracking result to obtain a first matching result. For the successfully matched motion trajectory and the first detection center in the first matching result, update their state information (such as position, speed, etc.) and appearance features (such as color histogram, feature vector, etc.), and update the initial tracking result of the corresponding pedestrian based on the first matching result, that is, splice the matched first detection center into the motion trajectory to obtain a new motion trajectory, that is, the target tracking result. The electronic device can also continue to perform target tracking and matching processes for subsequent frames based on the target tracking results.

[0057] In the second stage, the electronic device can increase the first preset threshold to obtain a second preset threshold (for example, set to 2.5), and match the unmatched detection center and motion trajectory based on the same embodiment as above, and further update the initial tracking result, and obtain the target tracking result in combination with the first matching result. It should be noted that the first preset threshold and the second preset threshold are set as empirical values.

[0058] Subsequently, in each time step, the electronic device can continue to predict the predicted center of each detection center in the next adjacent moment, and continuously update the appearance features of the motion trajectory to cope with potential appearance changes.

[0059] In some embodiments, if the motion trajectory and the detection center are successfully matched 5 times, the motion trajectory is confirmed to be a new target. In addition, if any unmatched detection center is also set as a new motion trajectory after the two-stage matching strategy is executed, the unmatched motion trajectory will be retained for 10 time steps so that the target can be re-identified when it appears in the subsequent time step.

[0060] In an optional embodiment, the electronic device can verify the above-mentioned target recognition and tracking algorithm using a deep learning framework. Specifically, the electronic device can install the Pytorch deep learning framework on a server with a Linux system, and train a recognition and tracking model based on the Pytorch deep learning framework, and run it on a server with a Linux system, using the NVIDIA A800 graphics card in the server for training and testing. First, prepare the Wildtrack dataset, which contains video data of 7 terminal devices, with the characteristics of high occlusion and high crowd density. Specifically, there are a total of 400 frames of video, with an average of about 25 people walking in each frame. Then, extract the video frames of the dataset into image format, and adjust the image size to a specified pixel size, such as 720×1280 pixels. Then, input the image into the recognition and tracking model for processing. In the data enhancement stage, the input image is randomly scaled and cropped, and the scaling ratio is within a specified range. For example, the scaling ratio range can be set to [0.8, 1.2], and the camera intrinsic parameter matrix of the terminal device is adjusted accordingly. In addition, noise is randomly added to the translation vector of the camera extrinsic parameters to simulate the uncertainty in real scenes to avoid overfitting of the decoder of the recognition tracking model. The loss function is trained with the Adam optimizer and a single-cycle learning rate scheduler with a maximum learning rate of . Training is performed for a total of 50 cycles with a batch size of 1, and weights are updated after accumulating gradients over multiple batches, so that the effective batch size is 16. That is, the electronic device can use the data loader to load data according to the batch size of 1, and perform data augmentation, forward propagate the input image through the network, obtain the prediction result, calculate the loss value based on the prediction result and the actual result, and perform back propagation based on the loss value to calculate the gradient. After accumulating gradients over multiple batches, the optimizer is used to update the network weights. At the end of each training cycle, the current model is saved for subsequent evaluation or testing.

[0061] In some embodiments, in the target recognition and tracking process, in order to adapt to the dynamic data flow in different scenarios and improve the robustness to abnormal situations such as camera failure or no data flow of some cameras, a strategy of randomly shielding some cameras is adopted during the training process. Specifically, the electronic device can collect video data including multiple camera perspectives to ensure that the data set covers a wide range of scenes and dynamic data flows, and pre-process the video data, such as frame extraction, image enhancement, size adjustment, etc., for subsequent processing. In each training batch, some cameras are randomly selected and their input images are replaced with pure black images, so that the recognition and tracking algorithm can be trained in the case of missing some data. Further, in order to verify the effectiveness of the training method, the electronic device can use two sets of data sets of shielded cameras as training test examples for experimental verification. Specifically, first, the data set of the first three cameras shielded and the data set of the last three cameras shielded are merged, and joint training is performed on the merged data set. After the training is completed, the data set without shielding the camera is tested to evaluate its tracking effect on the complete data. In this way, the recognition and tracking algorithm can not only effectively learn in the scenario of camera failure, but also still achieve good performance on data without failure. Secondly, by simulating abnormal situations in a multi-camera system, the algorithm’s fault tolerance and adaptability are improved. Therefore, whether the camera hardware fails in actual applications or some cameras lose data due to movement of people, the algorithm can maintain relatively stable performance.

[0062] Reference Figure 4 , which is a functional module diagram of a target tracking device based on dynamic data stream according to an embodiment of the present application.

[0063] In some embodiments, the target tracking device 40 based on dynamic data stream may include multiple functional modules composed of computer program segments. The computer program of each program segment of the target tracking device 40 based on dynamic data stream may be stored in the memory of the electronic device and executed by at least one processor to perform (see Figure 1 Description) The function of target tracking based on dynamic data stream. According to the functions performed, it can be divided into multiple functional modules. The functional modules may include: image acquisition module 401, feature extraction module 402, feature projection module 403, feature merging module 404, feature decoding module 405, feature processing module 406, recognition and tracking module 407 and preprocessing module 408. The module referred to in this application refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0064] The image acquisition module 401 is used to acquire the RGB image of the target object collected by each of the S terminal devices at the current moment, and obtain S RGB images, where S is an integer greater than 1.

[0065] The feature extraction module 402 is used to extract features from a target RGB picture to obtain picture features, wherein the target RGB picture is any one of the S RGB pictures.

[0066] The feature projection module 403 is used to project the image features into BEV features.

[0067] The feature merging module 404 is used to merge the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map.

[0068] The feature decoding module 405 is used to decode the first BEV feature map to obtain a second BEV feature map.

[0069] The feature processing module 406 is used to process the second BEV feature map to obtain a probability occupancy map.

[0070] The identification and tracking module 407 is used to obtain an initial tracking result of the target object based on the second BEV feature map and the probability occupancy map.

[0071] The identification and tracking module 407 is also used to: predict the first prediction center of the target object at the next moment based on the Kalman filter algorithm, wherein the next moment and the current moment are two adjacent moments; update the initial tracking result based on the first prediction center to obtain the target tracking result corresponding to the target object.

[0072] The identification and tracking module 407 is also specifically used to: determine a first detection result in the probability occupancy map where the occupancy probability value is greater than a first preset threshold; extract a first Reid feature of a first detection center of the first detection result; calculate a first Mahalanobis distance between the first prediction center and the first detection center; calculate a first cosine distance between the first prediction center and the first detection center based on the first Reid feature; match the first prediction center with the initial tracking result based on the first Mahalanobis distance and the first cosine distance to obtain a first matching result; and update the initial tracking result based on the first matching result to obtain the target tracking result.

[0073] The identification and tracking module 407 is also used to: determine a second detection result in the probability occupancy map where the occupancy probability value is greater than a second preset threshold; extract a second Reid feature of a second detection center of the second detection result; calculate a second Mahalanobis distance between the first prediction center and the second detection center; calculate a second cosine distance between the first prediction center and the second detection center based on the second Reid feature; match the first prediction center with the initial trajectory based on the second Mahalanobis distance and the second cosine distance to obtain a second matching result; and update the initial tracking result based on the first matching result and the second matching result to obtain the target tracking result.

[0074] The feature extraction module 402 is also specifically used to: input the target RGB image into the first convolution block of the ResNet-18 architecture, so that the first convolution block downsamples the target RGB image to obtain a first feature map; input the target RGB image into the second convolution block of the ResNet-18 architecture, so that the second convolution block downsamples the target RGB image to obtain a second feature map; input the target RGB image into the third convolution block of the ResNet-18 architecture, so that the third convolution block downsamples the target RGB image to obtain a third feature map; upsample the first feature map, the second feature map and the third feature map; and concatenate the first feature map, the second feature map and the third feature map after the upsampling process to obtain the image features.

[0075] The feature decoding module 405 is further specifically used for: inputting the first BEV feature map into the first convolution layer so that the first convolution layer downsamples the first BEV feature map to obtain a first downsampled feature; inputting the first BEV feature map into the second convolution layer so that the second convolution layer downsamples the first BEV feature map to obtain a second downsampled feature; inputting the first BEV feature map into the third convolution layer so that the third convolution layer downsamples the first BEV feature map to obtain a third downsampled feature; upsampling the first downsampled feature, the second downsampled feature and the third downsampled feature; connecting the first downsampled feature, the second downsampled feature and the third downsampled feature after the upsampling in the channel dimension to obtain the second BEV feature map.

[0076] The pre-processing module 408 is used to: perform centering and normalization processing on the target RGB image; and perform image enhancement on the target RGB image after the centering and normalization processing is completed.

[0077] It should be understood that the various variations and specific embodiments of the target tracking method based on dynamic data stream provided in the above embodiments are also applicable to the target tracking device based on dynamic data stream in the present embodiment. Through the above detailed description of the target tracking method based on dynamic data stream, those skilled in the art can clearly know the implementation method of the target tracking device based on dynamic data stream in the present embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0078] See also Figure 5 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. In a preferred embodiment of the present application, the electronic device 5 includes a memory 51 , at least one processor 52 and at least one communication bus 53 .

[0079] Those skilled in the art should understand that Figure 5 The structure of the electronic device 5 shown does not constitute a limitation of the embodiments of the present application, and can be either a bus structure or a star structure. The electronic device 5 can also include more or less other hardware or software than shown in the figure, or a different component arrangement.

[0080] In some embodiments, the electronic device 5 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The electronic device 5 may also include user equipment, which includes but is not limited to any electronic product that can interact with a user through a keyboard, mouse, remote control, touchpad, or voice control device, such as a personal computer, tablet computer, smart phone, digital camera, etc.

[0081] In the above embodiments provided in the present application, it should be understood that the disclosed methods, devices, computer-readable storage media, and electronic devices can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple components or modules can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or components or modules, which can be electrical, mechanical or other forms.

[0082] The components described as separate components may or may not be physically separated, and the components shown as components may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the components may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each component may exist physically separately, or two or more modules may be integrated into one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0084] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0085] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0086] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0087] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A target tracking method based on dynamic data stream, characterized in that: The method comprises: Obtain an RGB image of the target object collected by each of the S terminal devices at the current moment, and obtain S RGB images, where S is an integer greater than 1; Perform feature extraction on a target RGB image to obtain image features, wherein the target RGB image is any one of the S RGB images; Projecting the image features into BEV features; Merging the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map; Decoding the first BEV characteristic graph to obtain a second BEV characteristic graph; Processing the second BEV characteristic map to obtain a probability occupancy map; An initial tracking result of the target object is obtained based on the second BEV feature map and the probability occupancy map.

2. The target tracking method based on dynamic data stream according to claim 1 is characterized in that: The method further comprises: Predicting a first prediction center of the target object at a next moment based on a Kalman filter algorithm, wherein the next moment and the current moment are two adjacent moments; The initial tracking result is updated based on the first prediction center to obtain a target tracking result corresponding to the target object.

3. The target tracking method based on dynamic data stream according to claim 2 is characterized in that: The updating of the initial tracking result based on the first prediction center to obtain the target tracking result corresponding to the target object includes: Determine a first detection result in which the occupancy probability value in the probabilistic occupancy map is greater than a first preset threshold; Extracting a first Reid feature of a first detection center of the first detection result; Calculating a first Mahalanobis distance between the first prediction center and the first detection center; Calculate a first cosine distance between the first prediction center and the first detection center based on the first Reid feature; Matching the first predicted center with the initial tracking result based on the first Mahalanobis distance and the first cosine distance to obtain a first matching result; The initial tracking result is updated based on the first matching result to obtain the target tracking result.

4. The target tracking method based on dynamic data stream according to claim 3 is characterized in that: Before updating the initial tracking result based on the first matching result, the method further includes: Determine a second detection result in which the occupancy probability value in the probabilistic occupancy map is greater than a second preset threshold; Extracting a second ReID feature of a second detection center of the second detection result; Calculate a second Mahalanobis distance between the first prediction center and the second detection center; Calculate a second cosine distance between the first prediction center and the second detection center based on the second Reid feature; Matching the first predicted center with the initial trajectory based on the second Mahalanobis distance and the second cosine distance to obtain a second matching result; The initial tracking result is updated based on the first matching result and the second matching result to obtain the target tracking result.

5. The target tracking method based on dynamic data stream according to claim 1, characterized in that: The feature extraction of the target RGB image to obtain the image features includes: Input the target RGB image into a first convolution block of a ResNet-18 architecture, so that the first convolution block downsamples the target RGB image to obtain a first feature map; Inputting the target RGB image into the second convolution block of the ResNet-18 architecture, so that the second convolution block downsamples the target RGB image to obtain a second feature map; Inputting the target RGB image into the third convolution block of the ResNet-18 architecture, so that the third convolution block downsamples the target RGB image to obtain a third feature map; Performing upsampling processing on the first feature map, the second feature map, and the third feature map; The first feature map, the second feature map, and the third feature map after upsampling are concatenated to obtain the image features.

6. The target tracking method based on dynamic data stream according to claim 5, characterized in that: The decoding of the first BEV characteristic graph to obtain the second BEV characteristic graph comprises: Inputting the first BEV feature map into the first convolutional layer, so that the first convolutional layer performs downsampling processing on the first BEV feature map to obtain a first downsampling feature; Inputting the first BEV feature map into the second convolutional layer, so that the second convolutional layer performs downsampling processing on the first BEV feature map to obtain a second downsampling feature; inputting the first BEV feature map into the third convolutional layer, so that the third convolutional layer performs downsampling processing on the first BEV feature map to obtain a third downsampling feature; Performing upsampling processing on the first down-sampling feature, the second down-sampling feature and the third down-sampling feature; The first down-sampling feature, the second down-sampling feature and the third down-sampling feature after the up-sampling process are connected in the channel dimension to obtain the second BEV feature map.

7. The target tracking method based on dynamic data stream according to claim 1, characterized in that: Before extracting features from the target RGB image, the method further includes: Performing centering and normalization processing on the target RGB image; Image enhancement is performed on the target RGB image after the centering and normalization processing is completed.

8. A target tracking device based on dynamic data stream, characterized in that: The device comprises: An image acquisition module is used to acquire an RGB image of the target object collected by each of the S terminal devices at the current moment, and obtain S RGB images, where S is an integer greater than 1; A feature extraction module is used to extract features of a target RGB image to obtain image features, wherein the target RGB image is any one of the S RGB images; A feature projection module, used for projecting the image features into BEV features; A feature merging module, used for merging the BEV features corresponding to each RGB picture in the S RGB pictures to obtain a first BEV feature map; A feature decoding module, used for decoding the first BEV feature map to obtain a second BEV feature map; A feature processing module, used for processing the second BEV feature map to obtain a probability occupancy map; An identification and tracking module is used to obtain an initial tracking result of the target object based on the second BEV feature map and the probability occupancy map.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the target tracking method based on dynamic data stream as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the target tracking method based on dynamic data stream according to any one of claims 1 to 7 are implemented.