A large-scale feature comparison method and system based on matrix accelerator

Through the collaborative work of the matrix accelerator and the processor, the image and motion feature extraction model is used to compare the various features of the target object, which solves the problem of low monitoring efficiency in the existing technology and realizes efficient and accurate target object tracking and monitoring.

CN119888572BActive Publication Date: 2025-09-05CHINA UNIV OF MINING & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411979937.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-09-05
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing technologies have difficulty comparing multiple aspects of the target object's features, resulting in reduced monitoring efficiency and inability to fully utilize processing resources.

Method used

Through the collaborative work of the matrix accelerator and the processor, the adapted tasks and data are processed respectively, and the image feature extraction model and the motion feature extraction model are used to obtain the multi-faceted features of the target object for comparison. The model is trained through a comprehensive loss function to improve monitoring efficiency and accuracy.

Benefits of technology

It improves monitoring efficiency and accuracy, can effectively track target objects, fully utilize processing resources, and improves the training efficiency and accuracy of image and motion feature extraction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888572B_ABST
    Figure CN119888572B_ABST
Patent Text Reader

Abstract

The present invention provides a large-scale feature comparison method and system based on a matrix accelerator, relating to the field of computer technology. The method comprises: detecting a current video frame, determining a first target area, obtaining a first feature map using an image feature extraction model, and obtaining a first motion feature vector using a motion feature extraction model; detecting a next video frame, determining a second target area and a second feature map; determining a target area to be measured of the first target object that exists in both the next video frame and the current video frame, determining a comprehensive loss function, and training the image feature extraction model and the motion feature extraction model. According to the present invention, a matrix accelerator and a processor can respectively process adapted tasks and data, thereby fully utilizing processing resources, improving processing and monitoring efficiency, obtaining multiple features of the target object for comparison, and continuously improving monitoring efficiency and accuracy during actual monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a large-scale feature comparison method and system based on a matrix accelerator. Background Art

[0002] In related technology, CN119106374A discloses a big data analysis method, system, and cloud server for smart security. The method includes: acquiring image data, audio data, and sensor data, calculating similarity matrices and distance matrices to obtain similarity matrices and distance matrices; initializing a sparse matrix and performing heuristic clustering to obtain a sparse connected graph; extracting spatiotemporal depth features and graph structure features from the sparse connected graph to obtain spatiotemporal depth features and graph structure features, which are input into a pre-trained anomaly detection model to obtain an anomaly score; generating an alarm instruction when the anomaly score is determined to be greater than a preset anomaly threshold and sending it to the user end; wherein the anomaly detection model uses historical data and a standard anomaly score as input data to construct an initial anomaly detection model, and then training the initial anomaly detection model to obtain an anomaly detection model. This method can achieve real-time detection and accurate early warning of abnormal events, ensuring data security.

[0003] CN118861937A discloses an abnormality monitoring method and device based on discriminant feature cluster analysis, which can comprehensively extract the change characteristics of normal operation sampling data from the perspective of the discriminant features of each normal data. Specifically, the abnormality monitoring method involved in this scheme performs discriminant analysis on the sampling data under each normal operating condition, and then clusters the discriminant vectors using cluster analysis, thereby using the center vector conversion calculation of the representative cluster cluster to obtain the intermediate matrix for calculating the abnormality index and the threshold for identifying the abnormality. This scheme discloses further improved technical solutions for discriminant analysis, cluster analysis and conversion calculation respectively, ensuring the comprehensiveness of normal change feature analysis. Finally, this scheme also designs an abnormality monitoring device based on the same inventive concept, which saves and provides technical parameters in real time to realize abnormality monitoring of real-time sampled sample data and provide real-time monitoring images.

[0004] Therefore, the relevant technology can process a large amount of monitoring data and detect anomalies therein, but when monitoring the actions of the target object in the video frame, it is difficult to compare the characteristics of multiple aspects of the target object, making it difficult to track the target object, and it is impossible to fully call on processing resources for feature comparison, resulting in a decrease in monitoring efficiency.

[0005] The information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to those skilled in the art. Summary of the Invention

[0006] The present invention provides a large-scale feature comparison method and system based on a matrix accelerator, which can solve the technical problems that related technologies have difficulty in comparing multiple aspects of the target object's features, difficulty in fully mobilizing processing resources, and reduced monitoring efficiency.

[0007] According to a first aspect of the present invention, a large-scale feature comparison method based on a matrix accelerator is provided, comprising:

[0008] Detecting the current video frame through a matrix accelerator, determining a first target area where multiple target objects are located in the current video frame, and sending first position information of the first target area to a processor;

[0009] Extract features from each first target area using an image feature extraction model to obtain a first feature map of each first target area;

[0010] Processing the first feature map using a motion feature extraction model to obtain a first motion feature vector of a target object in a first target area, and sending the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction;

[0011] Detecting the next video frame through the matrix accelerator, determining a second target area where multiple target objects are located in the next video frame, and sending second position information of the second target area to the processor;

[0012] Obtaining a second feature map of each second target area;

[0013] Determine, based on the first feature map and the second feature map, a target area to be measured of the first target object that exists in both the next video frame and the current video frame, and send first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor;

[0014] The processor determines a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and sends the result to the matrix accelerator;

[0015] The image feature extraction model and the motion feature extraction model are trained according to the matrix accelerator to obtain a trained image feature extraction model and a trained motion feature extraction model.

[0016] According to a second aspect of the present invention, a large-scale feature comparison system based on a matrix accelerator is provided, comprising:

[0017] a detection module, configured to detect the current video frame through a matrix accelerator, determine a first target area where a plurality of target objects are located in the current video frame, and send first position information of the first target area to a processor;

[0018] A first feature map module, configured to extract features from each first target area using an image feature extraction model to obtain a first feature map of each first target area;

[0019] a first motion feature vector module, configured to process the first feature map using a motion feature extraction model to obtain a first motion feature vector of a target object in a first target area, and send the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction;

[0020] a sending module, configured to detect the next video frame through the matrix accelerator, determine a second target area where multiple target objects are located in the next video frame, and send second position information of the second target area to the processor;

[0021] A second feature map module, used to obtain a second feature map of each second target area;

[0022] a target area to be measured module, configured to determine a target area to be measured of a first target object that exists in both a next video frame and a current video frame based on the first feature map and the second feature map, and to send first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor;

[0023] a comprehensive loss function module, configured for the processor to determine a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and send the comprehensive loss function to the matrix accelerator;

[0024] The training module is used to train the image feature extraction model and the motion feature extraction model according to the matrix accelerator to obtain the trained image feature extraction model and the trained motion feature extraction model.

[0025] By adopting the above technical solution, the present invention can achieve the following technical effects:

[0026] According to the present invention, through communication between the matrix accelerator and the processor, the matrix accelerator and the processor can respectively process the adapted tasks and data, thereby fully utilizing processing resources and improving processing and monitoring efficiency. In addition, the image feature extraction model and the motion feature extraction model can be used to obtain various features of the target object for comparison, thereby improving the accuracy of the comparison and facilitating the tracking of the target object. The image feature extraction model and the motion feature extraction model can also be continuously trained to continuously improve the monitoring efficiency and accuracy during the actual monitoring process. When determining the motion loss function, the motion loss function can be determined by the total error between the reference motion feature vector and the first motion feature vector calculated based on the actual coordinates, thereby reducing the error of the first motion feature vector during the training process and improving the accuracy of the motion feature extraction model. When determining the motion rationality loss function, the two branches can be competitively trained using a conditional function. On the basis of continuously improving the accuracy of the second branch, the second branch is used to supervise the training of the first branch. That is, the training of the image feature extraction model and the motion feature extraction model is supervised by using an increasingly high supervision standard, thereby effectively improving the training efficiency and accuracy of the image feature extraction model and the motion feature extraction model, and improving the rationality of using the target object's body shape and movement to infer motion information. When determining the feature matching loss function, the rationality of the target object's motion can be determined by the relative error between the predicted motion feature vector and the reference motion feature vector, and the judgment condition can be set based on the rationality to determine whether there is a misjudgment of the first target object and whether the two video frames need to be manually labeled. If manual labeling is not required, the feature matching loss function is directly obtained through the appearance feature vector of the first feature map. Otherwise, the feature matching loss function is determined based on the area and position of the manually labeled second target object, which can improve the accuracy of training and improve the performance of the image feature extraction model and the motion feature extraction model.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and not limiting of the present invention. Other features and aspects of the present invention will become more apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without inventive work.

[0029] Figure 1A schematic flow chart of a large-scale feature comparison method based on a matrix accelerator according to an embodiment of the present invention is exemplarily shown;

[0030] Figure 2 A block diagram of a large-scale feature comparison system based on a matrix accelerator according to an embodiment of the present invention is exemplarily shown. DETAILED DESCRIPTION

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0032] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0033] Figure 1 A flow chart of a large-scale feature comparison method based on a matrix accelerator according to an embodiment of the present invention is exemplarily shown. The method includes:

[0034] Step S101: Detecting a current video frame through a matrix accelerator to determine a first target area where multiple target objects are located in the current video frame, and sending first position information of the first target area to a processor;

[0035] Step S102, extracting features from each first target region using an image feature extraction model to obtain a first feature map of each first target region;

[0036] Step S103: Processing the first feature map using a motion feature extraction model to obtain a first motion feature vector of the target object in the first target area, and sending the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction;

[0037] Step S104: Detect the next video frame using the matrix accelerator to determine a second target area where multiple target objects are located in the next video frame, and send second position information of the second target area to the processor;

[0038] Step S105, obtaining a second feature map of each second target area;

[0039] Step S106, determining a target area to be measured of the first target object that exists in both the next video frame and the current video frame based on the first feature map and the second feature map, and sending first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor;

[0040] Step S107: The processor determines a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and sends the result to a matrix accelerator.

[0041] Step S108 : training the image feature extraction model and the motion feature extraction model according to the matrix accelerator to obtain a trained image feature extraction model and a trained motion feature extraction model.

[0042] According to the large-scale feature comparison method based on a matrix accelerator in an embodiment of the present invention, the matrix accelerator and the processor can communicate with each other so that the matrix accelerator and the processor can respectively process adapted tasks and data, thereby fully calling processing resources and improving processing and monitoring efficiency. In addition, the image feature extraction model and the motion feature extraction model can be used to obtain various features of the target object for comparison, thereby improving the accuracy of the comparison and facilitating the tracking of the target object. The image feature extraction model and the motion feature extraction model can also be continuously trained to continuously improve the monitoring efficiency and accuracy during the actual monitoring process.

[0043] According to one embodiment of the present invention, in step S101, the matrix accelerator may include a processing device suitable for parallel computing, such as a GPU, and the processor may include a general-purpose processing device, such as a CPU. The two are respectively suitable for different types of operations. For example, the matrix accelerator is suitable for performing operations on deep learning neural network models and processing data of types such as images and matrices, while the processor is suitable for general-purpose calculations, such as function operations, etc. Therefore, the matrix accelerator and the processor can be fully called to perform operations on suitable data, thereby improving the overall computing efficiency.

[0044] According to one embodiment of the present invention, a matrix accelerator can monitor the current video frame using a convolutional neural network model, determine a first target region within the current video frame where multiple target objects are located, and send first position information of the first target region to a processor. The first target region is the region within a minimum rectangular box that selects the target objects, and the first position information may include the coordinates of the four vertices of the minimum rectangular box.

[0045] According to one embodiment of the present invention, in step S102, the image feature extraction model is a deep learning neural network model, including multiple network layers such as convolutional layers, pooling layers, and activation layers, which can extract features from the first target area to obtain a first feature map for each first target area. The resolution of the first feature map is lower than that of the first target area, but the number of first feature maps is large, that is, by observing the first target area from multiple perspectives, multiple first feature maps can be obtained. Multiple first feature maps can be obtained for each first target area.

[0046] According to one embodiment of the present invention, in step S103, a first motion feature vector of the target object in the first feature map can be extracted through a motion feature extraction model and sent to a processor. The motion feature extraction model is also a deep learning neural network model, which can process the first feature map through multiple network layers to obtain a first motion feature vector, which can be used to describe the probability of the target object moving in various directions, as well as the speed at which the target object moves in various directions. The first motion feature vector can be determined based on features such as the posture, action, and body shape of the target object. For example, the posture of the target object can be used to determine the forward direction of the target object, and the action and body shape of the target object can be used to determine the movement speed of the target object. For example, if the target object's action is running and the body is tall, its movement speed is faster. If the target object's action is walking and the body is short, its movement speed is slower.

[0047] According to one embodiment of the present invention, in step S104, the matrix accelerator may monitor the next video frame in a similar manner to the above, determine the second target area where multiple target objects are located in the next video frame, and send the second position information of the second target area to the processor.

[0048] According to one embodiment of the present invention, in step S105 , the second feature map may be obtained in a manner similar to the above method of obtaining the first feature map, which will not be described in detail herein.

[0049] According to one embodiment of the present invention, in step S106, matching processing can be performed based on the first feature map and the second feature map to determine the first target object that exists in both the next video frame and the current video frame, as well as the target area to be measured where the first target object is located, that is, the target area to be measured of the first target object in the current video frame and the target area to be measured of the first target object in the next video frame, and the first target position information of the target area to be measured in the current video frame and the second target position information of the target area to be measured in the next video frame can be determined, so that the first target position information and the second target position information can be sent to the processor.

[0050] According to one embodiment of the present invention, when performing matching processing, attention may be paid to the appearance features of the target object in the first feature map and the second feature map, while features such as posture and action may be ignored. That is, in different video frames, the posture, action and other features of the same target object may be different, but the appearance features should theoretically be the same. Therefore, the first feature map and the second feature map may be processed by a convolutional neural network model to obtain the appearance feature vectors of each target object in the current video frame and the appearance feature vectors of each target object in the next video frame. The similarity (for example, cosine similarity) between the appearance feature vectors of the target object in the current video frame and the target object in the next video frame is then determined. If the target object in the next video frame has a similarity to the target object that exceeds a preset similarity threshold, then the target object is the first target object that exists in both the current video frame and the next video frame.

[0051] According to one embodiment of the present invention, in step S107, the processor may determine the comprehensive loss function of the image feature extraction model and the motion feature extraction model based on the data sent to the processor above, thereby training the two models to continuously improve the monitoring accuracy in actual monitoring.

[0052] According to one embodiment of the present invention, the processor determines a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, first target position information, second target position information, a first feature map, a second feature map, the first position information, and the second position information, and sends the result to a matrix accelerator, including: determining a reference motion feature vector based on the first target position information and the second target position information; processing the first feature map through a body shape feature extraction model to obtain a first body shape feature vector of a first target object in a first target area; processing the first feature map through a motion feature extraction model to obtain a first motion feature vector of a first target object in the first target area; determining a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first feature map, the second feature map, the first body shape feature vector, the first motion feature vector, the reference motion feature vector, the first position information, and the second position information; and sending the comprehensive loss function to the matrix accelerator.

[0053] According to one embodiment of the present invention, the first motion feature vector can be used to describe the motion information of the target object, that is, the probability of the target object moving in various directions, and the speed at which the target object moves in various directions. The first motion feature vector has a specific form. For example, the first motion feature vector has 8 components, the first 4 components are the probability of the target object moving upward, downward, left, and right in the video frame, and the last 4 components are the speed of the target object moving upward, downward, left, and right in the video frame. The reference motion feature vector can also be set in the above form, and can be set according to the actual data of the first target position information and the second target position information. First, the centroid of the rectangular box corresponding to the first target position information can be used to refer to the first target position information, and the centroid of the rectangular box corresponding to the second target position information can be used to refer to the second target position information. Subsequently, the reference motion feature vector can be further calculated. For example, if the second target position information is located to the upper left of the first target position information, the probability of the target object moving to the left and the probability of moving upward can be set to 1, and the probability of moving in the other two directions can be set to 0. The speed of moving to the left is equal to the ratio of the horizontal coordinate difference between the first target position information and the second target position information to the time difference between the two video frames. The speed of moving upward is equal to the ratio of the total coordinate difference between the first target position information and the second target position information to the time difference between the two video frames. The speed of moving in the other two directions can be set to 0.

[0054] According to one embodiment of the present invention, the matrix accelerator may use a body shape feature extraction model to extract a first body shape feature vector of the first target object in the first feature map. The body shape feature extraction model is a deep learning neural network model. The first body shape feature vector may be used to describe the body shape of the first target object, for example, the height of the first target object, and the proportions between various body parts, etc., and the first body shape feature vector may be sent to the processor.

[0055] According to one embodiment of the present invention, the matrix accelerator may use a motion feature extraction model to extract the first motion feature vector of the first target object in the first feature map. The motion feature extraction model is a deep learning neural network model. The first motion feature vector may be used to describe the type of motion being performed by the first target object. For example, it may be used to describe the probability of the first target object performing various types of motions, for example, the probability of running is 60%, the probability of walking is 30%, the probability of jumping is 10%, etc. The first motion feature vector may also be used to describe the direction of motion of the first target object. For example, it may describe the probability of the first target object moving upward, downward, leftward, and rightward. After obtaining the first motion feature vector, it may be sent to the processor.

[0056] According to one embodiment of the present invention, the processor may calculate the comprehensive loss function of the image feature extraction model and the motion feature extraction model based on the information obtained above. The comprehensive loss function of the image feature extraction model and the motion feature extraction model is determined based on the first motion feature vector, the first feature map, the second feature map, the first body shape feature vector, the first action feature vector, the reference motion feature vector, the first position information, and the second position information, including: determining a motion loss function based on the first motion feature vector and the reference motion feature vector; determining a motion rationality loss function based on the first motion feature vector, the first body shape feature vector, the first action feature vector, and the reference motion feature vector; determining a feature matching loss function based on the first feature map, the second feature map, the first position information, and the second position information; and determining the comprehensive loss function based on the motion loss function, the motion rationality loss function, and the feature matching loss function.

[0057] According to one embodiment of the present invention, the first motion feature vector is a vector describing the motion of the target object obtained by the motion feature extraction model, and the reference motion feature vector is a vector describing the motion of the target object obtained based on the actual coordinate data of the first target position information and the second target position information. The reference motion feature vector is used as a reference to determine the error of the first motion feature vector, thereby obtaining the motion loss function, so as to minimize the motion loss function during the training process and improve the accuracy of the motion feature extraction model.

[0058] According to one embodiment of the present invention, determining a motion loss function according to the first motion feature vector and the reference motion feature vector includes: determining the motion loss function L according to formula (1): S ,

[0059]

[0060] Among them, S 1,i is the first motion feature vector of the i-th target area to be measured, S R,i is the reference motion feature vector of the i-th target area to be measured, n T is the number of target areas to be measured, i≤n T , and i and n T All are positive integers.

[0061] According to one embodiment of the present invention, |S 1,i -S R,i | is the error between the first motion feature vector and the reference motion feature vector, It is the total error between the first motion feature vector of the desired first target object and the reference motion feature vector in both video frames. The total error can be used as a motion loss function, thereby reducing the motion loss function during the training process and improving the accuracy of the motion feature extraction model.

[0062] In this way, the motion loss function can be determined by the total error between the reference motion feature vector and the first motion feature vector calculated based on the actual coordinates, thereby reducing the error of the first motion feature vector during training and improving the accuracy of the motion feature extraction model.

[0063] According to one embodiment of the present invention, a motion rationality loss function is determined based on the first motion feature vector, the first body shape feature vector, the first action feature vector and the reference motion feature vector, including: processing the first body shape feature vector and the first action feature vector according to a motion prediction model to obtain a predicted motion feature vector; and determining the motion rationality loss function based on the predicted motion feature vector, the first motion feature vector and the reference motion feature vector.

[0064] According to one embodiment of the present invention, the motion and speed of the target object can be determined based on its body shape and motion. Therefore, the motion and speed of the target object can be analyzed, thereby analyzing the rationality of the first motion feature vector and the reference motion feature vector of the target object, and then training the motion feature extraction model to obtain a more reasonable and accurate first motion feature vector.

[0065] According to one embodiment of the present invention, the motion prediction model is a deep learning neural network model, for example, a BP neural network model, which can process the first body shape feature vector and the first action feature vector to obtain a predicted motion feature vector. The predicted motion feature vector is a relatively reasonable motion feature vector calculated by the motion prediction model based on the body shape of the target object described by the first body shape feature vector and the action of the target object described by the first action feature vector, that is, a reasonable motion feature vector determined based on the body shape and action of the target object.

[0066] According to one embodiment of the present invention, determining a motion rationality loss function based on the predicted motion feature vector, the first motion feature vector and the reference motion feature vector includes: determining the motion rationality loss function L according to formula (2): R ,

[0067]

[0068] Among them, S 1,i is the first motion feature vector of the i-th target area to be measured, S R,i is the reference motion feature vector of the i-th target area to be measured, nT is the number of target areas to be measured, S p,i is the predicted motion feature vector of the i-th target area to be measured, w1 and w2 are preset weights, n T is the number of target areas to be measured, i≤n T , and i and n T All are positive integers, and if is a conditional function.

[0069] According to one embodiment of the present invention, the conditional function Indicates that in |S p,i -S R,i |≤|S 1,i -S R,i |, the conditional function value is Otherwise W1|S p,i -S R,i |+w2|S 1,i -S R,i As described above, the first motion feature vector is vector data obtained by processing the first target area using the image feature extraction model and the motion feature extraction model. The image feature extraction model and the motion feature extraction model serve as the first branch. The predicted motion feature vector is vector data obtained using the image feature extraction model, the body shape feature extraction model, the motion feature extraction model, and the motion prediction model. The image feature extraction model, the body shape feature extraction model, the motion feature extraction model, and the motion prediction model serve as the second branch. The above conditional function enables competitive training of the two branches, thereby improving training efficiency and the performance of the image feature extraction model and the motion feature extraction model.

[0070] According to one embodiment of the present invention, in |S p,i -S R,i |≤|S 1,i -S R,i |, indicating that the error of the first motion feature vector is large, that is, the error of the first branch is large. In this case, you can use As the conditional function value, that is, using The error between the first motion feature vector and the reference motion feature vector is amplified as a magnification factor |S 1,i -S R,i |, thereby improving the intensity and pertinence of training and quickly improving the accuracy of the first branch. Moreover, |S p,i -S R,i | is the error between the predicted motion feature vector and the reference motion feature vector, so, is the relative error between the predicted motion feature vector and the reference motion feature vector, It represents the similarity between the predicted motion feature vector and the reference motion feature vector, that is, the similarity between the predicted motion feature vector obtained based on the target object's body shape and movement inference and the actual reference motion feature vector. It can also represent the rationality of the inference based on body shape and movement. The higher the rationality, the smaller the magnification coefficient when this item is used as the denominator. Conversely, the lower the rationality, the greater the error. Therefore, the rationality can be used as the denominator as the magnification coefficient to efficiently train the first branch.

[0071] According to one embodiment of the present invention, in |S p,i -S R,i |>|S 1,i -S R,i |, it means that the error of the predicted motion feature vector is large, that is, the error of the second branch is large. In this case, w1|S p,i -S R,i |+w2|S 1,i -S R,i | is used as the conditional function value, thereby training both branches simultaneously. Therefore, when the error of the first branch is small, both branches can be trained simultaneously to improve the accuracy of both branches. Once the error of the first branch is large, the first branch can be trained in a focused manner. In other words, the second branch can be used to supervise the training of the first branch. As the accuracy of the second branch continues to improve, the accuracy of the first branch is monitored to see if it has effectively improved. If it has not, the first branch is trained in a focused manner, thereby effectively improving the training efficiency and accuracy of the first branch under increasingly higher supervision standards.

[0072] In this way, the two branches can be competitively trained through the conditional function. On the basis of continuously improving the accuracy of the second branch, the second branch is used to supervise the training of the first branch. That is, the training of the image feature extraction model and the motion feature extraction model is supervised using continuously improving supervision standards, thereby effectively improving the training efficiency and accuracy of the image feature extraction model and the motion feature extraction model, and improving the rationality of using the target object's body shape and movement to infer motion information.

[0073] According to one embodiment of the present invention, determining a feature matching loss function based on the first feature map, the second feature map, the first position information, and the second position information includes: determining the feature matching loss function L according to formula (3) M ,

[0074]

[0075] Among them, O 1i is the appearance feature vector obtained based on the first feature map of the i-th target area to be measured, O 2,iis the appearance feature vector obtained based on the second feature map of the i-th target area to be tested, K p is the preset coefficient threshold, O 1,AN,j is the appearance feature vector of the target area of ​​the j-th second target object in the current video frame that exists in both the next video frame and the current video frame based on the annotation information, (0 1,AN,j ) T O 1,AN,j The transposed vector of O 2,AN,j is the appearance feature vector of the target area of ​​the j-th second target object in the next video frame determined based on the annotation information and existing in both the next video frame and the current video frame, S AN,j is a labeled motion feature vector determined based on the first position information of the target area of ​​the j-th second target object in the current video frame and the second position information of the target area of ​​the j-th second target object in the next video frame, S 1,j The first motion feature vector of the j-th second target object is a zero vector if the j-th second target object is not the first target object, N is the number of second target objects, j≤N, and j and N are both positive integers.

[0076] According to one embodiment of the present invention, in formula (3), as described above, is the relative error between the predicted motion feature vector and the reference motion feature vector, is the average value of the relative error. If the average value is less than or equal to the preset coefficient threshold, it means that the movement of the first target object is relatively reasonable, and the total error between the appearance feature vector obtained by the first feature map and the appearance feature vector obtained by the second feature map can be used. As a conditional function value, it also serves as a feature matching loss function. As mentioned above, the posture, movement and other features of the same target object may be different, but the appearance features should theoretically be the same. Therefore, the appearance feature vectors of the same target object in two video frames should theoretically be the same. The sum of the errors between the two can be used as a feature matching loss function, thereby reducing the error between the two during the training process and improving the consistency of the appearance feature vectors of the same target object in the two video frames.

[0077] According to one embodiment of the present invention, in formula (3), if the average value of the relative error is greater than the preset coefficient threshold, it indicates that the movement of the first target object is unreasonable and there may be an error, that is, there is an error in determining the target object that exists in both the current video frame and the next video frame. Two different target objects may be mistakenly mistaken for the same target object, resulting in the position between the two being unreasonable compared to the size and movement of the target object in the current video frame. For example, the target object in the current video frame is short and moves slowly, but in the next video frame, the target object appears at a farther position, that is, it moves a large amount. Then, the position of the target object is unreasonable relative to its possible movement, and therefore, the target object may be misjudged. Therefore, in this case, the target object existing in both video frames can be manually marked. The target object existing in both video frames confirmed based on the marking information is the second target object, and the area where the second target object is located in the two video frames can also be determined.

[0078] According to one embodiment of the present invention, is the cosine similarity of the appearance feature vectors of the second target object in the two video frames based on the annotation information. Theoretically, the cosine similarity between the two is 1. However, if there is an error in the process of extracting the feature map by the image feature extraction model, the cosine similarity will be less than 1. Therefore, the error between the cosine similarity and 1 can be used as the term of the feature matching loss function. On the other hand, S AN,j The method for obtaining the tagged motion feature vector is based on the actual position of the second target object in the two video frames. The determination method is similar to the determination method of the reference motion feature vector described above and will not be repeated here. AN,j -S 1,j | is the error between the labeled motion feature vector and the first motion feature vector. If the jth second target object is not the first target object, that is, there is no first motion feature vector, then the first motion feature vector of the jth second target object is zero. If the jth second target object is the first target object, its first motion feature vector is directly substituted into this term to determine the error between the labeled motion feature vector and the first motion feature vector. Multiplying these two terms together yields the total error of the image feature extraction model and the motion feature extraction model for the jth second target object. Summing the total errors for all second target objects yields the feature matching loss function.

[0079] In this way, the rationality of the target object's motion can be determined by the relative error between the predicted motion feature vector and the reference motion feature vector, and the judgment condition can be set based on the rationality to determine whether there is a misjudgment of the first target object and whether the two video frames need to be manually labeled. If manual labeling is not required, the feature matching loss function is directly obtained through the appearance feature vector of the first feature map. Otherwise, the feature matching loss function is determined based on the area and position of the manually labeled second target object, which can improve the accuracy of training and improve the performance of the image feature extraction model and the motion feature extraction model.

[0080] According to one embodiment of the present invention, after obtaining the motion loss function, motion rationality loss function, and feature matching loss function, a weighted sum of the three can be used to obtain a comprehensive loss function. This comprehensive loss function is then sent to a matrix accelerator, which performs backpropagation and adjusts the parameters of each model.

[0081] According to one embodiment of the present invention, in step S108, the above-mentioned comprehensive loss function can be back-propagated, and the parameters of each model can be adjusted using the gradient descent method to obtain a trained image feature extraction model and a trained motion feature extraction model. These models can be used to monitor subsequent video frames, thereby tracking the position of each target object and predicting its motion trajectory.

[0082] According to the large-scale feature comparison method based on a matrix accelerator according to an embodiment of the present invention, the matrix accelerator and the processor can communicate with each other so that the matrix accelerator and the processor can respectively process the adapted tasks and data, thereby fully calling processing resources and improving processing and monitoring efficiency. In addition, the image feature extraction model and the motion feature extraction model can be used to obtain various features of the target object for comparison, thereby improving the accuracy of the comparison and facilitating the tracking of the target object. The image feature extraction model and the motion feature extraction model can also be continuously trained to continuously improve the monitoring efficiency and accuracy during the actual monitoring process. When determining the motion loss function, the motion loss function can be determined by the total error between the reference motion feature vector and the first motion feature vector obtained by calculation based on the actual coordinates, thereby reducing the error of the first motion feature vector during the training process and improving the accuracy of the motion feature extraction model. When determining the motion rationality loss function, the two branches can be competitively trained using a conditional function. On the basis of continuously improving the accuracy of the second branch, the second branch is used to supervise the training of the first branch. That is, the training of the image feature extraction model and the motion feature extraction model is supervised using continuously improving supervision standards, thereby effectively improving the training efficiency and accuracy of the image feature extraction model and the motion feature extraction model, and improving the rationality of using the target object's body shape and movement to infer motion information. When determining the feature matching loss function, the rationality of the target object's motion can be determined by the relative error between the predicted motion feature vector and the reference motion feature vector, and the judgment condition can be set based on this rationality to determine whether there is a misjudgment of the first target object and whether the two video frames need to be manually labeled. If manual labeling is not required, the feature matching loss function is directly obtained through the appearance feature vector of the first feature map. Otherwise, the feature matching loss function is determined based on the area and position of the manually labeled second target object. This can improve the accuracy of training and the performance of the image feature extraction model and the motion feature extraction model.

[0083] Figure 2 A block diagram of a large-scale feature comparison system based on a matrix accelerator according to an embodiment of the present invention is exemplarily shown, wherein the system includes:

[0084] a detection module, configured to detect the current video frame through a matrix accelerator, determine a first target area where a plurality of target objects are located in the current video frame, and send first position information of the first target area to a processor;

[0085] A first feature map module, configured to extract features from each first target area using an image feature extraction model to obtain a first feature map of each first target area;

[0086] a first motion feature vector module, configured to process the first feature map using a motion feature extraction model to obtain a first motion feature vector of a target object in a first target area, and send the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction;

[0087] a sending module, configured to detect the next video frame through the matrix accelerator, determine a second target area where multiple target objects are located in the next video frame, and send second position information of the second target area to the processor;

[0088] A second feature map module, used to obtain a second feature map of each second target area;

[0089] a target area to be measured module, configured to determine a target area to be measured of a first target object that exists in both a next video frame and a current video frame based on the first feature map and the second feature map, and to send first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor;

[0090] a comprehensive loss function module, configured for the processor to determine a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and send the comprehensive loss function to the matrix accelerator;

[0091] The training module is used to train the image feature extraction model and the motion feature extraction model according to the matrix accelerator to obtain the trained image feature extraction model and the trained motion feature extraction model.

[0092] Those skilled in the art will appreciate that the embodiments of the present invention described above and shown in the accompanying drawings are intended to be illustrative only and are not intended to limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functional and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Any variations or modifications may be made to the embodiments of the present invention without departing from the principles described.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large-scale feature comparison method based on a matrix accelerator, characterized in that: include: Detecting the current video frame through a matrix accelerator, determining a first target area where multiple target objects are located in the current video frame, and sending first position information of the first target area to a processor; Extract features from each first target area using an image feature extraction model to obtain a first feature map of each first target area; Processing the first feature map using a motion feature extraction model to obtain a first motion feature vector of a target object in a first target area, and sending the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction; Detecting the next video frame through the matrix accelerator, determining a second target area where multiple target objects are located in the next video frame, and sending second position information of the second target area to the processor; Obtaining a second feature map of each second target area; Determine, based on the first feature map and the second feature map, a target area to be measured of the first target object that exists in both the next video frame and the current video frame, and send first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor; The processor determines a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and sends the result to the matrix accelerator; Training the image feature extraction model and the motion feature extraction model according to the matrix accelerator to obtain a trained image feature extraction model and a trained motion feature extraction model; The processor determines a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and sends the result to the matrix accelerator, including: determining a reference motion feature vector according to the first target position information and the second target position information; Processing the first feature map using a body shape feature extraction model to obtain a first body shape feature vector of a first target object in a first target area; Processing the first feature map using a motion feature extraction model to obtain a first motion feature vector of a first target object in a first target area; determining a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first feature map, the second feature map, the first body shape feature vector, the first action feature vector, the reference motion feature vector, the first position information, and the second position information; Sending the comprehensive loss function to a matrix accelerator; Determining a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first feature map, the second feature map, the first body shape feature vector, the first action feature vector, the reference motion feature vector, the first position information, and the second position information includes: determining a motion loss function according to the first motion feature vector and the reference motion feature vector; determining a motion rationality loss function based on the first motion feature vector, the first body shape feature vector, the first action feature vector, and the reference motion feature vector; Determining a feature matching loss function based on the first feature map, the second feature map, the first position information, and the second position information; Determining the comprehensive loss function according to the motion loss function, the motion rationality loss function and the feature matching loss function; Determining a motion rationality loss function according to the first motion feature vector, the first body shape feature vector, the first action feature vector, and the reference motion feature vector includes: Processing the first body shape feature vector and the first motion feature vector according to the motion prediction model to obtain a predicted motion feature vector; determining a motion rationality loss function according to the predicted motion feature vector, the first motion feature vector, and the reference motion feature vector; Determining a motion rationality loss function according to the predicted motion feature vector, the first motion feature vector, and the reference motion feature vector includes: According to the formula Determine the motion rationality loss function L R , where S 1,i is the first motion feature vector of the i-th target area to be measured, S R,i is the reference motion feature vector of the i-th target area to be measured, n T is the number of target areas to be measured, S p,i is the predicted motion feature vector of the i-th target area to be measured, w1 and w2 are preset weights, n T is the number of target areas to be measured, i≤n T , and i and n T All are positive integers, and if is a conditional function.

2. The large-scale feature comparison method based on matrix accelerator according to claim 1, characterized in that: Determining a motion loss function according to the first motion feature vector and the reference motion feature vector includes: According to the formula Determine the motion loss function L S , where S 1,i is the first motion feature vector of the i-th target area to be measured, S R,i is the reference motion feature vector of the i-th target area to be measured, n T is the number of target areas to be measured, i≤n T , and i and n T All are positive integers.

3. The large-scale feature comparison method based on matrix accelerator according to claim 1, characterized in that: Determining a feature matching loss function according to the first feature map, the second feature map, the first position information, and the second position information includes: According to the formula Determine the feature matching loss function L M , where O 1,i is the appearance feature vector obtained based on the first feature map of the i-th target area to be measured, O 2,i is the appearance feature vector obtained from the second feature map based on the i-th target area to be tested, K p is the preset coefficient threshold, O 1,AN,j is the appearance feature vector of the target area of ​​the j-th second target object in the current video frame that exists in both the next video frame and the current video frame based on the annotation information, ( 1,AN,j ) T O 1,AN,j The transposed vector of O 2,AN,j is the appearance feature vector of the target area of ​​the j-th second target object in the next video frame determined based on the annotation information and existing in both the next video frame and the current video frame, S AN,j is a labeled motion feature vector determined based on the first position information of the target area of ​​the j-th second target object in the current video frame and the second position information of the target area of ​​the j-th second target object in the next video frame, S 1,j The first motion feature vector of the j-th second target object is a zero vector if the j-th second target object is not the first target object, N is the number of second target objects, j≤N, and j and N are both positive integers.

4. A large-scale feature comparison system based on a matrix accelerator, used to perform the method according to any one of claims 1 to 3, characterized in that: include: a detection module, configured to detect the current video frame through a matrix accelerator, determine a first target area where a plurality of target objects are located in the current video frame, and send first position information of the first target area to a processor; A first feature map module, configured to extract features from each first target area using an image feature extraction model to obtain a first feature map of each first target area; a first motion feature vector module, configured to process the first feature map using a motion feature extraction model to obtain a first motion feature vector of a target object in a first target area, and send the first motion feature vector to a processor, wherein the first motion feature vector is used to describe the probability of the target object moving in each direction and the speed of the target object moving in each direction; a sending module, configured to detect the next video frame through the matrix accelerator, determine a second target area where multiple target objects are located in the next video frame, and send second position information of the second target area to the processor; A second feature map module, used to obtain a second feature map of each second target area; a target area to be measured module, configured to determine a target area to be measured of a first target object that exists in both a next video frame and a current video frame based on the first feature map and the second feature map, and to send first target position information of the target area to be measured in the current video frame and second target position information of the target area to be measured in the next video frame to a processor; a comprehensive loss function module, configured for the processor to determine a comprehensive loss function of an image feature extraction model and a motion feature extraction model based on the first motion feature vector, the first target position information, the second target position information, the first feature map, the second feature map, the first position information, and the second position information, and send the comprehensive loss function to the matrix accelerator; The training module is used to train the image feature extraction model and the motion feature extraction model according to the matrix accelerator to obtain the trained image feature extraction model and the trained motion feature extraction model.

Citation Information

Patent Citations

  • Flame detection method, feature extraction model training method and device

    CN117274735A