Traffic video anomaly detection method based on enhanced object prompt

By introducing enhanced object prompt mechanism and instance-level loss in the traffic video abnormality detection method, the problems of dynamic background interference and spatial positioning under the on-board camera are solved, and the recognition ability and accuracy of abnormal behavior are improved.

CN120220086APending Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209341.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Under the on-board camera, the problems of dynamic background interference, difficult to identify abnormal behaviors at distant or edges, and difficult to achieve spatial positioning exist in the existing traffic video anomaly detection methods.

Method used

The traffic video anomaly detection method based on enhanced object prompts is adopted. By integrating the detected traffic objects into the frame-level traffic anomaly detection framework and introducing instance-level losses for supervision, the relationship between traffic objects and scenes, objects and objects is effectively captured.

Benefits of technology

It reduces dynamic background interference, improves the ability to identify abnormal behaviors at distant or edges, and achieves precise spatial positioning of abnormal behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220086A_ABST
    Figure CN120220086A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic video anomaly detection method based on enhanced object prompt. The invention provides a method for enhancing object prompt. A detected traffic object is integrated into a traffic anomaly detection framework comprising a frame-level feature extraction module. An object prompt network is introduced and comprises an object prompt encoder and two aggregation modules based on attention, and the two aggregation modules are an instance aggregation module used for information fusion between object instances and scenes and a relation aggregation module used for capturing the relation between objects. In addition, an instance-level loss is also designed to supervise object-level anomaly detection. According to the method provided by the invention, the interference of a dynamic background is effectively reduced, the detection capability of a long-distance or off-center abnormal object is improved, and the accurate spatial positioning of the abnormal object is realized. Experimental results on DoTA and DADA-2000 data sets show that the method provided by the invention reaches the most advanced performance level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to a traffic video anomaly detection method in the technical field of computer vision, and particularly relates to a traffic video anomaly detection method based on enhanced object cues. Background Art

[0002] Traffic video anomaly detection aims to use video analysis methods to identify abnormal behaviors or events in traffic video streams. And using computer vision technology to detect traffic video anomalies is a very promising method. In the pursuit of effective traffic video anomaly detection, especially from the perspective of the ego vehicle under an on-vehicle camera, there are three key problems to be solved:

[0003] 1) Dynamic background interference: The on-vehicle camera moves rapidly with the vehicle, and the rapidly changing background has a great impact on the model's recognition ability.

[0004] 2) Difficulty in recognizing abnormal behaviors in the distance or at the edge: The model often emphasizes objects in the center of the view, and is prone to ignoring objects in the distance or at the edge of the view, resulting in missed judgments of abnormal behaviors in the distance or at the edge of the view.

[0005] 3) Difficulty in achieving spatial positioning: Lack of accurate spatial positioning ability for abnormal behaviors.

[0006] In existing methods, dynamic background interference is alleviated by introducing language cue information or integrating high-frequency information, but few works deal with the problems of difficulty in recognizing abnormal behaviors in the distance or at the edge and difficulty in achieving spatial positioning with a clear design. Summary of the Invention

[0007] To solve the problems existing in the background art, the present invention provides a traffic video anomaly detection method based on enhanced object cues. The present invention effectively captures the relationships between traffic objects and scenes, and between objects and objects by integrating the detected traffic objects into the frame-level traffic anomaly detection framework through an object cue mechanism and introducing instance-level loss for supervision.

[0008] The technical solution adopted by the present invention is:

[0009] I. A traffic video anomaly detection method based on enhanced object cues

[0010] 1) Obtain a traffic video dataset;

[0011] 2) Construct a neural network model based on enhanced object cues. After training the neural network model based on enhanced object cues with the traffic video dataset, obtain a traffic video anomaly detection network model;

[0012] 3) Input the traffic video data to be detected into the traffic video anomaly detection network model, and the model outputs the traffic video anomaly detection result.

[0013] In the above (2), the neural network model based on enhanced object cues includes a frame-level feature extraction module, an object detector, an object cue encoder, an instance aggregation module, a relationship aggregation module, a long short-term memory network module, an instance-level anomaly detection head, and a frame-level anomaly detection head; each traffic video segment is used as the input of the frame-level feature extraction module, and the frame-level feature extraction module outputs frame embeddings and frame tokens, where the frame embeddings are used as the first input of the instance aggregation module, the last frame of the current traffic video segment is used as the input of the object detector, the object detector outputs the traffic object bounding box coordinates and uses them as the input of the object cue encoder, the object cue encoder outputs cue tokens and uses them as the second input of the instance aggregation module, the instance aggregation module outputs instance-enhanced frame embeddings and object tokens, generates the average frame embeddings according to the instance aggregation module and uses them as the first input of the relationship aggregation module, and uses the object tokens as the second input of the relationship aggregation module, the relationship aggregation module outputs enhanced object tokens and enhanced frame tokens, the enhanced object tokens are used as the input of the instance-level anomaly detection head, the instance-level anomaly detection head outputs the instance-level detection result, concatenates the frame tokens and the enhanced frame tokens and uses them together with the previous state output by the long short-term memory network module as the input of the long short-term memory network module, and the long short-term memory network module is connected to the frame-level anomaly detection head, and the frame-level anomaly detection head outputs the frame instance-level detection result.

[0014] The frame-level feature extraction module includes a backbone network, a feature pyramid network, and a dimensionality reducer connected in sequence, where the output of the feature pyramid network is denoted as frame embeddings, and the output of the dimensionality reducer is denoted as frame tokens.

[0015] The cue tokens include two position tokens and a learnable token, and the two position tokens are obtained by performing position encoding on the upper left coordinate and the lower right coordinate of the traffic object bounding box respectively.

[0016] The instance aggregation module includes L I cascaded instance-level blocks in sequence, and each instance-level block includes a multi-head self-attention layer, a first multi-head cross-attention layer, a multi-layer perceptron, and a second multi-head cross-attention layer connected in sequence; the cue tokens are used as the input of the multi-head self-attention layer of the first instance-level block, the frame embeddings are used as the input of the first multi-head cross-attention layer of the first instance-level block, the output of the second multi-head cross-attention layer of the previous instance-level block is used as the input of the first multi-head cross-attention layer of the next instance-level block, the output of the multi-layer perceptron of the previous instance-level block is used as the input of the multi-head self-attention layer of the next instance-level block, the output of the multi-layer perceptron of the last instance-level block is denoted as object tokens, and the output of the second multi-head cross-attention layer of the last instance-level block is denoted as instance-enhanced frame embeddings.

[0017] The relationship aggregation module includes L R cascaded relationship-level blocks in sequence. Each relationship-level block includes a first Cross-Transformer layer, a second Cross-Transformer layer, and a third Cross-Transformer layer connected in sequence; N K learnable query tokens and the object tokens output by the instance aggregation module are used as the input of the first Cross-Transformer layer of the first relationship-level block. The object tokens are used as the input of the first Cross-Transformer layer and the third Cross-Transformer layer. The average frame embedding is used as the input of the second Cross-Transformer layer of each relationship-level block. The output of the second Cross-Transformer layer of the previous relationship-level block is used as the input of the first Cross-Transformer layer of the next relationship-level block. The output of the third Cross-Transformer layer of the previous relationship-level block is used as the input of the first Cross-Transformer layer and the third Cross-Transformer layer of the next relationship-level block. The output of the second Cross-Transformer layer of the last relationship-level block is denoted as the enhanced frame token, and the output of the third Cross-Transformer layer of the last relationship-level block is denoted as the enhanced object token.

[0018] In 2), during the training process of the neural network model based on the enhanced object prompt, its loss function is a loss function based on cross-entropy, specifically including a frame-level loss and an instance-level loss. The formula is as follows:

[0019]

[0020] Where is the value of the loss function based on cross-entropy, is the frame-level loss value, is the instance-level loss, y t is the abnormal annotation of the f t frame image, y t ∈ {0, 1}, is the frame-level abnormal score of the image f t , is the set of true bounding boxes of the abnormal objects in the image f t , is the set of all traffic object bounding boxes obtained by the object detector, σ i represents the matching coefficient, σ i = 1 indicates that the predicted bounding box has been matched, σ i = 0 indicates that the bounding box is not matched, NO is the number of detected traffic objects, α is a scaling factor used to balance positive and negative samples, and s i is the image f t is the anomaly score of the i-th object in

[0021] The backbone network of the frame-level feature extraction module adopts the Video Swin Transformer model; the object detector adopts the YOLOv9 model; the object prompt encoder adopts the prompt encoder of the SAM model.

[0022] II. A computer device

[0023] The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method are implemented.

[0024] III. A computer-readable storage medium

[0025] A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method are implemented.

[0026] IV. A computer program product

[0027] The product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The present invention introduces an object prompt mechanism into the frame-level traffic anomaly detection framework, uses the detected traffic objects to guide the model's attention to the areas that need to be concerned, reduces the interference of dynamic backgrounds, and captures the relationship between objects and scenes at the same time.

[0030] 2. The present invention uses bipartite graph matching based on the Hungarian algorithm to classify the detected traffic objects into anomalies and non-anomalies, designs an instance-level loss function based on this, provides instance-level supervision information, and realizes accurate spatial localization of anomalies. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of the method according to an embodiment of the present invention.

[0032] Figure 2 is a schematic diagram of the network model in an embodiment of the present invention.

[0033] Figure 3 is a schematic diagram of the network structure of the instance aggregation module.

[0034] Figure 4Schematic diagram of the network structure of the relationship aggregation module. Detailed implementation manners

[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0036] The embodiments and the specific implementation process of the present invention are as follows:

[0037] The present invention proposes a traffic video anomaly detection method based on enhanced object cues, as Figure 1 shown, the method includes the following steps:

[0038] 1) Obtain a traffic video dataset;

[0039] 2) Construct a neural network model based on enhanced object cues. After training the neural network model based on enhanced object cues using the traffic video dataset, a traffic video anomaly detection network model is obtained;

[0040] As Figure 2 shown, the neural network model based on enhanced object cues includes a frame-level feature extraction module, an object detector, an object cue encoder, an instance aggregation module, a relationship aggregation module, a long short-term memory network module, an instance-level anomaly detection head, and a frame-level anomaly detection head; each traffic video segment is used as the input of the frame-level feature extraction module, and the frame-level feature extraction module outputs a frame embedding and a frame token, where the frame embedding is used as the first input of the instance aggregation module, the last frame of the current traffic video segment is used as the input of the object detector, the object detector outputs the traffic object bounding box coordinates of the last frame of the current traffic video segment and uses them as the input of the object cue encoder, the object cue encoder outputs a cue token and uses it as the second input of the instance aggregation module, the instance aggregation module outputs an instance-enhanced frame embedding and an object token, generates an average frame embedding according to the instance aggregation module and uses it as the first input of the relationship aggregation module, and uses the object token as the second input of the relationship aggregation module, the relationship aggregation module outputs an enhanced object token and an enhanced frame token, the enhanced object token is used as the input of the instance-level anomaly detection head, the instance-level anomaly detection head outputs an instance-level detection result (i.e., an instance-level anomaly score), concatenates the frame token and the enhanced frame token and uses them together with the previous state output by the long short-term memory network module as the input of the long short-term memory network module, and the long short-term memory network module is connected to the frame-level anomaly detection head, and the frame-level anomaly detection head outputs a frame-level detection result (i.e., a frame-level anomaly score).

[0041] As Figure 2 shown, input consecutive frames of a traffic video sequence, N F ×H F ×W F ×3, where N F , H F , W F, 3 represent the number of video frames, length, width, and RGB channels respectively. After passing through the backbone network (VST), features at four scales are obtained. where s ∈ [8, 16, 32, 32], D E is the feature dimension. Next, the multi-scale features pass through the Feature Pyramid Network (FPN). First, average pooling is applied to unify the time dimension to obtain features. Then, for each higher-level feature, a Multi-Layer Perceptron (MLP) is used to adjust the channel dimension, and the spatial dimension is aligned through upsampling to match the next-level feature. The aligned feature maps are merged through element-wise addition to obtain the frame embedding. Finally, the frame embedding passes through a reducer (Reducer) composed of average pooling to obtain the frame token. This process is represented as follows:

[0042] f t = FPN(VST(c t ))

[0043] f' t = Reducer(f t )

[0044] where c t represents the t-th video segment, is the frame embedding, is the frame token.

[0045] The frame-level feature extraction module includes a backbone network, a Feature Pyramid Network, and a reducer connected in sequence. Each traffic video segment serves as the input to the backbone network, where the output of the Feature Pyramid Network is denoted as the frame embedding, and the output of the reducer is denoted as the frame token.

[0046] The prompt token includes two position tokens and a learnable token. The two position tokens are obtained by performing position encoding on the upper-left coordinate and the lower-right coordinate of the traffic object bounding box respectively. Traffic objects include various vehicles, pedestrians, etc.

[0047] The last frame of the video segment obtains the traffic object bounding box through the object detector. Each bounding box includes the upper-left and lower-right coordinates, and then the prompt token is obtained through the object prompt encoder, O i = i represents the i-th object in the current frame, V i represents the learnable token, V i tl , V i br represent the tokens obtained by position encoding of the upper-left and lower-right coordinates respectively. If there is no object in the current frame, then another learnable token will be used to replace V i tl , Vi br 。

[0048] As Figure 3 shown, the instance aggregation module includes L I cascaded instance-level blocks in sequence. Each instance-level block includes a multi-head self-attention layer (MHSA), a first multi-head cross-attention layer (MHCA), a multi-layer perceptron (MLP), and a second multi-head cross-attention layer (MHCA) connected in sequence. The prompt token is input to the multi-head self-attention layer of the first instance-level block in the form of a key, the frame embedding is input to the first multi-head cross-attention layer of the first instance-level block in the form of a query, the output of the second multi-head cross-attention layer of the previous instance-level block is input to the first multi-head cross-attention layer of the next instance-level block, the output of the multi-layer perceptron of the previous instance-level block is input to the multi-head self-attention layer of the next instance-level block, the output of the multi-layer perceptron of the last instance-level block is denoted as the object token, and the output of the second multi-head cross-attention layer of the last instance-level block is denoted as the instance-enhanced frame embedding. In the l-th instance-level block, the whole process is represented as:

[0049]

[0050]

[0051] where LN(·) represents layer normalization, W q ,W k ,W b are the three weight matrices of the multi-head attention (MHA) module, represent the prompt token and the frame embedding of the l-th instance-level block respectively. The prompt token and the frame embedding of the first instance-level block come from the output f t of the feature pyramid network and the output O i of the object prompt encoder respectively. Finally, the L I -th instance-level block outputs the final instance-enhanced frame embedding and the object token

[0052] As Figure 4 shown, the relationship aggregation module includes L R cascaded relationship-level blocks in sequence. Each relationship-level block includes a first Cross-Transformer layer, a second Cross-Transformer layer, and a third Cross-Transformer layer connected in sequence; N KThe learnable query tokens and the object tokens output by the instance aggregation module are used as the inputs to the first Cross-Transformer layer of the first relational-level block, where the query tokens serve as the query and the object tokens serve as the key-value. The object tokens are used as the inputs to the first Cross-Transformer layer and the third Cross-Transformer layer, with the object tokens serving as the query, and the average frame embedding is used as the input to the second Cross-Transformer layer of each relational-level block. The output of the second Cross-Transformer layer of the previous relational-level block is used as the input to the first Cross-Transformer layer of the next relational-level block, and the output of the third Cross-Transformer layer of the previous relational-level block is used as the input to the first Cross-Transformer layer and the third Cross-Transformer layer of the next relational-level block; the output of the second Cross-Transformer layer of the last relational-level block is denoted as the enhanced frame token, and the output of the third Cross-Transformer layer of the last relational-level block is denoted as the enhanced object token. Among them, the Cross-Transformer layer contains multi-head cross-attention (MHCA) and a feed-forward network module (FFN), and the feed-forward network module contains two linear layers and the activation function ReLU between the two layers. The expression is as follows:

[0053]

[0054] Among them, LN(·) represents layer normalization; W q ,W k ,W v are the three weight matrices of the multi-head attention (MHA) module; x is the input; is the output of MHCA, which can be denoted as the intermediate value generated by the Cross-Transformer.

[0055] In the l-th relational-level block, the whole process is expressed as:

[0056]

[0057]

[0058] Among them, The query tokens, object tokens, and average frame embedding tokens of the l-th relational-level block. The query tokens and object tokens of the first relational-level block come from the learnable query tokens initialized separately and the output of the instance aggregation module and the average frame embedding tokens come from the output of the instance aggregation module Take the mean. Finally, the LR The relational level block obtains the final enhanced object token and the enhanced frame token

[0059] Concatenate the frame token and the enhanced frame token and input them into the long short-term memory network module and the frame-level anomaly detection head in sequence to obtain the frame-level anomaly score; the instance-level anomaly detection head includes an instance-level regression layer. During inference and prediction, only the enhanced object token is input into the instance-level anomaly detection head to obtain the anomaly score for each instance. This process is expressed as:

[0060]

[0061] where, f′ t is the frame embedding that passes through a reducer composed of average pooling to obtain the frame token, f″ t is the final feature obtained through LSTM, which is used for regression to obtain the anomaly score, [·||·] represents the concatenation operation, LSTM represents the long short-term memory network module, Regressor F represents the frame-level anomaly detection head, Regressor O represents the instance-level anomaly detection head. h t-1 is the previous state of the LSTM, is the image f t 's frame-level anomaly score, s represents the set of anomaly scores of all objects in the image f t

[0062] During the training process of the neural network model based on the enhanced object prompt, its loss function is a loss function based on cross-entropy, which specifically includes frame-level loss and instance-level loss; among them, the instance-level anomaly detection head performs bipartite graph matching based on the Hungarian algorithm on the traffic object bounding box coordinates and the input label box during training to obtain the corresponding matching result; the instance-level loss is calculated based on the matching result and the enhanced object token. The specific formula is as follows:

[0063]

[0064] where, is the value of the loss function based on cross-entropy, is the frame-level loss value, is the instance-level loss, y t is the anomaly annotation of the f t frame image, y t ∈{0, 1}, is the frame-level anomaly score of the image f t , is the set of true bounding boxes of the abnormal objects in the image f t , ​Given the set of all traffic object bounding boxes obtained by the target detector, the predicted bounding boxes are divided into two groups using the Hungarian algorithm: those that match the ground truth and those that do not. Use to represent, σ i ∈ {0, 1}, σ i represents the matching coefficient, σ i = 1 indicates that the predicted bounding box has been matched, σ i = 0 indicates that the bounding box is not matched, N O is the number of detected traffic objects, α is a scaling factor used to balance positive and negative samples, s i is the anomaly score of the i-th object in the image f t .

[0065] 3) Input the traffic video data to be detected into the traffic video anomaly detection network model, and the model outputs the traffic video anomaly detection result.

[0066] To verify the effectiveness of the present invention, the present invention is verified on the largest publicly available in-vehicle traffic video anomaly detection datasets DoTA and DADA-2000, and compared with the current state-of-the-art traffic anomaly detection methods. The DoTA dataset contains 4,677 videos with a resolution of 1280×720 pixels and is accompanied by time, space, and category annotations. To further evaluate the performance in scenarios involving small objects, the present invention constructs a dedicated subset by selecting scenarios where the anomaly target bounding box is less than 2,267 pixels, which corresponds to the size of the 10% smallest targets in the dataset. In this way, 334 videos are selected from the original 1,402 test videos to form the subset. DADA-2000 is a dataset designed for driver attention prediction and is mainly used for driving accident scenarios. It provides time annotations for accident intervals and space annotations based on driver eye movements. The dataset consists of 2,000 video sequences, each with a resolution of 1584×660 pixels, classified into 54 types of anomalies, and further divided into two main categories: self-vehicle involved and self-vehicle not involved.

[0067] The present invention uses the area under the ROC curve (AUC) and the area under the spatio-temporal ROC curve (STAUC) as evaluation metrics. To enable the network to achieve better performance, the parameter details are as follows: The training settings use the SGD optimizer with an initial learning rate of 0.002, a momentum of 0.9, a weight decay of 0.00001, and a batch size of 16. The model is trained for 200 epochs. The input video resolution is set to 640×480, and each video clip contains 4 frames. In addition, the weighting factor α = 0.7, D E = 256, the instance aggregation module contains L I = 4 instance-level blocks, and the relationship aggregation module contains L R= 4 relationship-level blocks. In addition, the number of learnable query tokens is set to N K = 12.

[0068] The experiment mainly includes two parts. The first part is a comparative experiment between the method of the present invention and the most advanced video anomaly detection methods currently available. The second part is a control variable experiment for each module in the network of the present invention to illustrate the effectiveness of each module in the present invention.

[0069] The first part: A comparative experiment with the most advanced video anomaly detection methods currently available to illustrate the superiority of the network performance of the present invention. Table 1 shows the comparison results of the network model of the present invention and other SOTA methods on the DoTA dataset. Note that TTHF uses median filtering and min-max normalization post-processing to rescale the anomaly scores, which makes it unable to be used online. To ensure a fair comparison, the experimental results of the present invention using the same post-processing technique are added, labeled as The results show that the present invention achieves state-of-the-art performance in both online detection and the variant with post-processing. In particular, the present invention significantly improves the STAUC, highlighting its strong precise spatio-temporal localization ability in traffic anomaly detection.

[0070] Table 1 is a comparison table of the metrics of the present invention and the published SOTA methods on DoTA

[0071]

[0072] Table 2 shows the AUC↑(%) of various types of accidents in the DOTA dataset for the network model of the present invention and other SOTA methods. N / A indicates that the AUC performance for the corresponding category is not available. The anomalies marked with * are non-self-vehicle anomalies, while the anomalies without * are self-vehicle involved anomalies. The model marked with uses the same post-processing technique. The best performance in the model without post-processing is shown in bold, while the better performance in the model with post-processing is underlined. Obviously, most methods perform better on self-vehicle involved anomalies, which may be because these anomalies have better visibility and larger anomaly regions. The present invention has achieved significant performance improvement in the online detection of most anomaly categories. In addition, by adding a post-processing step, the present invention shows a particularly strong improvement in the more challenging non-self-vehicle involved categories. This indicates that the present invention is particularly effective in anomaly scenarios involving distant objects or located outside the field of view center.

[0073] Table 2 is the AUC↑(%) of various types of accidents in the DOTA dataset

[0074]

[0075] Table 3 shows the specific meanings of the labels for various types of accidents in Table 2

[0076]

[0077] To evaluate the generalization ability of the present invention on unseen traffic anomaly types, the present invention is only trained on the DoTA dataset and then directly tested on the DADA-2000 dataset without accessing any training data of DADA-2000. The results are shown in Table 4, indicating that the present invention is superior to previous traffic anomaly detection methods, especially achieving significant performance improvement in the categories not involving the ego vehicle.

[0078] Table 4 is a comparison table of the metric AUC↑(%) of the present invention and the published SOTA methods on DADA-2000

[0079]

[0080] Second part: The control variable experiments of each module in the present invention are conducted to illustrate the effectiveness of each module in the present invention. Table 5 shows the experimental results of model variants with different combinations of key components of the present invention on the DoTA dataset, including the FPN module, the prompt mechanism (Prompt), and the instance-level loss (Ins.Loss). To further evaluate the performance improvement of the present invention on small anomaly objects, the model is also evaluated on the subset introduced above. The experiments show that each module contributes to the performance improvement. Especially the improvement on the subset is more significant, indicating that the present invention is particularly effective in dealing with anomalies at a distance or outside the field of view center.

[0081] Table 5 is a comparison table of the metric AUC↑(%) of model variants with different designs

[0082]

[0083]

[0084] In the prompt mechanism, in addition to the instance aggregation module that integrates the information between object instances and the scene, the present invention also utilizes a relation aggregation module to capture the relationships between objects. To evaluate the effectiveness of the relation aggregation module, the present invention explores other design schemes, that is, splicing the frame tokens obtained by processing the average frame embedding through a dimensionality reducer with the original frame tokens and applying the instance-level loss to the output of the instance aggregation module. The experimental results are shown in Table 6, indicating that the relation aggregation module significantly improves the performance, highlighting the importance of the relationships between objects in anomaly detection.

[0085] Table 6 is a comparison table of the metric AUC↑(%) of the relation aggregation module

[0086]

[0087] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a traffic video anomaly detection method based on enhanced object cues are implemented.

[0088] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a traffic video anomaly detection method based on enhanced object cues are implemented.

[0089] The present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of a traffic video anomaly detection method based on enhanced object cues are implemented.

[0090] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention and not to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered by the protection scope of the claims of the present invention.

Claims

1. A traffic video anomaly detection method based on enhanced object prompts, characterized in that: The following steps are involved: 1) Obtain traffic video dataset; 2) Construct a neural network model based on enhanced object prompts, and use the traffic video dataset to train the neural network model based on enhanced object prompts to obtain a traffic video anomaly detection network model; 3) The traffic video data to be detected is input into the traffic video anomaly detection network model, and the model outputs the traffic video anomaly detection result.

2. The method for detecting anomalies in traffic video based on enhanced object prompts according to claim 1, characterized in that: In the above 2), the neural network model based on enhanced object prompts includes a frame-level feature extraction module, a target detector, an object prompt encoder, an instance aggregation module, a relationship aggregation module, a long short-term memory network module, an instance-level anomaly detection head and a frame-level anomaly detection head; each traffic video clip is used as an input of the frame-level feature extraction module, and the frame-level feature extraction module outputs a frame embedding and a frame token, wherein the frame embedding is used as the first input of the instance aggregation module, the last frame of the current traffic video clip is used as the input of the target detector, the target detector outputs the traffic object bounding box coordinates and serves as the input of the object prompt encoder, and the object prompt encoder outputs the prompt token and serves as the instance aggregation module The second input of the block, the instance aggregation module outputs instance enhanced frame embedding and object token, the average frame embedding is generated according to the instance aggregation module and used as the first input of the relationship aggregation module, and the object token is used as the second input of the relationship aggregation module, the relationship aggregation module outputs enhanced object tokens and enhanced frame tokens, the enhanced object tokens are used as the input of the instance-level anomaly detection head, the instance-level anomaly detection head outputs instance-level detection results, the frame tokens and enhanced frame tokens are concatenated together with the previous state output by the long short-term memory network module as the input of the long short-term memory network module, the long short-term memory network module is connected to the frame-level anomaly detection head, and the frame-level anomaly detection head outputs frame-level detection results.

3. The traffic video anomaly detection method based on enhanced object prompting according to claim 2 is characterized in that: The frame-level feature extraction module includes a backbone network, a feature pyramid network and a dimension reducer connected in sequence, wherein the output of the feature pyramid network is recorded as a frame embedding, and the output of the dimension reducer is recorded as a frame token.

4. The method for detecting anomalies in traffic video based on enhanced object prompts according to claim 2, characterized in that: The prompt token includes two position tokens and one learnable token, and the two position tokens are obtained by respectively encoding the upper left coordinate and the lower right coordinate of the traffic object boundary box.

5. The method for detecting anomalies in traffic video based on enhanced object prompts according to claim 2, characterized in that: The instance aggregation module includes L I There are instance-level blocks cascaded in sequence, each instance-level block includes a multi-head self-attention layer, a first multi-head cross-attention layer, a multi-layer perceptron and a second multi-head cross-attention layer connected in sequence; the prompt token is used as the input of the multi-head self-attention layer of the first instance-level block, the frame embedding is used as the input of the first multi-head cross-attention layer of the first instance-level block, the output of the second multi-head cross-attention layer of the previous instance-level block is used as the input of the first multi-head cross-attention layer of the next instance-level block, the output of the multi-layer perceptron of the previous instance-level block is used as the input of the multi-head self-attention layer of the next instance-level block, the output of the multi-layer perceptron of the last instance-level block is recorded as the object token, and the output of the second multi-head cross-attention layer of the last instance-level block is recorded as the instance enhanced frame embedding.

6. The method for detecting anomalies in traffic video based on enhanced object prompts according to claim 2, characterized in that: The relationship aggregation module includes L R N sequentially cascaded relation-level blocks, each of which includes a first Cross-Transformer layer, a second Cross-Transformer layer, and a third Cross-Transformer layer connected in sequence; K The learnable query tokens and the object tokens output by the instance aggregation module are used as the input of the first Cross-Transformer layer of the first relation-level block, the object tokens are used as the input of the first and third Cross-Transformer layers, the average frame embeddings are used as the input of the second Cross-Transformer layer of each relation-level block, the output of the second Cross-Transformer layer of the previous relation-level block is used as the input of the first Cross-Transformer layer of the next relation-level block, and the output of the third Cross-Transformer layer of the previous relation-level block is used as the input of the first and third Cross-Transformer layers of the next relation-level block; the output of the second Cross-Transformer layer of the last relation-level block is recorded as the enhanced frame token, and the output of the third Cross-Transformer layer of the last relation-level block is recorded as the enhanced object token.

7. The method for detecting anomalies in traffic video based on enhanced object prompts according to claim 2, characterized in that: In the above 2), during the training process of the neural network model based on enhanced object prompts, its loss function is a loss function based on cross entropy, which specifically includes frame-level loss and instance-level loss, and the formula is as follows: in, is the loss function value based on cross entropy, is the frame-level loss value, is the instance-level loss, y t f t Anomaly annotation of frame image, y t ∈{0, 1}, For image f t The frame-level anomaly score of For image f t The set of ground-truth bounding boxes of the anomaly objects, Get the bounding box set of all traffic objects for the target detector, σ i represents the matching coefficient, σ i = 1 means the predicted bounding box has been matched, σ i = 0 means the bounding box is not matched, N O is the number of detected traffic objects, α is the scaling factor used to balance positive and negative samples, and s i is the image f t The anomaly score of the i-th object in .

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a method for detecting anomalies in traffic videos based on enhanced object prompts as described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for detecting anomalies in traffic videos based on enhanced object prompts as described in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of a method for detecting anomalies in traffic videos based on enhanced object prompts as described in any one of claims 1 to 7 are implemented.