Vehicle anomaly detection method and device, storage medium and electronic equipment
By segmenting and feature extraction of subway videos, combining selective structured state space model and full-connection layer network, real-time detection of vehicle operating status and accurate identification of abnormal events are achieved, and the problem of relying on manual judgment and optical signals in the existing technology is solved, and the accuracy and generalization ability of monitoring are improved.
Patent Information
- Application Number
- CN202510087640.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-27
AI Technical Summary
The existing subway abnormality monitoring technology relies on optical signals and manual judgments, which are costly and poorly generalized, making it difficult to accurately identify abnormal vehicle operation status and early warning potential problems.
The vehicle abnormality detection method is adopted to obtain the acquired video, perform video segmentation and feature extraction, and feature aggregation and mapping of the video clip features using a selective structured state space model and a fully connected layer network to generate an output feature sequence, and obtain a video score through the fully connected layer network to determine whether there is an abnormal score in the video clip.
Real-time detection of vehicle operating status and accurate identification of abnormal events are achieved, reducing the cost of manual judgment and improving the generalization ability and accuracy of monitoring.
Smart Images

Figure CN120047905A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of rail transit control technology, and in particular, to a vehicle anomaly detection method, a vehicle anomaly detection device, a storage medium, and an electronic device. Background Art
[0002] With the rapid development of urban public transportation, the subway, as a major means of transportation in large cities, its safe operation has received increasing attention. The mainstream subway anomaly monitoring is based on optical signals, which requires a large number of preset anomaly features to mine anomaly information through matching. It mostly relies on manual judgment, which has high labor costs and poor generalization, and it is difficult to accurately identify the abnormal operating state of vehicles and warn of potential problems. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure are expected to provide a vehicle anomaly detection method, a vehicle anomaly detection device, a storage medium, and an electronic device.
[0004] The technical solution of the present disclosure is implemented as follows:
[0005] In a first aspect, the present disclosure provides a vehicle anomaly detection method.
[0006] The vehicle anomaly detection method provided by the embodiments of the present disclosure includes:
[0007] Obtain a collected video during the operation of the vehicle to be detected;
[0008] Input the collected video into a video preprocessing module to perform video segmentation processing to obtain a plurality of video segments, and extract the video segment features of each of the video segments; wherein, the video segment features are used to characterize the spatial features and temporal features of the image changes in the video segment;
[0009] Perform feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal relevance;
[0010] Input the output feature sequence into a fully connected layer network to obtain the video scores of each video segment;
[0011] Based on the video scores of each of the video segments, determine whether there are abnormal scores for each of the video segments;
[0012] If there are abnormal scores among the video scores of each of the video segments, determine that there is an abnormal vehicle operation condition within the video segment corresponding to the abnormal score.
[0013] In some embodiments, aggregating the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features includes:
[0014] Aggregating the video segment features of each of the video segments through a selective structured state space model in combination with learnable vectors to obtain aggregated features; wherein, the learnable vectors are used to guide the vectors formed by the video segment features of each of the video segments to learn the temporal characteristics of video actions, so that the aggregated features have the temporal characteristics of the video actions.
[0015] In some embodiments, performing feature mapping on the aggregated features to obtain an output feature sequence includes:
[0016] Performing feature mapping on the aggregated features based on a mapping model to obtain an output feature sequence; wherein, the mapping model is:
[0017]
[0018] wherein,
[0019] are all discretized parameter matrices discretized according to the zeroth-order hold rule with a time step of Δ; f vi is the video segment feature in the aggregated feature of the i-th video segment; h i is the hidden state vector; y vi is the output feature sequence; wherein, 1 ≤ i ≤ M; M is the number of video segments.
[0020] In some embodiments, inputting the output feature sequence into a fully connected layer network to obtain the video scores of each video segment includes:
[0021] Building a fully connected layer network for video score analysis based on three fully connected layers and two tanh activation functions;
[0022] Inputting the output feature sequence into the fully connected layer network to obtain the video scores of each video segment.
[0023] Inputting the output feature sequence into a fully connected layer network and two tanh activation functions to obtain the video scores of each video segment; wherein,
[0024] s i =tanh(tanh(FFNN((FFN(FFN(y vi )))))s i is the video score of each video segment; FFN is the fully connected layer, and tanh is the activation function; yvi is the output feature sequence.
[0025] In some embodiments, determining whether there is an abnormal score for each of the video segments based on the video scores of the respective video segments includes:
[0026] Determine the judgment threshold for the abnormal score;
[0027] If the video score of the video segment is greater than the judgment threshold for the abnormal score, it is determined that the video segment has an abnormal score.
[0028] In some embodiments, before aggregating the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, the method includes:
[0029] Train the selective structured state space model and the fully connected layer to optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer;
[0030] Among them, training the selective structured state space model and the fully connected layer to optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer includes the following steps:
[0031] Obtain a sample video collected during the operation of the vehicle to be detected;
[0032] Divide the sample video into a positive packet video and a negative packet video based on whether there is an abnormal event in the video; wherein, the positive packet video is a sample video with at least one abnormal event, and the negative packet video is a sample video without an abnormal event;
[0033] Input the positive packet video and the negative packet video into a video preprocessing module, perform video segmentation processing to obtain a plurality of video segments, and extract the video segment features of each of the video segments;
[0034] Aggregate the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal relevance;
[0035] Input the output feature sequence into a fully connected layer network to obtain the video scores of each video segment;
[0036] With the goal of the video scores of each video segment satisfying the objective function, adjust the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer to obtain a selective structured state space model and a fully connected layer for vehicle anomaly detection; wherein, the objective function includes:
[0037] Where is the video score of the i-th video segment in the positive bag video, is the video score of the i-th video segment in the negative bag; B a is the positive bag video, B n is the negative bag video.
[0038] In some embodiments, the method includes:
[0039] Discretize the learnable parameters with a time step Δ according to the zeroth-order hold rule to obtain a discretized parameter matrix; wherein, the learnable parameters at least include a first learning parameter A, a second learning parameter B, and a third learning parameter C;
[0040] The discretized parameter matrix at least includes:
[0041] The first discretized parameter matrix The second discretized parameter matrix and the third discretized parameter matrix
[0042] Among them, the first learning parameter A is determined when the mapping model is initialized;
[0043]
[0044]
[0045]
[0046] B = m B (F v )
[0047] C = m C (F v )
[0048] Δ = log(1 + exp(m Δ (F v ) + L Δ ))
[0049] F v is the video segment feature in the aggregated feature.
[0050] Second aspect, the present disclosure provides a vehicle anomaly detection device, including:
[0051] A video acquisition module, configured to acquire an acquisition video during the operation of the vehicle to be detected;
[0052] A video segmentation module, configured to input the acquisition video into a video preprocessing module, perform video segmentation processing to obtain a plurality of video segments, and extract video segment features of each of the video segments; wherein, the video segment features are used to characterize the spatial features and temporal features of image changes in the video segment;
[0053] A feature aggregation module, configured to perform feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain an aggregated feature, and perform feature mapping on the aggregated feature to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal relevance;
[0054] A video score determination module, configured to input the output feature sequence into a fully connected layer network to obtain video scores of each of the video segments;
[0055] Anomaly score determination module, configured to determine whether there is an anomaly score for each of the video segments based on the video scores of each of the video segments;
[0056] Anomaly condition determination module, configured to, if there is an anomaly score among the video scores of each of the video segments, determine that there is a vehicle operation anomaly condition in the video segment corresponding to the anomaly score.
[0057] Third aspect, the present disclosure provides a computer-readable storage medium, on which a vehicle anomaly detection program is stored. When the vehicle anomaly detection program is executed by a processor, the vehicle anomaly detection method described in the first aspect above is implemented.
[0058] Fourth aspect, the present disclosure provides an electronic device, including a memory, a processor, and a vehicle anomaly detection program stored on the memory and executable on the processor. When the processor executes the vehicle anomaly detection program, the vehicle anomaly detection method described in the first aspect above is implemented.
[0059] The vehicle anomaly detection method provided by the embodiments of the present disclosure includes: acquiring a collected video during the operation of a vehicle to be detected; inputting the collected video into a video preprocessing module to perform video segmentation processing to obtain a plurality of video segments, and extracting video segment features of each video segment; wherein, the video segment features are used to characterize the spatial features and temporal features of image changes in the video segment; performing feature aggregation on the video segment features of each video segment through a selective structured state space model to obtain aggregated features, and performing feature mapping on the aggregated features to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal relevance; inputting the output feature sequence into a fully connected layer network to obtain video scores of each video segment; determining whether there are anomaly scores for each video segment based on the video scores of each video segment; if there are anomaly scores among the video scores of each video segment, determining that there is an abnormal vehicle operation condition within the video segment corresponding to the anomaly score. In this application, video segment features are obtained by slicing and feature extraction of the collected video, and then the video segment features are processed through a selective structured state space model and a fully connected layer. Finally, the video scores of each video segment are output. Based on the video scores of each video segment, it is determined whether there are anomaly scores for each video segment; if there are anomaly scores among the video scores of each video segment, it is determined that there is an abnormal vehicle operation condition within the video segment corresponding to the anomaly score. The entire process can be carried out in real time during the operation of the subway. By collecting subway operation images in real time and automatically analyzing video data through deep learning, various abnormal events can be accurately identified and potential safety problems can be warned in time.
[0060] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flowchart of a vehicle anomaly detection method shown according to an exemplary embodiment;
[0062] Figure 2 is a flowchart of a vehicle anomaly detection shown according to an exemplary embodiment;
[0063] Figure 3 is a schematic structural diagram of a vehicle anomaly detection device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure.
[0065] With the rapid development of urban public transportation, the subway, as the main means of transportation in large cities, its safe operation has received increasing attention. The mainstream subway anomaly monitoring is based on optical signals, which requires a large number of preset anomaly features and mines anomaly information through matching. It mostly relies on manual judgment. First, the labor cost is high. Second, the generalization ability is poor, and it is difficult to accurately identify the abnormal operation state of the vehicle and warn of potential problems.
[0066] In view of the above situation, the present disclosure provides a vehicle anomaly detection method. Figure 1 It is a flowchart of a vehicle anomaly detection method shown according to an exemplary embodiment. As Figure 1 shown, the vehicle anomaly detection method includes:
[0067] Step 10: Obtain the collected video during the operation of the vehicle to be detected;
[0068] Step 11: Input the collected video into a video preprocessing module, perform video segmentation processing to obtain a plurality of video segments, and extract the video segment features of each of the video segments; wherein, the video segment features are used to characterize the spatial features and temporal features of the image changes in the video segment;
[0069] Step 12: Aggregate the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal relevance;
[0070] Step 13: Input the output feature sequence into a fully connected layer network to obtain the video scores of each video segment;
[0071] Step 14: Based on the video scores of each of the video segments, determine whether there are abnormal scores in each of the video segments;
[0072] Step 15: If there are abnormal scores in the video scores of each of the video segments, determine that there is an abnormal vehicle operation condition in the video segment corresponding to the abnormal score.
[0073] In this exemplary embodiment, the operation state of the vehicle to be detected during operation can be video-captured by a station yard video capture device. Then, the captured video is sliced to obtain a plurality of video segments.
[0074] In this exemplary embodiment, the selective structured state space model can be pre-trained. The selective structured state space model can perform feature aggregation on the video segment features of each of the video segments to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence. Then, the output feature sequence is input into a fully connected layer network to obtain the video scores of each video segment. Based on the video scores of each video segment, it is determined whether there are abnormal scores in each of the video segments. If there are abnormal scores in the video scores of each video segment, it is determined that there is an abnormal vehicle operation condition in the video segment corresponding to the abnormal score. In this application, video segment features are obtained by slicing and feature extraction of the collected video, and then the video segment features are processed by the selective structured state space model and the fully connected layer. Finally, the video scores of each video segment are output. Based on the video scores of each video segment, it is determined whether there are abnormal scores in each of the video segments. If there are abnormal scores in the video scores of each video segment, it is determined that there is an abnormal vehicle operation condition in the video segment corresponding to the abnormal score. The entire process can be carried out in real time during the subway operation. By collecting subway operation images in real time and automatically analyzing video data through deep learning, various abnormal events can be accurately identified and potential safety problems can be warned in time.
[0075] In some embodiments, the performing feature aggregation on the video segment features of each of the video segments by the selective structured state space model to obtain aggregated features includes:
[0076] The selective structured state space model combines learnable vectors to aggregate the video segment features of each of the video segments to obtain aggregated features; wherein, the learnable vectors are used to guide the vectors formed by the video segment features of each of the video segments to learn the temporal characteristics of video actions, so that the aggregated features have the temporal characteristics of the video actions.
[0077] In this exemplary embodiment, the input video V collected by the vehicle is segmented into M video blocks by a video segmentation module The feature dimension of each block is wherein, H and W represent the height and width of each block, H = W = 256. T represents the number of frames of the video block, T = 8. Next, the video is processed by the I3D method and subsequent feature compression methods through a video block feature extraction module to extract the features of each video block
[0078] After obtaining the video segment features the video segment features can be aggregated by the video packet feature aggregation module in the selective structured state space model The aggregation method is to add a learnable vector At the end of the video features, the aggregated feature F is obtained v = [f v1 , f v2 , …, f vM , f cls .
[0079] In some embodiments, the feature mapping of the aggregated feature to obtain an output feature sequence includes:
[0080] Performing feature mapping on the aggregated feature based on a mapping model to obtain an output feature sequence; wherein, the mapping model is:
[0081]
[0082] Wherein,
[0083] are all discretized parameter matrices after being discretized with a time step Δ according to the zeroth-order hold rule; f vi is the video segment feature in the aggregated feature of the i-th video segment; h i is the hidden state vector; y vi is the output feature sequence; wherein, 1 ≤ i ≤ M; M is the number of video segments.
[0084] In this exemplary embodiment, this hidden state represents a latent state, representing the change manner from the input feature to the output feature. Among them, h i can start from the initialization of h0. Through the above mapping model, the association relationship between f vi and y vi can be established. Among them, the hidden state represents the change of video context anomaly information. Among them, the B matrix realizes the mapping from the input to the state, the C matrix realizes the mapping from the state to the output, and the A matrix realizes the state transition.
[0085] In some embodiments, the inputting the output feature sequence into a fully connected layer network to obtain the video scores of each video segment includes:
[0086] Building a fully connected layer network for video score analysis based on three fully connected layers and two tanh activation functions;
[0087] Inputting the output feature sequence into the fully connected layer network to obtain the video scores of each video segment.
[0088] Inputting the output feature sequence into a fully connected layer network and two tanh activation functions to obtain the video scores of each video segment; wherein,
[0089] s i=tanh(tanh(FFNN((FFN(FFN(y vi )))));s i is the video score of each video segment; FFN is a fully connected layer, and tanh is an activation function; y vi is the output feature sequence.
[0090] In this exemplary embodiment, the output feature sequence y vi is input into three fully connected layers FFN and two tanh activation functions to obtain the video score s of each video segment i ; among them, the three fully connected layers have 256, 64, and 1 unit respectively; among them, s i ∈[0, 1].
[0091] In this application, using three fully connected layers is beneficial to multi-layer feature combination, making the abnormal feature range more complete, and at the same time can increase the generalization ability of the model.
[0092] In some embodiments, determining whether there is an abnormal score for each of the video segments based on the video scores of the video segments includes:
[0093] Determine the judgment threshold for the abnormal score;
[0094] If the video score of the video segment is greater than the judgment threshold for the abnormal score, it is determined that the video segment has an abnormal score.
[0095] In this exemplary embodiment, in the deployment and inference stage of the model, this application first calculates the abnormal probability at the video level to indicate the possibility of abnormal occurrence in a given video. Secondly, this application sets a judgment threshold T to determine the abnormal to be located in the video. This application divides the video into N b equally long segments, where N v is the total length of a video. Further, overlapping segments are deleted. For each segment b i , this application uses a selective structured state space model and 3 fully connected layers to obtain an abnormal score If then this segment is abnormal. In this way, this application obtains the abnormal score and abnormal position of the video, and further obtains the abnormal video blocks and scores in the output video.
[0096] In some embodiments, before obtaining the aggregated features by aggregating the video segment features of each of the video segments through a selective structured state space model, the method includes:
[0097] Train the selective structured state space model and the fully connected layer, and optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer;
[0098] Among them, the training of the selective structured state space model and the fully connected layer to optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer includes the following steps:
[0099] Obtain the sample video collected during the operation of the vehicle to be detected;
[0100] Divide the sample video into positive packet videos and negative packet videos based on whether there are abnormal events in the video; among them, the positive packet video is a sample video with at least one abnormal event, and the negative packet video is a sample video without abnormal events;
[0101] Input the positive packet video and the negative packet video into the video preprocessing module, perform video segmentation processing to obtain multiple video segments, and extract the video segment features of each video segment;
[0102] Aggregate the video segment features of each video segment through the selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; among them, the features in the output feature sequence have temporal relevance;
[0103] Input the output feature sequence into the fully connected layer network to obtain the video scores of each video segment;
[0104] With the goal that the video scores of each video segment satisfy the objective function, adjust the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer to obtain the selective structured state space model and the fully connected layer for vehicle anomaly detection; among them, the objective function includes:
[0105] where is the video score of the i-th video segment in the positive packet video, is the video score of the i-th video segment in the negative packet; B a is the positive packet video, B n is the negative packet video.
[0106] In this exemplary embodiment, after performing video preprocessing on a sample video, the video segment features of the positive packet video and the negative packet video are input into a selective structured state space model and a fully connected layer. With the goal of the video scores of each video segment satisfying the objective function, the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer are adjusted and trained to obtain a selective structured state space model and a fully connected layer for vehicle anomaly detection. Among them,
[0107] where is the video score of the i-th video segment in the positive packet video, is the video score of the i-th video segment in the negative packet; B a is the positive packet video, B n is the negative packet video.
[0108] In this exemplary embodiment, by it is ensured that the video score in the positive packet video and the video score in the negative packet video satisfy that the minimum value of the video anomaly score in the positive packet video is higher than the maximum value of the video anomaly score in the negative packet video.
[0109] In some embodiments, the method includes: discretizing the learnable parameters with a time step Δ according to the zeroth-order hold rule to obtain a discretized parameter matrix; where the learnable parameters at least include a first learning parameter A, a second learning parameter B, and a third learning parameter C;
[0110] The discretized parameter matrix at least includes:
[0111] The first discretized parameter matrix The second discretized parameter matrix and the third discretized parameter matrix
[0112] Among them, the first learning parameter A is determined when initialized by the mapping model;
[0113]
[0114]
[0115]
[0116] B = m B (F v );
[0117] C = m C (F v );
[0118] Δ = log(1 + exp(m Δ (F v )) + L Δ ));
[0119] F v is the video segment feature in the aggregated features.
[0121] In this exemplary embodiment, A, B, C, and △ are all learnable parameters. Among them, B, C, and △ are the input features F v obtained through the learnable linear mappings m B , m c , m △ .
[0122] m B (·) = Linear(·);
[0123] m c (·) = Linear(·);
[0124] m c (·) = Broadcast(·);
[0125] Among them, L △ represents the regularization term of △.
[0126] This application is for vehicle anomaly detection based on the videos collected by the station cameras. The video block features are aggregated through the Selective Structured State Space Model and combined with the methods of the Selective Structured State Space Model and multi-instance learning to analyze whether there are anomaly scores in each video segment; and based on the anomaly scores existing in the video scores of each video segment, determine the abnormal vehicle running conditions existing within the video segment corresponding to the anomaly scores to accurately identify various abnormal events.
[0127] Figure 2 is the vehicle anomaly detection flowchart shown according to an exemplary embodiment. As Figure 2 shown, the vehicle anomaly detection process includes:
[0128] Step 20: Obtain positive packet videos;
[0129] Step 21: Obtain negative packet videos;
[0130] Step 22: The video preprocessing model preprocesses the positive packet videos and negative packet videos;
[0131] Step 23: The Selective Structured State Space Model processes the video segment features to obtain an output feature sequence;
[0132] Step 24: The fully connected layer FFN (with 256 neurons) processes the output feature sequence;
[0133] Step 25: The fully connected layer FFN (with 64 neurons) processes the output feature sequence.
[0134] Step 26: The fully connected layer FFN (with 1 neuron) processes the output feature sequence.
[0135] Step 27: Obtain the anomaly score in the positive bag.
[0136] Step 28: Obtain the anomaly score in the negative bag.
[0137] Step 29: Perform ranking loss based on the anomaly scores in the positive bag and the negative bag to amplify the difference between normal videos and abnormal videos.
[0138] This application is adapted to various types of video acquisition devices. This application will not cause the inability to perceive features due to the capacity limitations of video acquisition devices; only a small amount of expert knowledge is required, and the training data only needs to label whether the video level contains anomalies, without accurately marking the time and space positions of anomalies in the video; while better utilizing the selective structured state space long-context perception ability, the spatio-temporal attributes of video features are retained, further supporting weakly supervised anomaly detection; using the ability of deep learning to predict abnormal video blocks and anomaly scores provides basic hints for subsequent inspections and repairs, reducing the possibility of significant losses caused by sudden problems.
[0139] The present disclosure provides a vehicle anomaly detection device. Figure 3 It is a schematic structural diagram of a vehicle anomaly detection device shown according to an exemplary embodiment. As Figure 3 shown, the vehicle anomaly detection device includes:
[0140] A video acquisition module 30, configured to acquire an acquisition video during the operation of the vehicle to be detected;
[0141] A video segmentation module 31, configured to input the acquisition video into a video preprocessing module, perform video segmentation processing to obtain a plurality of video segments, and extract video segment features of each of the video segments; wherein, the video segment features are used to characterize the spatial features and temporal features of image changes in the video segment;
[0142] A feature aggregation module 32, configured to perform feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; wherein, the features in the output feature sequence have temporal correlation;
[0143] A video score determination module 33, configured to input the output feature sequence into a fully connected layer network to obtain video scores of each video segment;
[0144] Anomaly score determination module 34, configured to determine whether there is an anomaly score for each of the video segments based on the video scores of the video segments.
[0145] Anomaly condition determination module 35, configured to determine that there is a vehicle operation anomaly condition in the video segment corresponding to the anomaly score if there is an anomaly score among the video scores of the video segments.
[0146] In this exemplary embodiment, a station yard video acquisition device can be used to collect videos of the running state of the vehicle to be detected during the running process. Then, the collected video is sliced into multiple video segments.
[0147] In this exemplary embodiment, the selective structured state space model can be pre-trained. The selective structured state space model can perform feature aggregation on the video segment features of each of the video segments to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence. Then, the output feature sequence is input into a fully connected layer network to obtain the video scores of each video segment. Based on the video scores of each video segment, it is determined whether there is an anomaly score for each video segment; if there is an anomaly score among the video scores of each video segment, it is determined that there is a vehicle operation anomaly condition in the video segment corresponding to the anomaly score. In this application, video segment features are obtained by slicing and feature extraction of the collected video, and then the video segment features are processed by the selective structured state space model and the fully connected layer. Finally, the video scores of each video segment are output. Based on the video scores of each video segment, it is determined whether there is an anomaly score for each video segment; if there is an anomaly score among the video scores of each video segment, it is determined that there is a vehicle operation anomaly condition in the video segment corresponding to the anomaly score. The entire process can be carried out in real time during the subway operation process. By collecting subway operation images in real time and automatically analyzing video data through deep learning, various abnormal events can be accurately identified and potential safety problems can be warned in time.
[0148] In some embodiments, the feature aggregation module 32 is configured to
[0149] perform aggregation on the video segment features of each of the video segments through a selective structured state space model in combination with learnable vectors to obtain aggregated features; wherein, the learnable vectors are used to guide the vectors formed by the video segment features of each of the video segments to learn the temporal characteristics of video actions, so that the aggregated features have the temporal characteristics of the video actions.
[0150] In this exemplary embodiment, after obtaining the video segment features the video segment features can be aggregated by the video packet feature aggregation module in the selective structured state space model Perform aggregation, and the aggregation method is to add a learnable vector At the end of the video features, the aggregated feature F is obtained v = [f v1 , f v2 , …, f vM , f cls .
[0151] In some embodiments, the feature aggregation module 32 is used to
[0152] Perform feature mapping on the aggregated feature based on a mapping model to obtain an output feature sequence; wherein, the mapping model is:
[0153]
[0154] Wherein,
[0155] are all discretized parameter matrices after being discretized according to the zeroth-order hold rule with a time step Δ; f vi is the video segment feature in the aggregated feature; h i is the hidden state vector; y vi is the output feature sequence; wherein, 1 ≤ i ≤ M; M is the number of video segments.
[0156] In this exemplary embodiment, this hidden state represents a latent state, indicating the change manner from the input feature to the output feature. Among them, h i can be initialized starting from h0. Through the above mapping model, the association relationship between f vi and y vi can be established. Among them, the hidden state represents the change of video context anomaly information. Among them, the B matrix realizes the mapping from the input to the state, the C matrix realizes the mapping from the state to the output, and the A matrix realizes the state transition.
[0157] In some embodiments, the video score determination module 33 is used to
[0158] Input the output feature sequence into a fully connected layer network and two tanh activation functions to obtain the video scores of each video segment; wherein,
[0159] s i = tanh(tanh(FFN(FFN(FFN(y vi )))))); s i is the video score of each video segment; FFN is the fully connected layer, tanh is the activation function; y vi is the output feature sequence.
[0160] In this exemplary embodiment, the output feature sequence y vi is input into three fully connected layers FFN and two tanh activation functions to obtain the video score s of each video segment i ; among them, the three fully connected layers have 256, 64, and 1 unit respectively; among them, s i ∈[0, 1].
[0161] In some embodiments, the abnormal score determination module 34 is used to
[0162] determine the judgment threshold of the abnormal score;
[0163] If the video score of the video segment is greater than the judgment threshold of the abnormal score, it is determined that the video segment has an abnormal score.
[0164] In this exemplary embodiment, in the deployment and inference stage of the model, the present application first calculates the abnormal probability at the video level to indicate the possibility of abnormal occurrence in a given video. Secondly, the present application sets a judgment threshold t to determine the abnormality to be located in the video. The present application divides the video into N b equally long segments, where N v is the total length of a video. The overlapping segments are further deleted. For each segment b i , the present application uses the selective structured state space model and three fully connected layers to obtain the abnormal score If then this segment has an abnormality. In this way, the present application obtains the abnormal score and abnormal position of the video, and further obtains the abnormal video blocks and scores in the output video.
[0165] In some embodiments, the device includes a model training model; before aggregating the video segment features of each of the video segments through the selective structured state space model to obtain the aggregated features, the model training model is used to
[0166] train the selective structured state space model and the fully connected layers, and optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layers;
[0167] Among them, training the selective structured state space model and the fully connected layers and optimizing the model parameters of the selective structured state space model and the matrix parameters of the fully connected layers includes the following steps:
[0168] Obtain the sample videos collected during the operation of the vehicle to be detected;
[0169] Divide the sample video into positive-pack videos and negative-pack videos based on whether there are abnormal events in the video; among them, the positive-pack video is a sample video with at least one abnormal event, and the negative-pack video is a sample video without abnormal events;
[0170] Input the positive-pack videos and negative-pack videos into a video preprocessing module, perform video segmentation processing to obtain multiple video segments, and extract the video segment features of each of the video segments;
[0171] Perform feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and perform feature mapping on the aggregated features to obtain an output feature sequence; among them, the features in the output feature sequence have temporal relevance;
[0172] Input the output feature sequence into a fully connected layer network to obtain the video scores of each video segment;
[0173] With the goal that the video scores of each video segment satisfy the objective function, adjust the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer to obtain a selective structured state space model and a fully connected layer for vehicle anomaly detection; among them,
[0174]
[0175] Among them is the video score of the i-th video segment in the positive-pack video, is the video score of the i-th video segment in the negative-pack; B a is the positive-pack video, B n is the negative-pack video.
[0176] In this exemplary embodiment, perform video preprocessing on the sample video, and then input the video segment features of the positive-pack videos and negative-pack videos into the selective structured state space model and the fully connected layer. With the goal that the video scores of each video segment satisfy the objective function, adjust and train the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer to obtain a selective structured state space model and a fully connected layer for vehicle anomaly detection. Among them,
[0177] Among them is the video score of the i-th video segment in the positive-pack video, is the video score of the i-th video segment in the negative-pack; B a is the positive-pack video, B n is the negative-pack video.
[0178] In this exemplary embodiment, through Enable the video scores in the positive package videos and the video scores in the negative package videos to satisfy that the minimum value of the video anomaly score in the positive package videos is higher than the maximum value of the video anomaly score in the negative package videos.
[0179] The present disclosure provides a computer-readable storage medium, on which a vehicle anomaly detection program is stored. When the vehicle anomaly detection program is executed by a processor, the vehicle anomaly detection method described in each of the above embodiments is implemented.
[0180] The present disclosure provides an electronic device, including a memory, a processor, and a vehicle anomaly detection program stored on the memory and executable on the processor. When the processor executes the vehicle anomaly detection program, the vehicle anomaly detection method described in each of the above embodiments is implemented.
[0181] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can determine and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0182] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0183] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0184] In the description of the present disclosure, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present disclosure.
[0185] In addition, the terms "first", "second", etc. used in the embodiments of the present disclosure are only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the technical features indicated in this embodiment. Thus, the features defined with the terms "first", "second", etc. in the embodiments of the present disclosure can explicitly or implicitly indicate that at least one such feature is included in this embodiment. In the description of the present disclosure, the meaning of the word "plurality" is at least two or more, such as two, three, four, etc., unless otherwise specifically defined in the embodiment.
[0186] In this disclosure, unless otherwise clearly specified or limited in the embodiments, terms such as "installed", "connected", "coupled", and "fixed" in the embodiments shall be understood in a broad sense. For example, the connection can be a fixed connection, a detachable connection, or integrated. Understandably, it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be directly connected, or indirectly connected through an intermediate medium, or it can be the communication inside two components, or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in this disclosure can be understood according to specific implementation situations.
[0187] In this disclosure, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over", and "on top of" the second feature can be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath", and "underneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0188] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A vehicle abnormality detection method, characterized in that: include: Obtain the collected video during the operation of the vehicle to be tested; Input the captured video into a video preprocessing module, perform video segmentation processing to obtain multiple video segments, and extract video segment features of each of the video segments; wherein the video segment features are used to characterize the spatial features and temporal features of image changes in the video segments; Performing feature aggregation on the video clip features of each of the video clips through a selective structured state space model to obtain aggregated features, and performing feature mapping on the aggregated features to obtain an output feature sequence; wherein the features in the output feature sequence have temporal correlation; Input the output feature sequence into the fully connected layer network to obtain the video score of each video clip; Based on the video scores of the video clips, determining whether the video clips have abnormal scores; If there is an abnormal score in the video scores of the video clips, it is determined that an abnormal vehicle operation condition exists in the video clip corresponding to the abnormal score.
2. The vehicle abnormality detection method according to claim 1, characterized in that: The step of performing feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features includes: Through a selective structured state space model combined with a learnable vector, the video clip features of each of the video clips are aggregated to obtain an aggregate feature; wherein the learnable vector is used to guide the vector composed of the video clip features of each of the video clips to learn the timing characteristics of the video action, so that the aggregate feature has the timing characteristics of the video action.
3. The vehicle abnormality detection method according to claim 1, characterized in that: The performing feature mapping on the aggregated features to obtain an output feature sequence includes: The aggregated features are feature mapped based on a mapping model to obtain an output feature sequence; wherein the mapping model is: in, are all discretized parameter matrices after discretization with a time step length Δ according to the zeroth-order hold rule; f vi is the video segment feature in the aggregated feature of the i-th video segment; h i is the hidden state vector; y vi is the output feature sequence, where 1≤i≤M; M is the number of video clips.
4. The vehicle abnormality detection method according to claim 1, characterized in that: The output feature sequence is input into the fully connected layer network to obtain the video score of each video clip, including: A fully connected layer network for video score analysis is built based on three fully connected layers and two tanh activation functions; The output feature sequence is input into the fully connected layer network to obtain the video score of each video clip. The output feature sequence is input into the fully connected layer network and two tanh activation functions to obtain the video score of each video clip; s i =tanh(tanh(FFN(FFN(FFN(y vi )))));s i is the video score of each video clip; FFN is the fully connected layer, tanh is the activation function; y vi is the output feature sequence.
5. The vehicle abnormality detection method according to claim 1, characterized in that: The determining, based on the video scores of the video clips, whether the video clips have abnormal scores includes: Determine the judgment threshold of abnormal score; If the video score of the video clip is greater than the judgment threshold of the abnormal score, it is determined that the video clip has an abnormal score.
6. The vehicle abnormality detection method according to claim 1, characterized in that: Before performing feature aggregation on the video segment features of each of the video segments through the selective structured state space model to obtain aggregated features, the method includes: Training the selective structured state space model and the fully connected layer to optimize model parameters of the selective structured state space model and matrix parameters of the fully connected layer; The step of training the selective structured state space model and the fully connected layer to optimize the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer comprises the following steps: Obtain sample videos collected during the operation of the vehicle to be tested; The sample videos are divided into positive packet videos and negative packet videos based on whether there are abnormal events in the videos; wherein the positive packet videos are sample videos with at least one abnormal event, and the negative packet videos are sample videos without abnormal events; Inputting the positive packet video and the negative packet video into a video preprocessing module, performing video segmentation processing to obtain a plurality of video segments, and extracting video segment features of each of the video segments; Performing feature aggregation on the video clip features of each of the video clips through a selective structured state space model to obtain aggregated features, and performing feature mapping on the aggregated features to obtain an output feature sequence; wherein the features in the output feature sequence have temporal correlation; Input the output feature sequence into the fully connected layer network to obtain the video score of each video clip; With the goal that the video score of each video clip satisfies the objective function, the model parameters of the selective structured state space model and the matrix parameters of the fully connected layer are adjusted to obtain the selective structured state space model and the fully connected layer for vehicle anomaly detection; wherein the objective function includes: in is the video score of the i-th video clip in the positive packet video, is the video score of the i-th video clip in the negative bag; B a For the positive package video, B n For negative package video.
7. The vehicle abnormality detection method according to claim 3, characterized in that: The method comprises: Discretize the learnable parameters with a time step Δ according to a zeroth-order hold rule to obtain a discretized parameter matrix; wherein the learnable parameters at least include a first learning parameter A, a second learning parameter B, and a third learning parameter C; The discretization parameter matrix at least includes: The first discretization parameter matrix The second discretization parameter matrix And the third discretization parameter matrix Wherein, the first learning parameter A is determined when the mapping model is initialized; B=m B (F v ); C=m C (F v ); Δ=log(1+exp(m Δ (F v )+L Δ )); F v is the video segment feature in the aggregated feature.
8. A vehicle abnormality detection device, characterized in that: include: A video acquisition module is used to obtain the collected video during the operation of the vehicle to be detected; A video segmentation module, used to input the collected video into a video preprocessing module, perform video segmentation processing to obtain multiple video segments, and extract video segment features of each of the video segments; wherein the video segment features are used to characterize the spatial features and temporal features of image changes in the video segments; A feature aggregation module, used for performing feature aggregation on the video segment features of each of the video segments through a selective structured state space model to obtain aggregated features, and performing feature mapping on the aggregated features to obtain an output feature sequence; wherein the features in the output feature sequence have temporal correlation; A video score determination module is used to input the output feature sequence into the fully connected layer network to obtain the video score of each video clip; an abnormal score determination module, configured to determine whether each video segment has an abnormal score based on the video score of each video segment; The abnormal condition determination module is used to determine that an abnormal vehicle operation condition exists in the video segment corresponding to the abnormal score if there is an abnormal score in the video scores of the video segments.
9. A computer-readable storage medium, characterized in that: A vehicle abnormality detection program is stored thereon, and when the vehicle abnormality detection program is executed by a processor, the vehicle abnormality detection method described in any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The invention comprises a memory, a processor and a vehicle abnormality detection program stored in the memory and executable on the processor. When the processor executes the vehicle abnormality detection program, the vehicle abnormality detection method described in any one of claims 1 to 7 is implemented.