Method, device and vehicle for identifying abnormal events of networked vehicles

By generating a CAN signal matrix and aggregating it into an event sequence, and combining it with a pre-trained contrastive learning model, the problem of lag in the identification of abnormal connected vehicles is solved, and timely and robust identification of abnormal vehicle events is achieved.

CN122493660APending Publication Date: 2026-07-31AVITA INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AVITA INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2026-05-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing connected vehicle anomaly identification solutions rely on manually preset rules, which cannot identify abnormal events in a timely manner and have serious delays, resulting in damage to brand reputation.

Method used

By generating a CAN signal matrix based on a preset sampling frequency, aggregating it into an event sequence, and extracting a set of representation vectors using a pre-trained contrastive learning model, unsupervised abnormal event recognition is achieved, breaking the dependence on known abnormal scenarios.

Benefits of technology

It enables proactive and forward-looking identification of vehicle anomalies, avoiding brand reputation damage caused by delayed discovery of anomalies and enhancing the robustness and adaptability of the model in real-world connected vehicle scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493660A_ABST
    Figure CN122493660A_ABST
Patent Text Reader

Abstract

This invention relates to the field of vehicle safety technology and discloses a method, device, and vehicle for identifying abnormal events in connected vehicles. The method includes: acquiring target CAN signals of vehicle operating state segments based on a preset sampling frequency, and generating a CAN signal matrix corresponding to the target CAN signals based on the average frame count of the state segments; determining events in the CAN signal matrix and aggregating the events into an event sequence based on their duration; extracting a set of representation vectors corresponding to the event sequences based on a pre-trained contrastive learning model; and determining that an abnormal event exists in the vehicle if a subset of the representation vector set shows an abnormal detection result. Applying the technical solution of this invention enables proactive and pre-emptive identification of abnormal events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle safety technology, specifically to a method, device, and vehicle for identifying abnormal events in connected vehicles. Background Technology

[0002] Current connected vehicle anomaly detection solutions rely on manually preset threshold alarm rules for a single CAN (Controller Area Network) signal. An alarm is triggered when the vehicle's real-time operating signal matches the preset rules. However, this fixed alarm rule-based detection scheme only covers simple anomaly scenarios triggered by a single, predefined signal. The supplementation and iteration of these rules depend entirely on user complaints or post-event manual data review, failing to provide timely identification when an anomaly first occurs, resulting in a significant lag in anomaly detection. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention provide a method, device and vehicle for identifying abnormal events of connected vehicles, which solves the problem of lag in abnormal identification in the prior art.

[0004] According to one aspect of the present invention, a method for identifying abnormal events in connected vehicles is provided, the method comprising: The target CAN signal of the vehicle operating status segment is acquired based on a preset sampling frequency, and a CAN signal matrix corresponding to the target CAN signal is generated based on the average number of frames of the status segment. The events in the CAN signal matrix are identified, and the events are aggregated into an event sequence based on the duration of the events. The set of representation vectors corresponding to the event sequence is extracted based on a pre-trained contrastive learning model; If a subset of the representation vector set shows an abnormal detection result, it is determined that the vehicle has an abnormal event.

[0005] According to another aspect of the present invention, a device for identifying abnormal events in connected vehicles is provided, comprising: The signal processing module is used to acquire the target CAN signal of a vehicle operating state segment based on a preset sampling frequency, and generate a CAN signal matrix corresponding to the target CAN signal according to the average number of frames of the state segment. The event aggregation module is used to determine the event corresponding to each frame in the CAN signal matrix, and aggregate the events into an event sequence based on the duration of the events. The feature extraction module is used to extract the set of representation vectors corresponding to the event sequence based on a pre-trained contrastive learning model; An anomaly detection module is used to determine that the vehicle has an abnormal event if the detection result of a subset in the representation vector set is abnormal.

[0006] According to another aspect of the present invention, a vehicle is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform the operation of the networked vehicle abnormal event identification method as described above.

[0007] This invention acquires target CAN signals from vehicle operation state segments based on a preset sampling frequency and generates corresponding CAN signal matrices, thereby constructing a standardized vehicle operation data base in real time. Events within the CAN signal matrix are then identified, and based on their duration, they are aggregated into event sequences, transforming the previously fragmented CAN signals into event units that retain complete temporal logic. Subsequently, a pre-trained contrastive learning model extracts the representation vector set corresponding to the event sequence. This unsupervised extraction of the core temporal features of the current vehicle is achieved without relying on manually labeled abnormal samples or predefined anomaly rules, breaking the dependence on known abnormal scenarios. When an anomaly is detected in a subset of the representation vector set, an abnormal event in the vehicle can be identified, eliminating the need to wait for user complaints or public opinion to escalate, thus achieving proactive, pre-emptive identification of vehicle anomalies. Furthermore, it avoids brand reputation damage caused by delayed detection of anomalies, enhancing the robustness and adaptability of the model in real-world connected vehicle scenarios.

[0008] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0009] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a first embodiment of the method for identifying abnormal events in connected vehicles provided by the present invention is shown. Figure 2 A flowchart illustrating a second embodiment of the method for identifying abnormal events in connected vehicles provided by the present invention is shown. Figure 3 This diagram illustrates the changes in the loss variable during model iteration in the second embodiment of the present invention. Figure 4 The diagram illustrates an optional process obtained by combining various embodiments of the present invention. Figure 5 This diagram illustrates the visualization results of principal component analysis performed according to the present invention. Figure 6 A schematic diagram of the structure of a first embodiment of the networked vehicle abnormal event identification device provided by the present invention is shown; Figure 7 A structural schematic diagram of an embodiment of the vehicle provided by the present invention is shown. Detailed Implementation

[0010] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0011] Firstly, Figure 1 A flowchart illustrating a first embodiment of the method for identifying abnormal events in connected vehicles according to the present invention is shown. This method can be executed by the vehicle's local vehicle networking system or by a cloud device connected to the vehicle. Figure 1 As shown, the method includes the following steps: Step S10: Obtain the target CAN signal of the vehicle operating status segment based on the preset sampling frequency, and generate the CAN signal matrix corresponding to the target CAN signal according to the average number of frames of the status segment.

[0012] In this embodiment, CAN signals are specific individual data signals transmitted on the CAN bus, such as vehicle speed, engine speed, braking status, vehicle speed, engine speed, throttle, brakes, and light status—real-time signals that can be collected in real time by various electronic control units (ECUs) of the vehicle and sent to the vehicle networking system or cloud devices. The CAN signal sampling rate pre-configured in the vehicle networking system is measured in frames per second (fps), i.e., the number of CAN signal frames collected per second. CAN signals include signals strongly related to driving safety, power control, and chassis operation, as well as signals that do not affect vehicle operation and safety. The target CAN signal is a key CAN status signal strongly related to the connected vehicle and status monitoring, such as brake pedal status, steering wheel intervention status, ABS (Anti-lock Braking System) activation status, AEB (Autonomous Emergency Braking) activation status, EBD (Electronic Brake-force Distribution) activation status, charging status, door opening status, and steering wheel grip status. Different CAN signals share the same sampling frequency.

[0013] The CAN signal matrix is ​​a two-dimensional data structure consisting of target CAN signals as rows and sampling frames as columns. It is used to store the timing data of CAN signals within different state segments. Its matrix dimension is m×d, where m is the number of target CAN signals and d is the average number of frames in a state segment. A state segment is a continuous time interval with clear business significance during vehicle operation, such as a charging state segment, a vehicle power-on state segment, and a door opening state segment.

[0014] To accurately capture core data from different state segments such as charging, power-on, and door opening, avoiding subsequent processing chaos caused by scattered CAN signals, and providing structured data support for identifying abnormal correlations between braking, steering wheel, and ABS / AEB / EBD during power-on, a CAN signal matrix can be constructed after acquiring the target CAN signal based on a preset sampling frequency. First, the quotient between the duration of the vehicle's operating state segment and the preset sampling rate is determined as the average frame count. Then, a CAN signal matrix is ​​constructed with the number of target CAN signals as rows and the average frame count as columns. Each state segment consists of three elements: state segment number (state_id), state start time (t_start), and state end time (t_end). When the CAN signal sampling frequency is f (frames / s), the number of frames from the state start time to the state end time is... .

[0015] As an optional implementation, taking charging status as an example, the vehicle's local control system automatically generates a unique state_id for the current charging status segment, collects and records the state start time t_start and state end time t_end, and then uses the formula d=(t_end) The average number of frames is calculated using t_start / f. For example, if the sampling frequency f is 10 frames / second and the duration of a state segment is 30 seconds, then d = 30 × 10 = 300 frames. The m target CAN signals collected in real time according to the sampling frequency include charging status, brake pedal, and ABS / EBD status signals. Similarly, after collecting the data corresponding to all state segments, a CAN signal matrix of m × d can be constructed with the m target CAN signals as rows and the current average frame as columns.

[0016] As an alternative implementation, the cloud device assigns a unique state segment number (state_id) to the vehicle's runtime segment. It then locks the start time (t_start) and end time (t_end) of the vehicle's power-on state segment. If the state segment duration (t_start - t_end) is 6 seconds and the sampling frequency (f) is 50 frames per second, the average number of frames (d) is calculated to be 300 frames. During this process, the cloud device can determine the target CAN signal and a sub-segment signal matrix within the CAN signal matrix composed of the target CAN signal from the vehicle's real-time feedback data.

[0017] Understandably, for each event number, a [database name] can be obtained. The CAN signal matrix. Since the start and end times of different states can be different, even exponentially different, the size of the CAN signal matrix will vary at different times. For example... , or Therefore, the CAN signal matrix has d time points over a period of time, and each time point has m signal values. In other words, each time point corresponds to a combination state, that is, the state of each target CAN signal.

[0018] Step S20: Determine the events in the CAN signal matrix and aggregate the events into an event sequence based on the duration of the events.

[0019] In this embodiment, an event refers to a combination of events in which the combined states of all target CAN signals remain unchanged across multiple consecutive frames; that is, an event includes both the combined state of the target CAN signals and its duration. An event sequence is a set of multiple events arranged chronologically. The value range of a signal s in the CAN signal can be defined as a sample space: , Indicates signal A valid value for . Let . for The combined sample space of a CAN signal is the Cartesian product of the sample spaces of each signal: , That is, each element A signal combination state represents The combination of sampled values ​​of a signal at the same time. Therefore, for the signal obtained in step S10... The CAN signal matrix can be converted into a time-ordered array. event Duration The event sequence. For each event sequence element , indicating the state Down, arrive Within the time interval, the event is always satisfied. That is, in a certain vehicle state context, such as driving, charging, or stationary, there exists at least one vehicle that continuously satisfies the event for a certain period of time. .

[0020] When aggregating events into an event sequence, we can first determine the event segment containing the combined states of CAN signals corresponding to each frame in the CAN signal matrix. Then, we merge multiple consecutive event segments with consistent combined states of CAN signals into an event and record the duration of the consecutive frames. Finally, we arrange all events in the event order to obtain multiple event sequences corresponding to all vehicle operating state segments. Specifically, we can label each time point, determine which event E each time point's combined state belongs to, and then package and merge consecutive identical events. We merge consecutive time points belonging to the same event E into a segment, while recording the event type and duration. Finally, when generating the event sequence, the original long string of 100 scattered time points is transformed into an event sequence of "Event 1 (duration 5 seconds) → Event 2 (duration 3 seconds) → Event 3 (duration 2 seconds)...".

[0021] For example, the resulting event sequence can be shown in the following table:

[0022] The event sequence contains two events: the first event is that all status bits remain at 0 for 5 seconds, corresponding to the vehicle being powered on; the second event is that all status bits except the "hands on steering wheel" status bit remain at 0 for 1098 seconds, corresponding to the vehicle being in motion.

[0023] Understandably, the number of columns (frame number d) in the CAN signal matrix for different state segments can span several orders of magnitude; for example, the matrix number of a previous state segment might be 50 columns, while the matrix number of a subsequent state segment might be 180,000 columns. Directly inputting this data into the model leads to problems such as high training / processing difficulty and imbalanced feature extraction. Therefore, it is necessary to aggregate consecutive identical signal combinations into event sequences, thereby transforming sparse time series of unequal lengths into event sequences of equal structure.

[0024] Therefore, by aggregating events, all event segments are structured as signal combinations plus durations, thereby transforming sparse time series of unequal lengths into event sequences of multiple events with equal structures in the form of (signal 1~n, duration), alleviating the difficulty in model processing caused by the order-of-magnitude difference in CAN signal lengths.

[0025] Step S30: Extract the set of representation vectors corresponding to the event sequence based on the pre-trained contrastive learning model.

[0026] In this embodiment, although the event sequences have reduced length differences, they are still non-standardized time-series data. Therefore, they need to be transformed into fixed-length representation vectors through a pre-trained contrastive learning model to further eliminate the influence of remaining length differences and generate a quantifiable feature set for comparison. The contrastive learning model is an unsupervised model obtained by pre-training event sequences of normal state segments. It includes an encoder, decoder, and projector head, and can transform event sequences into fixed-dimensional representation vectors. The representation vector set consists of fixed-length feature vectors from one or more event sequences, with each vector containing the core operational features of the corresponding state segment.

[0027] In this embodiment, the contrastive learning model can uniformly fill the event sequence into a fixed-dimensional tensor, thereby compensating for the differences in length between the event sequences. The data is then input into the encoder of the pre-trained model to extract a 128-dimensional fixed-length representation vector. Subsequently, the vector corresponding to each event is input into the representation vector set. Specifically, the encoder extracts features from the input event sequence, transforming the (signal 1~n, duration) structural information of each event into a preliminary feature representation. The decoder then further processes and reconstructs the features output by the encoder to capture the temporal dependencies and deep semantic information in the event sequence. The projection head is responsible for mapping the features output by the decoder to a fixed-dimensional vector space, ultimately obtaining representation vectors of consistent length.

[0028] It should be noted that the above parameters are for illustrative purposes only and are not intended to limit this application.

[0029] Step S40: If the detection result of a subset in the vector set is abnormal, it is determined that there is an abnormal event in the vehicle.

[0030] In this embodiment, the vector subset represents the vector unit corresponding to a single state_id state segment in the vector set, representing the operational characteristics of a single state segment, without length-related bias. The detection result includes normal and abnormal.

[0031] As an alternative implementation, the Isolation Forest algorithm can be used to detect anomalies based on the representation vectors of different event sequences, dividing the input samples (representation vectors) into normal and abnormal categories. For example, for each subset of representation vectors, the model calculates an anomaly score based on the sparsity and outlierness of the vectors. When the anomaly score of a vector unit exceeds a preset threshold, the detection result of the state segment corresponding to that vector unit is determined to be abnormal. Since each vector unit corresponds to a state segment with a single state_id, by independently detecting each vector unit, it is possible to accurately locate which specific state segment is abnormal, thus providing a clear basis for subsequently identifying abnormal events in connected vehicles.

[0032] As another optional implementation, the similarity between the representation vector and the normal reference vector can also be calculated, and if the similarity meets the preset condition, such as the similarity being greater than a threshold, it is judged as normal; otherwise, it is judged as abnormal.

[0033] Understandably, when an abnormal event occurs in the vehicle, safety alerts can be directly issued through the vehicle network system, such as informing the user of a potential safety hazard via audio. In non-driving states, such as charging, safety alerts can be issued through devices connected to / paired with the vehicle's infotainment system, such as pushing charging anomaly warnings to the user's mobile app, displaying the specific anomaly type and suggested handling methods. Similarly, when identifying safety anomalies through cloud devices, similar warning actions can be taken. Furthermore, for detected anomalies, the system can automatically record relevant event sequence data, including the status and duration of each signal at the time of the anomaly, forming a detailed anomaly event log to help technicians accurately pinpoint the root cause of the problem. Simultaneously, if a safety anomaly is detected in autonomous driving mode, appropriate emergency measures can be taken based on the severity of the anomaly, such as gradually reducing vehicle speed, activating hazard lights, and, ensuring safety, smoothly parking the vehicle on the roadside, while simultaneously sending a distress signal and vehicle location information to the user and relevant safety monitoring platforms.

[0034] This embodiment aggregates CAN signals of unequal length into a unified event sequence, generates a fixed-length representation vector based on this event sequence, eliminates the impact of length deviation on feature extraction and anomaly detection, and encodes different events. In its encoding space, events with similar characteristics are brought closer together, while events with different characteristics are pushed apart. Finally, the detection of abnormal events is achieved by measuring the distance between different events in the encoding space. Thus, anomalies can be identified in a timely manner under normal vehicle usage conditions, realizing the proactive identification of abnormal events.

[0035] Based on any of the above embodiments, in Embodiment 2 of this application, please refer to... Figure 2Before acquiring the target CAN signal of a vehicle operating state segment based on a preset sampling frequency, and generating the CAN signal matrix corresponding to the target CAN signal according to the average number of frames of the state segment, steps S50~S90 are also included: Step S50: Obtain the training CAN signal matrix generated based on the training CAN signals of normal samples, and aggregate the training CAN signal matrix into a training event sequence.

[0036] In this embodiment, normal samples include vehicle operation data without driving safety anomalies, covering normal state segments of different durations and scenarios. The training CAN signal matrix is ​​a structured matrix generated based on the CAN signals of the normal samples, and its structure is consistent with the target CAN signal matrix. The training event sequence is an event sequence transformed from the training CAN signal matrix according to the aggregation rules of S20. The training CAN signals of the normal samples are typically stored in a Hive table. In practical applications, CAN signals generated by the vehicle during normal driving are continuously collected according to preset time intervals or event triggering mechanisms and sent to the backend server via a data transmission link. Subsequently, these raw CAN signal data are cleaned, converted, and formatted before being stored in an orderly manner in a Hive table. The storage structure of the Hive table is usually designed according to the attributes of the CAN signals, such as including fields for signal identifier (ID), timestamp, signal value, and signal length, to facilitate efficient querying and extraction of the training CAN signals later.

[0037] As an optional implementation, X normal CAN signals can be retrieved from the vehicle network data lake via a cloud-based big data platform. Each sample corresponds to a unique state_id, t_start, and t_end. The frame number d is calculated using f (frames per second), generating an m×d training CAN signal matrix. Simultaneously, consecutive identical signal combinations within each training CAN signal matrix are aggregated into events of (signal 1~n, duration) using aggregation rules. Finally, the event segments are arranged chronologically to generate a training event sequence. The number of event segments in each sequence can be controlled within a preset range, such as 2~10, thereby significantly reducing the length difference of the original CAN signal matrix and adapting to the needs of batch model training.

[0038] As an alternative implementation, CAN signals can be retrieved from locally stored lightweight normal samples, and then consecutive identical signals can be aggregated into an event sequence using aggregation rules.

[0039] By constructing training event sequences that are structurally consistent with the actual detection scenario and have controllable length differences, a standardized data source is provided for subsequent sample augmentation, ensuring that the features learned by the model match the real scenario.

[0040] Step S60: Generate at least two augmented samples of the training event sequence.

[0041] In this embodiment, augmented samples refer to variants generated after data augmentation of training samples. They retain core features but have slight differences in details. Common augmentation methods include duration jitter, random masking, and event pruning. Sub-events in the training event sequence are considered as one training sample.

[0042] As an optional implementation, different training samples can be processed through different enhancement methods to obtain at least two enhanced samples. For example, the enhancement methods are duration jitter and random masking. In this case, the duration of the first training event in the training event sequence can be scaled based on a preset scaling ratio to obtain the first enhanced sample. The signal state of at least one CAN signal of the second training event in the training event sequence can be masked to obtain the second enhanced sample.

[0043] As another optional implementation, the same training sample can be processed using different augmentation methods to obtain at least two augmented samples. That is, when the first training event and the second training event are the same, it is equivalent to performing different augmentation processes on the same sample data. In addition, different augmentation conditions can be set, such as different augmentation methods corresponding to different numbers of training samples. The selection of augmentation methods and the setting of conditions can be set according to actual needs.

[0044] As another alternative implementation, two augmented samples can be generated for each training sample in the event sequence. Different training samples can also be processed using different augmentation methods to obtain at least two augmented samples.

[0045] For example, the training event sequence is shown in the table below:

[0046] This training event sequence contains only two events, meaning only two training samples. The duration of the training samples can be randomly adjusted by ±10% using a duration jitter method. For example, the 1098s in the second event can be amplified by a factor of 1.1, resulting in 1098 * 1.1 = 1207.8s, yielding an enhanced sample [0,0,0,0,0,0,0,0,0,0,1,1207.8] to simulate the duration error of CAN signal sampling in a real-world scenario. Simultaneously, the second event can be randomly masked using either a random or fixed masking method. For instance, the "hands on steering wheel" column can be masked, setting its signal value to 0. This transforms the training event representation in the second row from [0,0,0,0,0,0,0,0,0,0,1,1098] to the enhanced sample [0,0,0,0,0,0,0,0,0,0,0,1098]. The duration jitter magnitude is: , Let be a random variable that follows a normal distribution. In a random mask, a mask vector can be generated. , The probability of obedience is The Bernoulli distribution, that is, probability It is 0, and By using a random mask, it can be ensured that at least one item is present. Not equal to 0, that is, if Then in the mask vector Random selection ,make .

[0047] It should be noted that the above enhancement process may cause some samples with insufficient data to become invalid. In actual processing, there are usually enough training samples. In this case, the above process will enhance the model's generalization ability, enabling the model to learn the core invariance of normal features, rather than just memorizing specific parameters.

[0048] Step S70: The encoder extracts the training representation vector of the training event sequence based on the initial alignment model, and the decoder projects the training representation vector to obtain the projection result.

[0049] In this embodiment, all sample data, including augmented samples, need to be converted into computable training representation vectors, and then projected onto the contrastive learning-specific feature space through a decoder to provide a suitable feature form for subsequent loss calculation.

[0050] After all samples are augmented, the sample becomes (number of samples, m, maximum number of events). In this format, the sample is input into the initial alignment model. The sample is processed by an encoder with a tensor dimension of (64, 128) to obtain a 128-dimensional training representation vector. The decoder's projection tensor dimension is... 64,32 This allows the extraction of a 64-dimensional training representation vector, which is then input into the decoder and projected onto a 32-dimensional contrast space to obtain the projection result. The hyperparameters used for sample encoding and decoding are: batch size BATCH_SIZE=64, number of iterations EPOCHS=50, learning rate LR=1e-3, feature dimension (used for loss function calculation) FEAT_DIM=32, and class prior probability (the probability of randomly drawing a sample of the same class from the anchor sample k) CLASS_P=0.5.

[0051] Step S80: Calculate the loss scalar corresponding to the projection result based on the contrast loss function.

[0052] In this embodiment, the loss scalar is the normalized scalar value of the loss vector, which is the core basis for model parameter iteration; that is, the smaller the scalar, the higher the similarity of the positive sample pairs. The loss vector, on the other hand, represents the error vector calculated dimension by dimension of the projection results of different augmented samples, reflecting the matching degree of features in each dimension.

[0053] As an optional implementation, the two augmented samples corresponding to the training event in the projection result can be identified as positive sample pairs, and the augmented samples of other training events as negative sample pairs. Then, the cosine similarity of the projection results of the positive sample pairs is calculated using the debiased NT-Xent loss function, and the similarity of the negative sample pairs is calculated. After calculating the error of the positive sample pairs dimension by dimension, a 64-dimensional loss vector is generated. The loss vector is then normalized to obtain a loss scalar; the smaller the scalar, the more similar the core features of the positive sample pairs are. For example, in a model training, for a training event of "autonomous driving activation state," two augmented samples A and B are generated, which constitute a positive sample pair; simultaneously, 10 augmented samples are randomly selected from other training events as negative sample pairs. First, the cosine similarity of the projection results of the positive sample pairs A and B is calculated, yielding a value of 0.85, indicating that they have high similarity in the 32-dimensional contrast space; the cosine similarity of the negative sample pairs with the "autonomous driving activation state" samples is 0.32, lower than that of the positive sample pairs. Subsequently, the errors of positive sample pairs A and B are calculated dimension by dimension. For example, on the 5th feature dimension, the projected value of A is 0.72, and the projected value of B is 0.68, with an error of 0.04; on the 18th feature dimension, A is 0.29, and B is 0.31, with an error of 0.02, and so on, to obtain a 64-dimensional loss vector. After normalizing this vector using the L2 norm, the loss scalar is obtained as 0.12.

[0054] As another alternative implementation, the Euclidean distance between the positive samples and the projection results can be calculated, and a 32-dimensional loss vector (such as [0.01,0.03,...,0.02]) can be generated dimension by dimension. Finally, the loss scalar is obtained by normalization, which simplifies the calculation logic and improves the speed of vehicle training.

[0055] Step S90: Iterate the encoder and decoder parameters of the initial comparison model through backpropagation based on the loss scalar until the model converges, and obtain the pre-trained contrastive learning model.

[0056] In this embodiment, the parameter gradients of the encoder and decoder can be calculated using the backpropagation algorithm, and the Adam optimizer is used to adjust the parameters and reduce the loss scalar. Then, the process of S60~S80 is repeated to perform batch iteration on the collected training event sequence. When the number of iterations reaches a preset number, the iteration stops and the encoder and decoder parameters are saved to obtain the pre-trained contrastive learning model.

[0057] For example, please refer to Figure 3 The number of iterations was 50, the initial loss variable was 0.8, and after 50 iterations the loss scalar decreased to below 0.1, and the subsequent batches fluctuated little. At this point, the model was determined to have converged, and the pre-trained contrastive learning model was finally obtained.

[0058] This embodiment compares the unsupervised pre-training of the learning model, first constructs a training event sequence with controllable length differences, then uses sample augmentation to enable the model to learn the core invariance of normal features, and finally iterates the parameters to obtain a converged pre-trained model, thereby enabling the pre-trained model to have strong generalization ability and accurately identify abnormal information in all scenarios such as emergency braking, high-speed driving, and charging, realizing the pre-emptive and standardized identification of anomalies.

[0059] Based on the second embodiment described above, in the third embodiment of this application, the calculation of the loss scalar corresponding to the projection results of different samples in the same training event based on the contrastive loss function further includes steps S81 to S82: Step S81: Calculate the expected negative sample distribution of the training samples anchored by the projection result, and calculate the negative sample index based on the expected distribution.

[0060] Step S82: Calculate the loss vector of the training samples based on the negative sample index and the loss vector of the training samples.

[0061] In this embodiment, the observed negative sample can be calculated using the positive sample distribution, negative sample distribution, and the law of total probability of the training samples. Then, the expected distribution of the negative sample can be calculated using the data distribution of the observed negative sample.

[0062] Specifically, for the sample anchored by the projection result Its observed negative samples originate from a mixture of two latent distributions, namely the positive sample distribution. and negative sample distribution The data distribution of observed negative samples can be calculated using the law of total probability: , in, To randomly select a sample from the dataset that is exactly the same as the anchor sample The probability of belonging to the same category. By estimating the distribution of the true negative samples from all observed negative samples, the expected distribution of the true negative samples is obtained as follows: , Furthermore, after obtaining the expected distribution, it is necessary to learn the similarity distribution between samples. When positive samples are mixed in with negative samples, the sum of the exponents of the true negative samples needs to be corrected. Here, the sample is the normalized similarity between samples, and the sum of the exponents of the true negative samples... for: , in, For each sample Number of true negative samples , For similarity function, , This is the temperature coefficient. Positive sample index (sample) Its data augmentation samples (index) The sum of all sample indices minus the indices of self-similar and positive samples.

[0063] After bias removal based on the above formula, each sample is obtained. Loss function: , in, Since it is an extremely small floating-point number, the stability of exponential function calculations can be guaranteed. Also known as the bias removal loss value of sample k.

[0064] Step S83: Calculate the energy weight of the training sample based on the duration of the training sample, and weight the loss vector based on the energy weight to obtain the loss scalar.

[0065] In this embodiment, for each sample Calculate the sum of the durations of the corresponding events: , Where L is the total number of time steps in the event sequence. This represents the duration of the event in sample k at the i-th time step.

[0066] Next, the statistics for the duration of events in the batch sample are calculated. In this process, the mean of the event duration is calculated first: , Where B is half the size of the training batch, and 2B is the total number of samples in the batch.

[0067] Then calculate the standard deviation of the event duration: , Finally, the event duration was standardized to obtain a sample. Standardized energy weights: , in, For very small floating-point numbers The standard deviation of the event duration ensures the stability of the exponential function calculation.

[0068] Finally, using samples The loss is weighted by the standardized energy weights to obtain the final loss scalar, i.e., the final weighted loss value: , In order to reduce the erroneous rejection between positive samples and make full use of the physical characteristics of event encoding, this embodiment optimizes the traditional NTXentLoss loss function by performing biased estimation, and uses event duration to weight the loss, thereby improving the model training efficiency, making the training loss decrease faster, and increasing the robustness of the model.

[0069] Based on any of the above embodiments, in Embodiment 4 of this application, before generating at least two augmented samples of the training subsequence in the training event sequence, the number of samples in the training event sequence can also be obtained. If the number of samples is less than the preset number of events, the tensor of the training event sequence is filled with empty events to obtain multiple empty samples. That is, when the number of samples in the training event sequence is less than the preset number of events, empty samples are generated by filling the tensor with empty events, thereby supplementing the training sample scale required for contrastive learning and avoiding model training non-convergence, overfitting, or insufficient feature learning due to insufficient sample quantity. At the same time, by filling the tensor of the unified training event sequence with empty events, the tensor dimension and input format are adapted to the requirements of the contrastive learning model for fixed-dimensional input, without adjusting the model structure due to differences in sample length and quantity. In addition, without destroying the core structure of the event sequence, the supplemented events can be augmented through augmentation methods to expand the diversity of training samples, improve the model's adaptability and generalization ability to sparse, short-duration, and low-sample-volume scenarios, thereby ensuring the stability of the model training process, accelerating model convergence, and reducing the cost of data preprocessing and sample selection in real vehicle networking scenarios.

[0070] Based on any of the above embodiments, in Embodiment 5 of this application, after extracting the set of representation vectors corresponding to the event sequence based on the pre-trained contrastive learning model, the pre-trained contrastive learning model can also be evaluated using an evaluation metric. Specifically, the standard deviation of the evaluation metric for sample points when the detection result is normal is obtained. Then, when the standard deviation exceeds a preset standard deviation, the pre-trained contrastive learning model is iteratively updated. Specifically, the evaluated model can be deployed in a production environment, and the evaluation metric MSDN can be used as a periodic monitoring item for the model. A statistical threshold control method is used to continuously monitor the MSDN metric. The mean and standard deviation of the MSDN metric are obtained based on the metrics within one month of stable operation. Subsequently, when the metric exceeds a preset threshold, it indicates that the model performance has deteriorated and retraining with collected samples is required. That is, the training actions of Embodiment 2 are re-executed.

[0071] Based on this, the implementation process obtained by combining the above embodiments can be as follows: Figure 4 As shown in steps S101-S105, specifically, during the training phase, the CAN signal matrix of all events is first acquired to construct a standardized data base. Then, the CAN signal matrix is ​​aggregated into event sequences to alleviate the training difficulties caused by the order-of-magnitude span of CAN signal lengths across different state segments. Subsequently, a contrastive learning method is used to extract representation vectors from the event sequences, and the representation vectors are evaluated, along with iterative model processing. The model is then evaluated to see if it meets expectations; if not, it iterates again; if it meets expectations, the subsequent model deployment and maintenance phase proceeds. During maintenance, real-time vehicle data is continuously monitored and anomaly identification is performed. If an indicator deviates from a threshold during anomaly identification, the model is retrained and iterated again, forming a closed-loop lifecycle of data acquisition → feature extraction → model validation → deployment and maintenance → iterative optimization. This solves the pain points of delayed anomaly identification and reliance on manual annotation, and improves the model's generalization ability and robustness in real-world vehicle networking scenarios through length normalization and unsupervised contrastive learning.

[0072] Based on any of the above embodiments, in Embodiment Six of this application, during model training, in addition to using the isolated forest algorithm to detect anomalies in the representation vectors of different event sequences, Principal Component Analysis (PCA) can be used to reduce the representation vectors of different event sequences from 128 dimensions to 2 dimensions, changing the tensor shape to... Sample size, 2 This allows for the projection of two-dimensional vectors onto a plane, enabling intuitive observation of the distribution characteristics of different event sequences and rapid determination of any anomalies. Furthermore, the system can automatically extract the results of principal component analysis through automatic identification.

[0073] For example, the visualization results output after principal component analysis are as follows: Figure 5 As shown, the coordinate axes are constructed using the two principal component dimensions of Principal Component Analysis (PCA). The horizontal axis represents Principal Component 1, with values ​​ranging from 0 to 600. The vertical axis represents Principal Component 2, with values ​​ranging from -125 to 50. Dots represent normal samples, and crosses represent abnormal driving samples. It can be seen that normal samples are highly clustered in the 0-50 range of Principal Component 1 and the -25-0 range of Principal Component 2, indicating that the normal driving behavior representation vectors extracted by the model have strong homogeneity. Crosses, on the other hand, are widely distributed outside the normal sample clustering area, exhibiting a significant discrete distribution characteristic. This visualization clearly demonstrates the effectiveness of the model's feature extraction; that is, the representation vectors of normal driving behavior form a compact core cluster in the latent space, while the representation vectors of abnormal driving behavior deviate significantly from this clustering area. Furthermore, different types of abnormal samples exhibit differentiated discrete distributions in the latent space, indicating that the model can accurately distinguish the features of normal and abnormal driving behaviors, meeting the core requirement of anomaly identification.

[0074] Understandably, during dimensionality reduction, PCA retains the most important information from the original data—the direction of maximum variance—by calculating the covariance matrix and extracting eigenvalues ​​and eigenvectors. The high-dimensional representation vector is then projected onto a plane formed by the first two principal components. Thus, the representation vectors of normal event sequences typically cluster in specific regions, while the representation vectors of anomalous event sequences may deviate significantly from these clusters due to their significant differences from normal samples. Visual analysis of the projected two-dimensional scatter plot allows for the rapid identification of potential anomalies, providing preliminary judgment for anomaly detection and improving its efficiency and accuracy.

[0075] Secondly, Figure 6 A schematic diagram of an embodiment of the networked vehicle abnormal event identification device of the present invention is shown. Figure 6 As shown, the device 300 includes: a signal processing module 310, an event aggregation module 320, a feature extraction module 330, and an anomaly recognition module 340.

[0076] The signal processing module 310 is used to acquire the target CAN signal of a vehicle operating state segment based on a preset sampling frequency, and generate a CAN signal matrix corresponding to the target CAN signal according to the average number of frames of the state segment.

[0077] The event aggregation module 320 is used to determine the event corresponding to each frame in the CAN signal matrix and aggregate the events into an event sequence based on the duration of the events.

[0078] The feature extraction module 330 is used to extract the set of representation vectors corresponding to the event sequence based on a pre-trained contrastive learning model.

[0079] The anomaly detection module 340 is used to determine that the vehicle has an abnormal event if the detection result of a subset in the representation vector set is abnormal.

[0080] In an alternative embodiment, the signal processing module 310 is further configured to acquire the target CAN signal based on the preset sampling frequency; determine the quotient between the duration of the vehicle operating state segment and the preset sampling rate as the average frame count; and construct the CAN signal matrix with the number of target CAN signals as rows and the average frame count as columns.

[0081] In an optional manner, the event aggregation module 320 is further configured to determine the event segment containing the CAN signal combination state corresponding to each frame in the CAN signal matrix; merge the event segments of multiple consecutive frames with the CAN signal combination state remaining consistent into an event, and record the duration of the multiple consecutive frames; and arrange all the events in chronological order to obtain multiple event sequences corresponding to all the vehicle operating state segments.

[0082] In one optional embodiment, the connected vehicle abnormal event identification device 300 further includes a training module 350, configured to acquire a training CAN signal matrix generated from training CAN signals based on normal samples, and aggregate the training CAN signal matrix into a training event sequence; generate at least two augmented samples of the training event sequence; extract training representation vectors of the training event sequence based on the encoder of the initial comparison model, and perform projection processing on the training representation vectors through the decoder to obtain a projection result; calculate the loss scalar corresponding to the projection result based on the contrastive loss function; and iterate the encoder parameters and decoder parameters of the initial comparison model through backpropagation based on the loss scalar until the model converges to obtain the pre-trained contrastive learning model.

[0083] In an alternative embodiment, the training module 350 is further configured to calculate the expected negative sample distribution of the training samples anchored by the projection result, and calculate the negative sample index sum based on the expected distribution; calculate the loss vector of the training samples based on the negative sample index sum; calculate the energy weight of the training samples based on the duration of the training samples, and weight the loss vector based on the energy weight to obtain the loss scalar.

[0084] In an alternative approach, the training module 350 is further configured to scale the duration of the first training event in the training event sequence based on a preset scaling ratio to obtain a first augmented sample. The signal state of at least one CAN signal in the second training event in the training event sequence is masked to obtain the second enhanced sample.

[0085] In one alternative approach, the training module 350 is further configured to obtain the number of samples in the training event sequence; if the number of samples is less than a preset number of events, the tensor of the training event sequence is filled with empty events based on the preset number of events.

[0086] In one alternative approach, the connected vehicle abnormal event recognition device 300 further includes an iterative update module 360, which is used to obtain the standard deviation of the evaluation index of sample points whose detection results are normal; if the standard deviation is greater than a preset standard deviation, the pre-trained contrastive learning model is iteratively updated.

[0087] Figure 7 The diagram shows a structural schematic of an embodiment of the vehicle of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the vehicle.

[0088] like Figure 7 As shown, the vehicle may include: a processor processor 402. Communication Interface CommunicationsInterface 404, Memory memory 406, and communication bus 408.

[0089] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other network elements, such as clients or other servers. Processor 402 executes program 410, specifically performing the relevant steps in the above-described embodiment of the method for identifying abnormal events in connected vehicles.

[0090] Specifically, program 410 may include program code, which includes computer-executable instructions.

[0091] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The vehicle may include one or more processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0092] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0093] Thirdly, embodiments of the present invention provide a computer-readable storage medium storing at least one executable instruction that, when executed on a vehicle / vehicle control device, causes the vehicle / connected vehicle abnormal event identification device to perform the connected vehicle abnormal event identification method in any of the above method embodiments.

[0094] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, the embodiments of this invention are not directed to any particular programming language.

[0095] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.

[0096] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.

[0097] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for identifying abnormal events in connected vehicles, characterized in that, Applied to vehicles, the method includes: The target CAN signal of the vehicle operating status segment is acquired based on a preset sampling frequency, and a CAN signal matrix corresponding to the target CAN signal is generated based on the average number of frames of the status segment. The events in the CAN signal matrix are identified, and the events are aggregated into an event sequence based on the duration of the events. The set of representation vectors corresponding to the event sequence is extracted based on a pre-trained contrastive learning model; If a subset of the representation vector set shows an abnormal detection result, it is determined that the vehicle has an abnormal event.

2. The method as described in claim 1, characterized in that, The process of acquiring the target CAN signal of a vehicle operating state segment based on a preset sampling frequency, and generating a CAN signal matrix corresponding to the target CAN signal based on the average frame number of the state segment, includes: The target CAN signal is acquired based on the preset sampling frequency; The quotient between the duration of the vehicle operation state segment and the preset sampling rate is determined as the average number of frames; The CAN signal matrix is ​​constructed using the number of target CAN signals as rows and the average number of frames as columns.

3. The method as described in claim 1, characterized in that, The step of determining the events in the CAN signal matrix and aggregating the events into an event sequence based on the duration of the events includes: Determine the event segment containing the combined state of CAN signals corresponding to each frame in the CAN signal matrix; The event segments of multiple consecutive frames with consistent CAN signal combination states are merged into an event, and the duration of the multiple consecutive frames is recorded. Based on the chronological order of all the events, multiple event sequences corresponding to all the vehicle operation state segments are obtained.

4. The method as described in claim 1, characterized in that, Before aggregating events with identical CAN signal parameters within consecutive time periods into an event sequence based on the CAN signal matrix corresponding to vehicle operating state segments, the process also includes: Obtain the training CAN signal matrix generated based on the training CAN signals of normal samples, and aggregate the training CAN signal matrix into a training event sequence; Generate at least two augmented samples of the training event sequence; The encoder extracts the training representation vector of the training event sequence based on the initial alignment model, and the decoder projects the training representation vector to obtain the projection result. The loss scalar corresponding to the projection result is calculated based on the contrastive loss function; The encoder and decoder parameters of the initial alignment model are iterated by backpropagation based on the loss scalar until the model converges, thus obtaining the pre-trained contrastive learning model.

5. The method as described in claim 4, characterized in that, The step of calculating the loss scalar corresponding to the projection result based on the contrastive loss function includes: Calculate the expected negative sample distribution of the training samples anchored by the projection result, and calculate the negative sample index based on the expected distribution. Calculate the loss vector of the training samples based on the negative sample index; The energy weights of the training samples are calculated based on their duration, and the loss vector is weighted based on these energy weights to obtain the loss scalar.

6. The method as described in claim 4, characterized in that, The at least two augmented samples used to generate the training event sequence include: The duration of the first training event in the training event sequence is scaled based on a preset scaling ratio to obtain the first augmented sample. The signal state of at least one CAN signal in the second training event in the training event sequence is masked to obtain the second enhanced sample.

7. The method as described in claim 4, characterized in that, The generation of the training event sequence, before at least two augmented samples of the training subsequence, further includes: Obtain the number of samples in the training event sequence; If the number of samples is less than the preset number of events, the tensor of the training event sequence is filled with empty events based on the preset number of events.

8. The method as described in claim 1, characterized in that, After the pre-trained contrastive learning model extracts the set of representation vectors corresponding to the event sequence, it further includes: Obtain the standard deviation of the evaluation index for sample points with normal test results; If the standard deviation is greater than the preset standard deviation, the pre-trained contrastive learning model is iteratively updated.

9. A device for identifying abnormal events in networked vehicles, characterized in that, The device includes: The signal processing module is used to acquire the target CAN signal of a vehicle operating state segment based on a preset sampling frequency, and generate a CAN signal matrix corresponding to the target CAN signal according to the average number of frames of the state segment. An event aggregation module is used to determine the event corresponding to each frame in the CAN signal matrix and aggregate the events into an event sequence based on the duration of the events. The feature extraction module is used to extract the set of representation vectors corresponding to the event sequence based on a pre-trained contrastive learning model; An anomaly detection module is used to determine that the vehicle has an abnormal event if the detection result of a subset in the representation vector set is abnormal.

10. A vehicle, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the method for identifying abnormal events of connected vehicles as described in any one of claims 1-8.