Vehicle engine abnormal detection method, device, electronic device and storage medium
By combining the feature extraction and fusion of on-board images and engine voiceprint sequences, the problem of low accuracy in engine abnormality detection is solved, and more efficient abnormality detection is achieved.
Patent Information
- Application Number
- CN202011584113.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-12-28
AI Technical Summary
In the prior art, the accuracy of automobile engine abnormality detection is greatly affected by subjective factors, and all factors of engine abnormality during driving cannot be considered objectively and comprehensively, resulting in a low detection accuracy.
By acquiring the on-board image sequence and engine voiceprint sequence after the engine is started, the image feature extraction network and the voiceprint feature extraction network extracts the spatiotemporal and temporal and frequency soundprint features, and fuses it, and finally performs abnormal detection through the classification prediction network.
The accuracy of engine abnormality detection is improved, all factors of engine abnormality can be more comprehensively abstracted during driving, the threshold setting is converted to classification problems, and the complex information of the on-board image and voiceprint sequence are used for feature extraction.
Smart Images

Figure CN114693945B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a vehicle engine abnormality detection method, device, electronic equipment and storage medium. Background Art
[0002] In recent years, with the rapid development of the economy, more and more families have chosen cars as their means of transportation. As one of the three major components of a car, the engine is a crucial component that provides the vehicle with power. When an engine anomaly occurs while driving, the driver typically relies on hearing to detect the anomaly. However, at the onset of the anomaly, the driver's judgment is often inaccurate and inadequate. For electronic components monitoring the engine, the accuracy of anomaly detection is significantly affected by subjective factors, as the parameter data of these components is typically set based on anomaly thresholds. These thresholds rely on the subjective experience of professionals, and the selected parameter types are relatively limited, failing to objectively and comprehensively consider all factors that may contribute to the anomaly during driving. This results in low anomaly detection accuracy. Summary of the Invention
[0003] An embodiment of the present invention provides a vehicle engine anomaly detection method, which can classify and predict vehicle engine anomaly detection by using a vehicle-mounted image sequence and an engine soundprint sequence after the engine is started. Since the anomaly type is determined based on the anomaly that has occurred, the threshold setting problem in anomaly detection can be converted into a classification problem. In addition, the vehicle-mounted image sequence is used to assist the engine soundprint sequence in detection, making full use of the complex information of the vehicle-mounted image for feature extraction. Furthermore, all factors of the vehicle engine anomaly during driving can be abstracted into high-level features, thereby improving the accuracy of vehicle engine anomaly detection.
[0004] In a first aspect, an embodiment of the present invention provides a method for detecting abnormality in a vehicle engine, the method comprising:
[0005] Acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously;
[0006] Extracting features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood;
[0007] Extracting features of the engine voiceprint sequence through a preset voiceprint feature extraction network to obtain the engine's time-frequency voiceprint features;
[0008] fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine;
[0009] The fusion features are classified and predicted by a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
[0010] Optionally, obtaining the engine soundprint sequence after the engine is started includes:
[0011] Obtain the first engine soundprint sequence after the engine is started;
[0012] De-noising the first engine soundprint sequence to obtain a second engine soundprint sequence;
[0013] The second engine soundprint sequence is converted from time domain information to frequency domain information, and an engine soundprint sequence is obtained according to the time domain information and the frequency domain information.
[0014] Optionally, converting the second engine voiceprint sequence from time domain information to frequency domain information includes:
[0015] performing frame processing on the second engine voiceprint sequence to obtain a frame-processed second engine voiceprint sequence;
[0016] performing windowing processing on the second engine voiceprint sequence after the frame processing to obtain a windowed second engine voiceprint sequence;
[0017] Performing fast Fourier transform point processing on the windowed second engine voiceprint sequence to convert the second engine voiceprint sequence from time domain information to frequency domain information.
[0018] Optionally, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine have the same feature dimension, and fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine includes:
[0019] Performing a dot product on the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine to obtain a first fusion feature;
[0020] The first fusion features are accumulated to obtain a second fusion feature as a fusion feature of the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine.
[0021] Optionally, the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are trained using the same data set, and the training steps include:
[0022] Constructing a first data set, the first data set including a sample vehicle image sequence, a sample engine soundprint sequence, and corresponding fault type labels, wherein the sample vehicle image sequence is a vehicle image sequence after the engine is started in an engine abnormality condition, the sample engine soundprint sequence is an engine soundprint sequence after the engine is started in an engine abnormality condition, and one sample vehicle image sequence corresponds to one sample engine soundprint sequence;
[0023] The image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are jointly trained using the first data set.
[0024] Optionally, the jointly training the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network using the data set includes:
[0025] Extracting features from the sample vehicle image sequence using the image feature extraction network to be trained to obtain sample spatiotemporal features of the engine hood;
[0026] Extracting features from the sample engine voiceprint sequence using a voiceprint feature extraction network to be trained to obtain sample time-frequency voiceprint features of the engine;
[0027] fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the sample spatiotemporal features of the engine hood and the sample time-frequency soundprint features of the engine;
[0028] The fusion feature is subjected to sample classification prediction through the classification prediction network to be trained to obtain the sample classification prediction result. The error loss between the sample classification prediction result and the corresponding fault type label is calculated through a preset loss function. The network parameters of the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are adjusted according to the error loss.
[0029] Optionally, after the training is completed, the method further includes:
[0030] Acquire a normal vehicle image sequence and a corresponding normal engine soundprint sequence for each vehicle type after engine startup under normal engine conditions, and an abnormal vehicle image sequence and a corresponding abnormal engine soundprint sequence for the corresponding vehicle after engine startup under abnormal engine conditions as a second data set, wherein in the second data set, one normal sample vehicle image sequence corresponds to one normal sample engine soundprint sequence, and one abnormal sample vehicle image sequence corresponds to one abnormal sample engine soundprint sequence, one second data set corresponds to one vehicle type, and samples in the second data set are not repeated with those in the first data set;
[0031] When initializing the user's vehicle, obtain a vehicle image sequence and a corresponding engine soundprint sequence of the user's vehicle after the engine is started under normal conditions;
[0032] A corresponding second data set is matched based on the on-board image sequence of the user vehicle after the engine is started under normal engine conditions and the corresponding engine soundprint sequence, and the network parameters of the image feature extraction network, the soundprint feature extraction network, and the classification prediction network are fine-tuned based on the corresponding second data set.
[0033] In a second aspect, an embodiment of the present invention further provides a vehicle engine abnormality detection device, the device comprising:
[0034] A first acquisition module is configured to acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously;
[0035] a first extraction module, configured to extract features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood;
[0036] a second extraction module, configured to extract features from the engine voiceprint sequence using a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature of the engine;
[0037] a first fusion module, configured to fuse the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine;
[0038] The detection module is used to perform classification prediction on the fusion features through a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
[0039] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the vehicle engine abnormality detection method provided in an embodiment of the present invention are implemented.
[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the vehicle engine abnormality detection method provided in the embodiment of the invention are implemented.
[0041] In an embodiment of the present invention, a vehicle-mounted image sequence and an engine soundprint sequence are obtained after the engine is started, the vehicle-mounted image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle-mounted image sequence and the engine soundprint sequence are obtained simultaneously; feature extraction is performed on the vehicle-mounted image sequence through a preset image feature extraction network to obtain the spatiotemporal features of the engine hood; feature extraction is performed on the engine soundprint sequence through a preset soundprint feature extraction network to obtain the time-frequency soundprint features of the engine; the spatiotemporal features of the engine hood are fused with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine; classification prediction is performed on the fused features through a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result. The vehicle engine anomaly detection and classification prediction can be performed by using the vehicle image sequence and engine soundprint sequence after the engine is started. Since the anomaly type is determined based on the anomaly that has occurred, the threshold setting problem in anomaly detection can be converted into a classification problem. In addition, the vehicle image sequence is used to assist the engine soundprint sequence in detection, which makes full use of the complex information of the vehicle image for feature extraction. In addition, all factors of the vehicle engine anomaly during driving can be abstracted into high-level features, thereby improving the accuracy of vehicle engine anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1 This is a flow chart of a vehicle engine abnormality detection method provided by an embodiment of the present invention;
[0044] Figure 2 This is a flow chart of a method for obtaining an engine voiceprint sequence provided by an embodiment of the present invention;
[0045] Figure 2a This is a schematic diagram of the relationship between time domain information and frequency domain information provided by an embodiment of the present invention;
[0046] Figure 3 This is a flow chart of converting time domain information into frequency domain information provided by an embodiment of the present invention;
[0047] Figure 4 This is a training flow chart of a vehicle engine anomaly detection model provided by an embodiment of the present invention;
[0048] Figure 5 1 is a schematic structural diagram of a vehicle engine abnormality detection device provided by an embodiment of the present invention;
[0049] Figure 6 is a structural diagram of a first acquisition module provided by an embodiment of the present invention;
[0050] Figure 7 This is a schematic structural diagram of a conversion submodule provided by an embodiment of the present invention;
[0051] Figure 8 This is a schematic structural diagram of a fusion module provided by an embodiment of the present invention;
[0052] Figure 9 1 is a schematic structural diagram of another vehicle engine abnormality detection device provided by an embodiment of the present invention;
[0053] Figure 10 is a structural diagram of a training module provided by an embodiment of the present invention;
[0054] Figure 11 1 is a schematic structural diagram of another vehicle engine abnormality detection device provided by an embodiment of the present invention;
[0055] Figure 12 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0057] See Figure 1 , Figure 1 This is a flow chart of a vehicle engine abnormality detection method provided by an embodiment of the present invention. Figure 1 As shown, the following steps are included:
[0058] 101. Obtain a vehicle image sequence and an engine soundprint sequence after the engine is started.
[0059] In an embodiment of the present invention, the vehicle image sequence includes an engine hood image, the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously.
[0060] After the engine is started, the engine can be operated while the vehicle is stopped or while the vehicle is moving. The vehicle can be a motor vehicle that requires an engine. The engine refers to a vehicle power component that converts the internal energy of a working fluid into mechanical energy for output through combustion, such as a gasoline internal combustion engine or a diesel internal combustion engine.
[0061] The above-mentioned vehicle-mounted image sequence can be understood as a vehicle-mounted video. The above-mentioned vehicle-mounted image sequence can be obtained by shooting with a camera set in the vehicle. The above-mentioned camera is set to aim at the engine hood of the vehicle so that the camera can capture the vehicle-mounted image sequence including the engine hood.
[0062] The above-mentioned engine soundprint sequence can be obtained through a sound sensor installed in the engine compartment. The above-mentioned engine soundprint sequence can be framed according to the framing rules of the vehicle-mounted image sequence. For example, if the vehicle-mounted image sequence is 29 frames per second, the engine soundprint can be divided into 29 segments per second to form a corresponding engine soundprint sequence.
[0063] It should be noted that the above-mentioned vehicle-mounted image sequence and the engine soundprint sequence have the same preset duration. For example, if the vehicle-mounted image sequence is 3 seconds, the engine soundprint sequence is also 3 seconds. It can be further understood that the number of frames of the above-mentioned vehicle-mounted image sequence is the same as the number of frames of the engine soundprint sequence, and the interval time between two adjacent frames in the vehicle-mounted image sequence is the same as the interval time between two adjacent frames in the engine soundprint sequence.
[0064] The time domain information is the voiceprint information originally collected by the sound sensor and can be understood as voiceprint information described on the time axis. The frequency domain information can be obtained by converting the time domain information and can be understood as voiceprint information described on the frequency axis. Dynamic voiceprint information can be converted from the time domain to the frequency domain using Fourier series and Fourier transform.
[0065] Optional, see Figure 2 , Figure 2 This is a flow chart of a method for obtaining an engine voiceprint sequence provided by an embodiment of the present invention. Figure 2 As shown, the conversion method includes the following steps:
[0066] 201. Obtain a first engine voiceprint sequence after the engine is started.
[0067] In an embodiment of the present invention, after the engine of the vehicle is started, the sound print information in the engine compartment can be collected in real time or periodically by a sound sensor as a first engine sound print sequence.
[0068] 202. De-noise the first engine soundprint sequence to obtain a second engine soundprint sequence.
[0069] In the embodiment of the present invention, the first engine soundprint sequence may be denoised by autocorrelation denoising to eliminate the environmental noise of the first engine soundprint sequence and obtain the second engine soundprint sequence.
[0070] 203. Convert the second engine voiceprint sequence from time domain information to frequency domain information, and obtain the engine voiceprint sequence according to the time domain information and frequency domain information of the second engine voiceprint sequence.
[0071] In an embodiment of the present invention, the engine soundprint sequence may include both time domain information and frequency domain information. The conversion of the second engine soundprint sequence from time domain information to frequency domain information may be performed by short-time Fourier transform. After obtaining the frequency domain information of the second engine soundprint sequence, the frequency domain information of the second engine soundprint sequence is fused with the time domain information of the second soundprint sequence. The fused engine soundprint sequence has both time and frequency dimensions. The engine soundprint sequence may be as follows: Figure 2a shown.
[0072] For details, see Figure 3 , Figure 3 This is a flow chart of converting time domain information into frequency domain information provided by an embodiment of the present invention. Figure 3 As shown, the following steps are included:
[0073] 301. Perform frame processing on the second engine voiceprint sequence to obtain a frame-processed second engine voiceprint sequence.
[0074] In an embodiment of the present invention, the framing process may be to divide the second voiceprint sequence into frames according to the number of frames in the vehicle-mounted image sequence. That is, the second voiceprint sequence is divided into as many frames as the number of frames in the vehicle-mounted image sequence. For example, if the vehicle-mounted image sequence includes 100 frames, the second voiceprint sequence is divided into 100 frames. It should be noted that the vehicle-mounted image sequence and the second voiceprint sequence have the same duration.
[0075] Furthermore, the length of the second voiceprint sequence can be expressed by the following formula:
[0076] N=f s ×t
[0077] Where N is the length of the second voiceprint sequence, f s is the sampling frequency of the second voiceprint sequence (the sampling frequency of the sound sensor), and t is the sampling duration of the second voiceprint sequence.
[0078] After framing, the second soundprint sequence of the engine is:
[0079] {y1,y2,…,y m}
[0080] Among them, the above y m is the mth voiceprint frame in the second voiceprint sequence of the engine after the frame processing, and m is the total number of voiceprint frames in the second voiceprint sequence of the engine after the frame processing.
[0081] In one possible embodiment, in order to achieve a smooth transition between voiceprint frames in the second engine voiceprint sequence after framing and maintain the continuity of the voiceprint data, the second engine voiceprint sequence can be framed using an overlapping segmentation method. The overlapping segmentation method can be understood as the overlapping portion between the current voiceprint frame and the previous voiceprint frame, and the overlapping portion between the current voiceprint frame and the next voiceprint frame. In this case, assuming that the length of each voiceprint frame in the second engine voiceprint sequence after framing is nfft, and the overlapping length between two adjacent voiceprint frames in the second engine voiceprint sequence after framing is overlap, the length L of the second engine voiceprint sequence after framing can be expressed by the following formula:
[0082] L=(nfft-overlap)×m
[0083] 302. Perform windowing processing on the second engine voiceprint sequence after the frame processing to obtain a windowed second engine voiceprint sequence.
[0084] In the embodiment of the present invention, the initial window type, window length, and sliding step size can be set to determine the window parameters to be added. Further, the second engine voiceprint sequence after frame processing is windowed to obtain the windowed second engine voiceprint sequence. Specifically, the window function y can be calculated according to the adaptive time scale rule. window And the window length, multiply the framed voiceprint data by the window function to obtain the windowed second engine voiceprint sequence:
[0085] y i =y i ×y window
[0086] Among them, y i is the i-th voiceprint frame data after framing, y i is the corresponding i-th windowed voiceprint frame.
[0087] 303. Perform fast Fourier transform point processing on the windowed second engine voiceprint sequence to convert the second engine voiceprint sequence from time domain information to frequency domain information.
[0088] In an embodiment of the present invention, after obtaining a windowed second engine soundprint sequence, the windowed second engine soundprint sequence can be subjected to Fast Fourier Transform (FFT) processing. This Fast Fourier Transform (FFT) processing converts the second engine soundprint sequence from time domain information to frequency domain information, obtaining the frequency and amplitude information corresponding to the second engine soundprint sequence at each moment, thereby obtaining the time-frequency characteristics of the second engine soundprint sequence in this embodiment of the present invention. Specifically, the soundprint data is first processed using an initial window type, window length, sliding step size, and Fast Fourier Transform (FFT) point number. The window type can be selected to provide good sidelobe suppression, the sliding step size can be 100%, and the FFT point number can be selected to a smaller value. This allows for rapid processing of the voiceprint data, improving processing speed. Next, the window type, window length, sliding step size, and FFT point number are adjusted based on the frequency of the voiceprint data to meet the requirements of time-frequency analysis. Specifically, adjustments can be made by using a window with a narrower main lobe width, a 40% sliding step size, and an increased number of Fast Fourier Transform points. After adjusting the adaptive time scale described above, the window type, window length, sliding step size, and number of Fast Fourier Transform points that meet the requirements of time-frequency analysis are obtained. Short-time Fourier transform processing is then performed on the voiceprint data based on the window type, window length, sliding step size, and number of Fast Fourier Transform points that meet the requirements of time-frequency analysis.
[0089] 102. Feature extraction is performed on the vehicle image sequence through a preset image feature extraction network to obtain the spatiotemporal features of the engine hood.
[0090] In an embodiment of the present invention, the above-mentioned preset image feature extraction network can be understood as a pre-trained image feature extraction network, and the above-mentioned preset image feature extraction network can be constructed based on a convolutional neural network. The above-mentioned image feature extraction network may include a three-dimensional convolution module for extracting the spatiotemporal features of the engine hood in the vehicle-mounted image sequence. Due to the jitter of the engine itself, the vehicle-mounted image sequence will produce jitter between frames, which is reflected as the image jitter of the engine hood. After this jitter is abstracted into high-level semantics (spatiotemporal features), it has a certain auxiliary role in expressing the state of the engine. Therefore, the vehicle-mounted image sequence can be used as a means to assist in engine abnormality detection.
[0091] 103. The engine voiceprint sequence is subjected to feature extraction through a preset voiceprint feature extraction network to obtain the engine's time-frequency voiceprint features.
[0092] In an embodiment of the present invention, the preset voiceprint feature extraction network can be understood as a pretrained voiceprint feature extraction network, which can be constructed based on a convolutional neural network. The voiceprint feature extraction network can include a three-dimensional convolution module for extracting the engine's time-frequency voiceprint features. By extracting features from the engine soundprint sequence using the voiceprint feature extraction network, instead of relying on the driver's auditory perception, the engine information in the engine soundprint sequence can be more fully utilized to detect engine anomalies, potentially even detecting anomalies that the driver cannot perceive.
[0093] 104. Fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain fusion features of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine.
[0094] In an embodiment of the present invention, the above-mentioned fusion can be multiplicative fusion and additive fusion. Furthermore, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine can be dot-producted to obtain a first fusion feature; the first fusion feature can be accumulated to obtain a second fusion feature as a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine.
[0095] Furthermore, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine have the same feature dimension. In this way, it is possible to perform a dot product of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine, and then accumulate all the dot product results to obtain a fusion feature.
[0096] Of course, in some possible embodiments, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine may also have different feature dimensions. In this case, the above-mentioned fusion may be a splicing fusion. For example, the spatiotemporal features of the engine hood are 64-dimensional features, and the time-frequency soundprint features of the engine are 128-dimensional features. They can be spliced into a first fusion feature of 192 dimensions, and then the 192-dimensional first integrated feature is linearly transformed to obtain a second fusion feature of 128 dimensions, or a second fusion feature of 256 dimensions is obtained as the fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine.
[0097] 105. The fusion features are classified and predicted through a preset classification prediction network, and the classification prediction results are obtained as the vehicle engine abnormality detection results.
[0098] In an embodiment of the present invention, the above-mentioned preset classification prediction network can be understood as a pre-trained fully connected classification network, and the above-mentioned preset classification prediction network can be constructed based on a fully connected neuron classification network. The above-mentioned classification prediction network may include input neurons of the same dimension as the fusion feature, and output neurons corresponding to the abnormality type. For example, when the abnormality type is a binary classification, it may include two neurons, one for abnormality and the other for normal. For example, if the output classification prediction result is 1, the vehicle engine abnormality detection result is an abnormality, and if the output classification prediction result is 0, the vehicle engine abnormality detection result is normal. When the abnormality type is multi-classification, such as k abnormality types, the output neurons can be k+1, and the extra one is the classification of the normal type.
[0099] In an embodiment of the present invention, a vehicle-mounted image sequence and an engine soundprint sequence are obtained after the engine is started, the vehicle-mounted image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle-mounted image sequence and the engine soundprint sequence are obtained simultaneously; feature extraction is performed on the vehicle-mounted image sequence through a preset image feature extraction network to obtain the spatiotemporal features of the engine hood; feature extraction is performed on the engine soundprint sequence through a preset soundprint feature extraction network to obtain the time-frequency soundprint features of the engine; the spatiotemporal features of the engine hood are fused with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine; classification prediction is performed on the fused features through a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result. The vehicle engine anomaly detection and classification prediction can be performed by using the vehicle image sequence and engine soundprint sequence after the engine is started. Since the anomaly type is determined based on the anomaly that has occurred, the threshold setting problem in anomaly detection can be converted into a classification problem. In addition, the vehicle image sequence is used to assist the engine soundprint sequence in detection, which makes full use of the complex information of the vehicle image for feature extraction. In addition, all factors of the vehicle engine anomaly during driving can be abstracted into high-level features, thereby improving the accuracy of vehicle engine anomaly detection.
[0100] It should be noted that the vehicle engine anomaly detection method provided in the embodiment of the present invention can be applied to devices such as mobile phones, monitors, computers, servers, etc. that can perform vehicle engine anomaly detection.
[0101] Optional, see Figure 4 , Figure 4This is a training flow chart of a vehicle engine anomaly detection model provided by an embodiment of the present invention. The vehicle engine anomaly detection model includes an image feature extraction network, a voiceprint feature extraction network, and a classification prediction network. The output of the image feature extraction network is connected to the input of the classification prediction network, and the output of the voiceprint feature extraction network is connected to the input of the classification prediction network. Furthermore, the output of the image feature extraction network is connected to the input of the classification prediction network through a preset fusion layer, and the output of the voiceprint feature extraction network is also connected to the input of the classification prediction network through a preset fusion layer. Figure 4 As shown, the following steps are included:
[0102] 401. Build a first data set.
[0103] In an embodiment of the present invention, the above-mentioned first data set includes a sample vehicle-mounted image sequence, a sample engine soundprint sequence and a corresponding fault type label. The above-mentioned sample vehicle-mounted image sequence is a vehicle-mounted image sequence after the engine of the vehicle is started under an engine abnormality condition. The above-mentioned sample engine soundprint sequence is an engine soundprint sequence after the engine of the vehicle is started under an engine abnormality condition. The above-mentioned one sample vehicle-mounted image sequence corresponds to one sample engine soundprint sequence.
[0104] The above sample vehicle image sequence and the sample engine soundprint sequence have the same duration. Figure 1 The vehicle image sequences in the embodiment have the same duration. The above sample engine soundprint sequence includes time domain information and frequency domain information. Specifically, the above sample engine soundprint sequence can be Figure 2 The frequency domain information in the above sample engine voiceprint sequence can be obtained by the method of the embodiment. Figure 3 The conversion method of the embodiment is used to convert.
[0105] 402. Using the first data set, jointly train the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network.
[0106] In an embodiment of the present invention, the image feature extraction network to be trained may be used to perform feature extraction on the sample vehicle image sequence to obtain sample spatiotemporal features of the engine hood.
[0107] The sample engine voiceprint sequence is subjected to feature extraction by the voiceprint feature extraction network to be trained to obtain the sample time-frequency voiceprint features of the engine.
[0108] The spatiotemporal features of the engine hood are fused with the time-frequency soundprint features of the engine to obtain fusion features of the sample spatiotemporal features of the engine hood and the sample time-frequency soundprint features of the engine.
[0109] The classification prediction network to be trained performs sample classification prediction on the fused features to obtain a sample classification prediction result. A preset loss function is used to calculate the error loss between the sample classification prediction result and the corresponding fault type label. Based on the error loss, the network parameters of the image feature extraction network, voiceprint feature extraction network, and classification prediction network are adjusted. The loss function may be a cross-entropy loss function, and the network parameters may be adjusted during backpropagation using a gradient descent method.
[0110] Optionally, in an embodiment of the present invention, a normal vehicle-mounted image sequence and a corresponding normal engine soundprint sequence of each vehicle type after the engine is started under normal engine conditions, as well as an abnormal vehicle-mounted image sequence and a corresponding abnormal engine soundprint sequence of the above-mentioned corresponding vehicle after the engine is started under abnormal engine conditions can also be obtained as a second data set. In the above-mentioned second data set, one normal sample vehicle-mounted image sequence corresponds to one normal sample engine soundprint sequence, and one abnormal sample vehicle-mounted image sequence corresponds to one abnormal sample engine soundprint sequence. The above-mentioned second data set corresponds to one vehicle type, and the samples in the above-mentioned second data set are not repeated with those in the above-mentioned first data set. When initializing the user vehicle, the vehicle-mounted image sequence and the corresponding engine soundprint sequence of the user vehicle after the engine is started under normal engine conditions are obtained. The corresponding second data set is matched according to the vehicle-mounted image sequence and the corresponding engine soundprint sequence of the user vehicle after the engine is started under normal engine conditions, and the network parameters of the above-mentioned image feature extraction network, soundprint feature extraction network, and classification prediction network are fine-tuned according to the above-mentioned corresponding second data set. The above-mentioned normal sample engine soundprint sequence includes time domain information and frequency domain information. Specifically, the above-mentioned normal sample engine soundprint sequence can be obtained through Figure 2 The frequency domain information in the above normal sample engine voiceprint sequence can be obtained by Figure 3 The conversion method of the embodiment is used to convert.
[0111] The above-mentioned initialization of the vehicle can be automatically performed when the vehicle leaves the factory, or can be manually initialized by relevant staff when the user purchases the vehicle.
[0112] During fine-tuning of the network parameters of the aforementioned image feature extraction network, voiceprint feature extraction network, and classification prediction network using the second dataset, the inclusion of vehicle image sequences and engine soundprint sequences of vehicles with normal engines allows for more comprehensive fine-tuning of network parameters. Furthermore, the inclusion of vehicle image sequences and engine soundprint sequences of vehicles with normal engines in the second dataset allows for indexing of specific vehicles to the corresponding second dataset. This is primarily because new car engines are generally in normal condition, rarely exhibiting abnormalities, and only the vehicle image sequences and engine soundprint sequences under normal conditions are available. Furthermore, due to spatiotemporal relationships, vehicles from the same batch may be sold at different times and locations. Fine-tuning the network parameters of the image feature extraction network, voiceprint feature extraction network, and classification prediction network by matching the corresponding second datasets to vehicles sold at different times and locations makes the vehicle anomaly detection model more targeted, thereby improving the accuracy of vehicle anomaly detection.
[0113] See Figure 5 , Figure 5 FIG. 1 is a schematic structural diagram of a vehicle engine abnormality detection device provided by an embodiment of the present invention. Figure 5 As shown, the device includes:
[0114] A first acquisition module 501 is configured to acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint. The vehicle image sequence and the engine soundprint sequence are acquired simultaneously.
[0115] A first extraction module 502 is configured to extract features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood;
[0116] The second extraction module 503 is configured to extract features from the engine voiceprint sequence using a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature of the engine;
[0117] A first fusion module 504 is configured to fuse the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine;
[0118] The detection module 505 is used to perform classification prediction on the fusion features through a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
[0119] Optional, such as Figure 6 As shown, the first acquisition module 501 includes:
[0120] The acquisition submodule 5011 is used to acquire the first engine voiceprint sequence after the engine is started;
[0121] a noise reduction submodule 5012 for reducing noise on the first engine soundprint sequence to obtain a second engine soundprint sequence;
[0122] The conversion submodule 5013 is configured to convert the second engine voiceprint sequence from time domain information to frequency domain information, and obtain an engine voiceprint sequence according to the time domain information and the frequency domain information.
[0123] Optional, such as Figure 7 As shown, the conversion submodule 5013 includes:
[0124] The framing unit 50131 is configured to perform framing processing on the second engine voiceprint sequence to obtain a frame-processed second engine voiceprint sequence;
[0125] A windowing unit 50132 is configured to perform windowing processing on the second engine voiceprint sequence after the frame processing to obtain a windowed second engine voiceprint sequence;
[0126] The transformation unit 50133 is configured to perform fast Fourier transform point processing on the windowed second engine voiceprint sequence, so as to transform the second engine voiceprint sequence from time domain information to frequency domain information.
[0127] Optional, such as Figure 8 As shown, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine have the same feature dimension. The fusion module 504 includes:
[0128] A first fusion submodule 5041 is configured to perform a dot product between the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine to obtain a first fusion feature;
[0129] The second fusion submodule 5042 is configured to accumulate the first fusion features to obtain a second fusion feature as a fusion feature of the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine.
[0130] Optional, such as Figure 9 As shown, the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are trained using the same data set. The device also includes:
[0131] A construction module 506 is configured to construct a first data set, the first data set including a sample vehicle image sequence, a sample engine soundprint sequence, and corresponding fault type labels, wherein the sample vehicle image sequence is a vehicle image sequence after the engine is started in an engine abnormality condition, the sample engine soundprint sequence is an engine soundprint sequence after the engine is started in an engine abnormality condition, and each sample vehicle image sequence corresponds to one sample engine soundprint sequence;
[0132] The training module 507 is used to jointly train the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network using the first data set.
[0133] Optional, such as Figure 10 As shown, the training module 507 includes:
[0134] A first extraction submodule 5071 is configured to extract features from the sample vehicle image sequence using a to-be-trained image feature extraction network to obtain sample spatiotemporal features of the engine hood;
[0135] The second extraction submodule 5072 is configured to extract features from the sample engine voiceprint sequence using a voiceprint feature extraction network to be trained, thereby obtaining a sample time-frequency voiceprint feature of the engine;
[0136] The third fusion submodule 5073 is configured to fuse the spatiotemporal features of the engine hood with the time-frequency voiceprint features of the engine to obtain a fusion feature of the sample spatiotemporal features of the engine hood and the sample time-frequency voiceprint features of the engine;
[0137] The classification submodule is used to perform sample classification prediction on the fusion features through the classification prediction network to be trained to obtain the sample classification prediction result, calculate the error loss between the sample classification prediction result and the corresponding fault type label through a preset loss function, and adjust the network parameters of the image feature extraction network, voiceprint feature extraction network, and classification prediction network according to the error loss.
[0138] Optional, such as Figure 11 As shown, the device also includes:
[0139] A second acquisition module 508 is configured to acquire, for each vehicle type, a normal vehicle-mounted image sequence and a corresponding normal engine soundprint sequence after engine startup under normal engine conditions, and an abnormal vehicle-mounted image sequence and a corresponding abnormal engine soundprint sequence after engine startup under abnormal engine conditions as a second data set, wherein in the second data set, one normal sample vehicle-mounted image sequence corresponds to one normal sample engine soundprint sequence, and one abnormal sample vehicle-mounted image sequence corresponds to one abnormal sample engine soundprint sequence. One second data set corresponds to one vehicle type, and samples in the second data set are not repeated with those in the first data set.
[0140] The third acquisition module 509 is used to acquire a vehicle image sequence and a corresponding engine soundprint sequence of the user's vehicle after the engine is started under normal conditions when the user's vehicle is initialized;
[0141] The fine-tuning module 510 is used to match the corresponding second data set according to the vehicle image sequence of the user vehicle after the engine is started under normal engine conditions and the corresponding engine soundprint sequence, and fine-tune the network parameters of the image feature extraction network, the soundprint feature extraction network, and the classification prediction network according to the corresponding second data set.
[0142] It should be noted that the vehicle engine anomaly detection device provided in the embodiment of the present invention can be applied to devices such as mobile phones, monitors, computers, servers, etc. that can perform vehicle engine anomaly detection.
[0143] The vehicle engine abnormality detection device provided in the embodiment of the present invention can implement each process implemented by the vehicle engine abnormality detection method in the above method embodiment and can achieve the same beneficial effects. To avoid repetition, it will not be described here.
[0144] See also Figure 12 , Figure 12 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, such as Figure 12 As shown, it includes: a memory 1202, a processor 1201, and a computer program stored in the memory 1202 and executable on the processor 1201, wherein:
[0145] The processor 1201 is configured to call the computer program stored in the memory 1202 and execute the following steps:
[0146] Acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously;
[0147] Extracting features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood;
[0148] Extracting features of the engine voiceprint sequence through a preset voiceprint feature extraction network to obtain the engine's time-frequency voiceprint features;
[0149] fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine;
[0150] The fusion features are classified and predicted by a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
[0151] Optionally, the processor 1201 executes the step of obtaining the engine soundprint sequence after the engine is started, including:
[0152] Obtain the first engine soundprint sequence after the engine is started;
[0153] De-noising the first engine soundprint sequence to obtain a second engine soundprint sequence;
[0154] The second engine soundprint sequence is converted from time domain information to frequency domain information, and an engine soundprint sequence is obtained according to the time domain information and the frequency domain information.
[0155] Optionally, the processor 1201 converts the second engine voiceprint sequence from time domain information to frequency domain information, including:
[0156] performing frame processing on the second engine voiceprint sequence to obtain a frame-processed second engine voiceprint sequence;
[0157] performing windowing processing on the second engine voiceprint sequence after the frame processing to obtain a windowed second engine voiceprint sequence;
[0158] Performing fast Fourier transform point processing on the windowed second engine voiceprint sequence to convert the second engine voiceprint sequence from time domain information to frequency domain information.
[0159] Optionally, the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine have the same feature dimension, and the processor 1201 performs the fusion of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine to obtain the fusion features of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine, including:
[0160] Performing a dot product on the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine to obtain a first fusion feature;
[0161] The first fusion features are accumulated to obtain a second fusion feature as a fusion feature of the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine.
[0162] Optionally, the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are trained using the same data set, and the training steps performed by the processor 1201 include:
[0163] Constructing a first data set, the first data set including a sample vehicle image sequence, a sample engine soundprint sequence, and corresponding fault type labels, wherein the sample vehicle image sequence is a vehicle image sequence after the engine is started in an engine abnormality condition, the sample engine soundprint sequence is an engine soundprint sequence after the engine is started in an engine abnormality condition, and one sample vehicle image sequence corresponds to one sample engine soundprint sequence;
[0164] The image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are jointly trained using the first data set.
[0165] Optionally, the processor 1201 performs the joint training of the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network using the data set, including:
[0166] Extracting features from the sample vehicle image sequence using the image feature extraction network to be trained to obtain sample spatiotemporal features of the engine hood;
[0167] Extracting features from the sample engine voiceprint sequence using a voiceprint feature extraction network to be trained to obtain sample time-frequency voiceprint features of the engine;
[0168] fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the sample spatiotemporal features of the engine hood and the sample time-frequency soundprint features of the engine;
[0169] The fusion feature is subjected to sample classification prediction through the classification prediction network to be trained to obtain the sample classification prediction result. The error loss between the sample classification prediction result and the corresponding fault type label is calculated through a preset loss function. The network parameters of the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are adjusted according to the error loss.
[0170] Optionally, after the training is completed, the method executed by the processor 1201 further includes:
[0171] Acquire a normal vehicle image sequence and a corresponding normal engine soundprint sequence for each vehicle type after engine startup under normal engine conditions, and an abnormal vehicle image sequence and a corresponding abnormal engine soundprint sequence for the corresponding vehicle after engine startup under abnormal engine conditions as a second data set, wherein in the second data set, one normal sample vehicle image sequence corresponds to one normal sample engine soundprint sequence, and one abnormal sample vehicle image sequence corresponds to one abnormal sample engine soundprint sequence, one second data set corresponds to one vehicle type, and samples in the second data set are not repeated with those in the first data set;
[0172] When initializing the user's vehicle, obtain a vehicle image sequence and a corresponding engine soundprint sequence of the user's vehicle after the engine is started under normal conditions;
[0173] A corresponding second data set is matched based on the on-board image sequence of the user vehicle after the engine is started under normal engine conditions and the corresponding engine soundprint sequence, and the network parameters of the image feature extraction network, the soundprint feature extraction network, and the classification prediction network are fine-tuned based on the corresponding second data set.
[0174] It should be noted that the above-mentioned electronic devices may be mobile phones, monitors, computers, servers and other devices that can be used to detect vehicle engine anomalies.
[0175] The electronic device provided by the embodiment of the present invention can implement each process implemented by the vehicle engine abnormality detection method in the above method embodiment and can achieve the same beneficial effects. To avoid repetition, it will not be described here.
[0176] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the vehicle engine abnormality detection method provided in the embodiment of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0177] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0178] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A vehicle engine abnormality detection method, characterized in that: The following steps are involved: Acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously; Extracting features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood; Extracting features from the engine soundprint sequence using a preset soundprint feature extraction network to obtain a time-frequency soundprint feature of the engine, wherein the spatiotemporal features of the engine hood and the time-frequency soundprint feature of the engine have the same feature dimension; Performing a dot product of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine to obtain a first fusion feature; accumulating the first fusion features to obtain a second fusion feature as a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine; The fusion features are classified and predicted by a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
2. The method according to claim 1, wherein The step of obtaining the engine soundprint sequence after the engine is started includes: Obtain the first engine soundprint sequence after the engine is started; De-noising the first engine soundprint sequence to obtain a second engine soundprint sequence; The second engine voiceprint sequence is converted from time domain information to frequency domain information, and an engine voiceprint sequence is obtained according to the time domain information and the frequency domain information of the second engine voiceprint sequence.
3. The method according to claim 2, wherein The converting the second engine soundprint sequence from time domain information to frequency domain information includes: performing frame processing on the second engine voiceprint sequence to obtain a frame-processed second engine voiceprint sequence; performing windowing processing on the second engine voiceprint sequence after the frame processing to obtain a windowed second engine voiceprint sequence; Performing fast Fourier transform point processing on the windowed second engine voiceprint sequence to convert the second engine voiceprint sequence from time domain information to frequency domain information.
4. The method according to any one of claims 1 to 3, characterized in that The image feature extraction network, voiceprint feature extraction network, and classification prediction network are trained using the same data set, and the training steps include: Constructing a first data set, the first data set including a sample vehicle image sequence, a sample engine soundprint sequence, and corresponding fault type labels, wherein the sample vehicle image sequence is a vehicle image sequence after the engine is started in an engine abnormality condition, the sample engine soundprint sequence is an engine soundprint sequence after the engine is started in an engine abnormality condition, and one sample vehicle image sequence corresponds to one sample engine soundprint sequence; The image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are jointly trained using the first data set.
5. The method according to claim 4, wherein The image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are jointly trained using the data set, including: Extracting features from the sample vehicle image sequence using the image feature extraction network to be trained to obtain sample spatiotemporal features of the engine hood; Extracting features from the sample engine voiceprint sequence using a voiceprint feature extraction network to be trained to obtain sample time-frequency voiceprint features of the engine; fusing the spatiotemporal features of the engine hood with the time-frequency soundprint features of the engine to obtain a fusion feature of the sample spatiotemporal features of the engine hood and the sample time-frequency soundprint features of the engine; The fusion feature is subjected to sample classification prediction through the classification prediction network to be trained to obtain the sample classification prediction result. The error loss between the sample classification prediction result and the corresponding fault type label is calculated through a preset loss function. The network parameters of the image feature extraction network, the voiceprint feature extraction network, and the classification prediction network are adjusted according to the error loss.
6. The method according to claim 5, wherein After the training is completed, the method further includes: Acquire a normal vehicle image sequence and a corresponding normal engine soundprint sequence for each vehicle type after engine startup under normal engine conditions, and an abnormal vehicle image sequence and a corresponding abnormal engine soundprint sequence for the corresponding vehicle after engine startup under abnormal engine conditions as a second data set, wherein in the second data set, one normal sample vehicle image sequence corresponds to one normal sample engine soundprint sequence, and one abnormal sample vehicle image sequence corresponds to one abnormal sample engine soundprint sequence, one second data set corresponds to one vehicle type, and samples in the second data set are not repeated with those in the first data set; When initializing the user's vehicle, obtain a vehicle image sequence and a corresponding engine soundprint sequence of the user's vehicle after the engine is started under normal conditions; A corresponding second data set is matched based on the on-board image sequence of the user vehicle after the engine is started under normal engine conditions and the corresponding engine soundprint sequence, and the network parameters of the image feature extraction network, the soundprint feature extraction network, and the classification prediction network are fine-tuned based on the corresponding second data set.
7. A vehicle engine abnormality detection device, characterized in that: The device comprises: A first acquisition module is configured to acquire a vehicle image sequence and an engine soundprint sequence after the engine is started, wherein the vehicle image sequence includes an engine hood image, and the engine soundprint sequence includes time domain information and frequency domain information of the engine soundprint, and the vehicle image sequence and the engine soundprint sequence are acquired simultaneously; a first extraction module, configured to extract features from the vehicle image sequence using a preset image feature extraction network to obtain spatiotemporal features of the engine hood; a second extraction module, configured to extract features from the engine soundprint sequence using a preset soundprint feature extraction network to obtain a time-frequency soundprint feature of the engine, wherein the spatiotemporal feature of the engine hood and the time-frequency soundprint feature of the engine have the same feature dimension; a first fusion module, configured to perform a dot product of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine to obtain a first fusion feature; and to accumulate the first fusion features to obtain a second fusion feature as a fusion feature of the spatiotemporal features of the engine hood and the time-frequency soundprint features of the engine; The detection module is used to perform classification prediction on the fusion features through a preset classification prediction network to obtain a classification prediction result as the vehicle engine abnormality detection result.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the vehicle engine abnormality detection method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the vehicle engine abnormality detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Circuit breaker fault type judgment method and device, electronic equipment and storage medium
CN110926782A
Engine surge fault prediction system and method based on fusion neural network model
CN112131673A