Vehicle fault prediction method and device, vehicle, and storage medium
By extracting operational audio features that do not contain voice content from vehicle fault prediction, and activating only sensors of abnormal components, the problems of high energy consumption and privacy leakage are solved, achieving efficient and accurate fault prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-07-03
AI Technical Summary
Existing vehicle fault prediction technologies have high energy consumption and privacy risks, especially when continuously collecting vehicle operating audio data, which affects battery life and privacy security.
The vehicle-mounted terminal extracts operating audio features that do not contain the speech content of the speaker, filters out selected sensors, activates only the sensors related to abnormal components, and extracts key predictive features and uploads them to the server for fault prediction.
It reduces the risk of user privacy leaks, saves vehicle energy consumption, reduces data transmission volume, and improves the accuracy of fault prediction and network efficiency.
Smart Images

Figure CN122340153A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent connected vehicles, and in particular to a vehicle fault prediction method and device, vehicle, and storage medium. Background Technology
[0002] In related technologies, vehicle fault prediction is accomplished using vehicle operating audio data. Specifically, the vehicle operating audio data is input into a pre-trained neural network model for fault prediction. The model identifies malfunctioning parts by comparing the vehicle operating audio data with abnormal component audio data, and then obtains the monitoring parameters of the malfunctioning parts to complete the fault prediction, making vehicle fault prediction more intelligent. However, during the fault prediction process, continuously collecting monitoring parameters of various vehicle components increases vehicle energy consumption and affects the vehicle's driving range. Furthermore, the vehicle operating audio data involves audio content recognition and processing, posing a risk of privacy leakage.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main objective of this application is to propose a vehicle fault prediction method and device, vehicle and storage medium, which aims to reduce vehicle energy consumption and reduce the risk of privacy leakage.
[0005] To achieve the above objectives, one aspect of this application proposes a vehicle fault prediction method, executed by an on-board terminal, which is communicatively connected to a server. The method includes the following steps: Obtain the raw operating audio data of the target vehicle; Selected running audio features are extracted from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; Selected sensors are selected from preset candidate sensors based on preset abnormal audio characteristics and the selected operating audio characteristics; wherein, the selected sensors are used to collect component monitoring data associated with abnormal components; Activate the selected sensor and receive the component monitoring data collected by the selected sensor; Predictive key features are extracted from the selected operating audio features and the component monitoring data, and the predicted key features are sent to the server. The server also receives fault prediction information from the server. The server performs fault prediction on the target vehicle based on the predicted key features to obtain the fault prediction information.
[0006] In some embodiments, extracting selected runtime audio features from the original runtime audio data includes: The original running audio data is preprocessed to obtain an original running audio frame sequence; wherein the original running audio frame sequence includes at least one original running audio frame; Feature extraction is performed on each of the original running audio frames to obtain the original running audio features; Speech activity detection is performed on the original running audio features to obtain speech activity detection information; Non-speech frames are filtered out from the original running audio frames based on the speech activity detection information; The original running audio features are filtered based on the non-speech frames to obtain the selected running audio features.
[0007] In some embodiments, the step of selecting a selected sensor from a preset pool of candidate sensors based on preset abnormal audio features and the selected operating audio features includes: Obtain candidate anomaly categories for the preset abnormal audio features; The similarity between the preset abnormal audio features and the selected running audio features is calculated to obtain the audio similarity. Selected anomaly categories are filtered from the candidate anomaly categories based on the audio similarity. Selected trigger signals are selected from preset trigger signals based on preset voice commands and audio similarity. The selected sensor is selected from the candidate sensors based on the selected anomaly category and the selected trigger signal.
[0008] In some embodiments, extracting predictive key features from the selected running audio features and the component monitoring data includes: Feature extraction is performed on the component monitoring data to obtain component monitoring features; The first feature extraction model is used to extract the key features for fault prediction from the component monitoring features and the selected operating audio features; wherein, the first feature extraction model is adjusted according to the model parameters of the second feature extraction model on the server.
[0009] In some embodiments, the step of extracting key predictive features for fault prediction from the component monitoring features and the selected operating audio features using a preset first feature extraction model includes: The monitoring features of the component are subjected to time-domain feature extraction to obtain the first time-domain feature; Temporal feature extraction is performed on the selected audio features to obtain the second temporal feature; Frequency domain features are extracted from the selected audio features to obtain the first frequency domain features; The first time-domain feature, the second time-domain feature, and the first frequency-domain feature are concatenated to obtain the concatenated feature; The first feature extraction model is used to perform feature inference on the spliced features to obtain the predicted key features.
[0010] In some embodiments, after extracting the key features for fault prediction from the component monitoring features and the selected operating audio features using a preset first feature extraction model, the method further includes: Based on the predicted key features, the component monitoring data and the original operating audio data are erased, and the erase detection is performed on the component monitoring data and the original operating audio data to obtain erase detection information; If the erasure detection information indicates that the component monitoring data and the original operating audio data have been completely erased, erasure proof information is generated.
[0011] In some embodiments, the server performs fault prediction on the target vehicle based on the predicted key features to obtain the fault prediction information, including: The server obtains the vehicle model information of the target vehicle; The server selects a chosen anomaly knowledge graph from a preset candidate anomaly knowledge graph based on the vehicle model information; wherein, the chosen anomaly knowledge graph is constructed based on the anomaly features of the anomaly components in the target vehicle; The server performs fault prediction on the target vehicle using the selected anomaly knowledge graph and the predicted key features, thereby obtaining the fault prediction information.
[0012] To achieve the above objectives, another aspect of this application provides a vehicle fault prediction device, which is disposed on an in-vehicle terminal that is communicatively connected to a server. The device includes: The data acquisition module is used to acquire the raw operating audio data of the target vehicle; The feature extraction module is used to extract selected running audio features from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; The sensor screening module is used to select a sensor from a preset candidate sensor based on preset abnormal audio characteristics and the selected operating audio characteristics; wherein, the selected sensor is used to collect component monitoring data of abnormal component operation; An activation module is used to perform an activation operation on the selected sensor and receive the component monitoring data collected by the selected sensor. The sending module is used to extract predictive key features from the selected operating audio features and the component monitoring data, send the predictive key features to the server, and receive fault prediction information sent by the server; wherein, the server obtains the fault prediction information by performing fault prediction on the target vehicle based on the predictive key features.
[0013] To achieve the above objectives, another aspect of this application provides a vehicle including an on-board terminal and at least one sensor; wherein the on-board terminal is communicatively connected to a server, and the on-board terminal is used to execute the vehicle fault prediction method as described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a vehicle fault prediction method and apparatus, a vehicle, and a storage medium. This solution predicts faults in a target vehicle in advance, and the fault prediction process utilizes the operating audio characteristics of the target vehicle during operation. These operating audio characteristics do not include the speech content characteristics of the speaker, reducing the risk of user privacy leakage. Once the operating audio characteristics are determined, sensors are activated selectively in conjunction with preset abnormal audio characteristics and the operating audio characteristics. This eliminates the need to activate all sensors, accurately collecting component monitoring data of abnormal parts while saving vehicle energy consumption. Furthermore, before the server performs fault prediction, the on-board terminal extracts key prediction features from the component monitoring data and operating audio characteristics, uploading only these key features to the server to complete the fault prediction. This not only saves data transmission volume and reduces the probability of network congestion but also ensures the accuracy of the target vehicle fault prediction. Attached Figure Description
[0016] Figure 1 This is a system architecture diagram of the vehicle fault prediction method provided in the embodiments of this application; Figure 2 This is a flowchart of the vehicle fault prediction method provided in the embodiments of this application; Figure 3 yes Figure 2 The flowchart of step S202 in the document; Figure 4 This is a schematic diagram of the sensor arrangement on the target vehicle in this application embodiment; Figure 5 yes Figure 2 The flowchart of step S203 in the process; Figure 6This is a schematic diagram illustrating the mapping relationship between abnormal audio features and abnormal categories in an embodiment of this application; Figure 7 This is a flowchart of a vehicle fault prediction method provided in another embodiment of this application; Figure 8 yes Figure 7 The flowchart of step S702 in the process; Figure 9 This is a flowchart of feature distillation in the vehicle fault prediction method provided in this application embodiment; Figure 10 This is a flowchart of a vehicle fault prediction method provided in another embodiment of this application; Figure 11 This is a flowchart of a vehicle fault prediction method provided in another embodiment of this application; Figure 12 This is a schematic diagram of the vehicle fault prediction device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0022] 1) Audio fingerprinting technology is a core technology that generates identifiers by extracting unique digital features from audio signals, which are then used to identify or match database content. Its identification process is unaffected by audio format, encoding, or compression methods, and it replaces traditional hashing or metadata comparison with high-precision feature matching.
[0023] 2) Mel-frequency cepstral coefficients (MFCC) are a speech feature extraction method designed based on human auditory characteristics. By simulating the nonlinear perception of frequency by the human ear, speech signals are converted into low-dimensional feature representations and are widely used in speech recognition, speaker recognition and speech synthesis.
[0024] 3) Voice Activity Detection (VAD) is a technique used to detect the presence of human speech in an audio stream. It is widely used in speech recognition, communication and audio processing.
[0025] 4) Dynamic Time Warping (DTW) algorithm is used to measure the similarity between two time series of different lengths. It stretches or shrinks (compresses) the unknown quantity until it matches the length of the reference template. In this process, the unknown sequence will be distorted or bent so that its features correspond to the standard pattern.
[0026] 5) Feature distillation is a technique that transfers the feature representations of complex models to simpler models, aiming to improve model performance without significantly increasing computational costs.
[0027] In related technologies, fault detection technology for intelligent connected vehicles is shifting from traditional rule-based diagnostic models to AI-based intelligent diagnostic models. Acoustic signal-based vehicle fault detection technology, due to its non-invasiveness, real-time performance, and cost-effectiveness, is gradually becoming an important means of vehicle health monitoring. However, existing technical solutions have many shortcomings in areas such as privacy protection, energy efficiency, data processing, and early warning.
[0028] For example, vehicle fault detection first acquires data from the vehicle's vehicle control unit (VCU) to parse fault information for various vehicle components. It then collects audio data of the vehicle's operation during driving and inputs it into a trained neural network model for training and analysis. The model outputs a classification report to predict the faults of each component. By comparing this predicted fault information with the individual component fault information, the predicted faults for each component are determined, thus achieving vehicle fault prediction. However, during the vehicle fault prediction process, the uploaded audio data may contain the speech content of the speaker, posing a privacy risk. Furthermore, the continuous collection of VCU data and audio data can lead to excessive energy consumption, affecting the vehicle's driving range. Additionally, directly uploading the collected audio data and VCU data to a server for analysis can easily cause network congestion and incur high storage costs.
[0029] In view of this, this application provides a vehicle fault prediction method and apparatus, a vehicle, and a storage medium. This solution predicts faults in the target vehicle in advance, and the fault prediction process uses the operating audio features of the target vehicle during operation. These operating audio features do not include the speech content features of the speaker, reducing the risk of user privacy leakage. Once the operating audio features are determined, sensors are activated selectively based on preset abnormal audio features and the operating audio features. This eliminates the need to activate all sensors, accurately collecting component monitoring data of abnormal parts while saving vehicle energy consumption. Furthermore, before the server performs fault prediction, the on-board terminal extracts key prediction features from the component monitoring data and operating audio features, uploading only these key features to the server to complete the fault prediction. This not only saves data transmission volume and reduces the probability of network congestion but also ensures the accuracy of the target vehicle fault prediction.
[0030] The vehicle fault prediction method provided in this application relates to the field of intelligent connected vehicle technology. The vehicle fault prediction method provided in this application can be applied to a vehicle's in-vehicle terminal, or it can be software running on the in-vehicle terminal. In some embodiments, the in-vehicle terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or other in-vehicle terminal, but is not limited to these. The in-vehicle terminal communicates with a server. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application that implements the vehicle fault prediction method, but is not limited to the above forms.
[0031] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0032] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0033] Figure 1 This is a system architecture diagram of the vehicle fault prediction method provided in the embodiments of this application. It includes an in-vehicle terminal and a server. The server is a cloud server, and a cloud subsystem is set on the cloud server. An in-vehicle subsystem is set on the in-vehicle terminal.
[0034] The vehicle-mounted subsystem is used to collect the operating audio features of the target vehicle that do not contain speech during operation. It activates the sensors by using the operating audio features and receives the component monitoring data collected by the sensors. It selects the key predictive features for fault prediction from the component monitoring data and the operating audio features, and then sends the key predictive features to the cloud subsystem, which then completes the fault prediction.
[0035] The cloud subsystem is used to receive predicted key features and complete the fault prediction of the target vehicle based on the predicted key features, so as to realize early fault prediction, allow the target vehicle to perform abnormal repairs in advance, and reduce the vehicle failure rate.
[0036] In some embodiments, the vehicle subsystem adopts a three-layer architecture design, specifically including a continuous monitoring layer, a dynamic acquisition layer, and an edge processing layer. The continuous monitoring layer includes a low-power microphone array, an acoustic scene perception chip, and an abnormal audio memory. The low-power microphone array uses a 4-channel microphone with a sampling frequency of 8kHz and power consumption <1mW. The acoustic scene perception chip is a dedicated ASIC chip based on the RISC-V architecture, integrating an FFT accelerator and a lightweight network inference unit, with an operating power consumption <10mW. The abnormal audio memory is a NOR Flash memory with a pre-built feature library of preset abnormal audio features. The dynamic acquisition layer includes a main controller, a sensor network, a directional microphone module, and a data buffer. The main controller is an automotive-grade MCU responsible for sensor scheduling and data buffer management. The sensor network uses distributed MEMS accelerometers with a sampling rate of 5kHz, connected via a CAN bus. The directional microphone module is a rotatable microphone supporting beamforming, with a sampling frequency of 16kHz. Different sampling frequencies can be selected as needed; this embodiment does not restrict the selection of microphones. The data buffer uses a circular write mechanism to complete data caching. The edge processing layer includes a neural network accelerator, a feature distillation model memory, and a hardware security module. The neural network accelerator is used for inference of the first feature extraction model, the feature distillation model memory is used to store the model parameters of the first feature extraction model, and the hardware security module is responsible for data encryption and secure erasure, saving data storage space.
[0037] In some embodiments, the cloud subsystem includes a feature vector parser, a fault prediction model, an audio feature update server, and a repair service interface. The feature vector parser is used to parse the predicted key features uploaded by the vehicle subsystem and convert the predicted key features into a fault prediction format. The fault prediction model is a large-scale diagnostic model based on the Transformer architecture, combined with a vehicle model knowledge graph to complete fault prediction. It should be noted that the fault prediction model can also be other deep learning models or large language models; this embodiment does not limit the specific type of fault prediction model. The audio feature update server continuously optimizes the feature library of preset abnormal audio features based on feedback from fault prediction information, making abnormal audio analysis more accurate. The repair service interface is used to connect with the after-sales service of the target vehicle to provide repair or after-sales service to the target vehicle.
[0038] In summary, the embodiments of this application provide a system architecture diagram for an in-vehicle fault prediction method, which combines the in-vehicle subsystem and the cloud subsystem to complete vehicle fault prediction, making vehicle fault prediction more efficient, and improving privacy protection, energy efficiency, data transmission and early warning.
[0039] Figure 2 This is an optional flowchart of the vehicle fault prediction method provided in the embodiments of this application. Figure 2 The method may include, but is not limited to, steps S201 to S205.
[0040] Step S201: Obtain the original operating audio data of the target vehicle; Step S202: Extract selected running audio features from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; Step S203: Select a sensor from the preset candidate sensors based on the preset abnormal audio characteristics and the selected operating audio characteristics; wherein, the selected sensor is used to collect component monitoring data associated with the abnormal component; Step S204: Activate the selected sensor and receive the component monitoring data collected by the selected sensor; Step S205: Extract the predicted key features from the selected operating audio features and component monitoring data, send the predicted key features to the server, and receive the fault prediction information sent by the server based on the predicted key features.
[0041] Steps S201 to S205 of this embodiment involve acquiring the original operating audio data of the target vehicle, extracting selected operating audio features from the original operating audio data that do not contain the speech content of the speaker, and ensuring that the selected operating audio features can characterize the abnormal noise characteristics of the target vehicle during operation without disclosing the privacy of the users inside the target vehicle, thus improving privacy and security. Then, sensors are activated specifically based on the selected operating audio features and preset abnormal audio features, achieving intelligent sensor triggering. This eliminates the need for sensors to be running for extended periods; they only activate and collect component monitoring data when a component malfunctions, ensuring accurate collection of component monitoring data during malfunctions and saving sensor energy. Furthermore, before the server completes fault prediction, the vehicle terminal extracts key prediction features from the component monitoring data and the selected operating audio features, and uploads these key prediction features to the server to complete fault prediction. This not only saves data transmission volume and reduces network congestion but also ensures accurate fault prediction. In summary, this application embodiment, while protecting user privacy, uses abnormal noises during vehicle operation to locate abnormal components, then specifically triggers sensors to collect sensor data related to the abnormal components, and combines the sensor data and key features in the abnormal noises to complete fault prediction. This not only saves vehicle energy consumption and reduces data transmission, but also provides early fault warnings and reduces the failure rate.
[0042] In step S201 of some embodiments, the target vehicle is the vehicle monitored by the vehicle-mounted terminal. It should be noted that the vehicle-mounted terminal can be a terminal configured on the target vehicle, a mobile terminal authorized by the target vehicle, or a terminal configured on other vehicles. This embodiment does not limit the specific type of vehicle-mounted terminal. This embodiment provides various settings for the vehicle-mounted terminal to facilitate fault monitoring of the target vehicle by vehicle owners or other users, improving the convenience of target vehicle fault monitoring. For example, with the increase in car rental ride-hailing vehicles, drivers do not constantly activate vehicle fault prediction to save vehicle energy, leading to an increased fault rate. This embodiment sets up a mobile terminal to complete vehicle fault prediction, allowing car rental operators to remotely complete fault prediction for the target vehicle, not only saving energy but also reducing the fault rate.
[0043] In some embodiments, the raw operating audio data is the in-vehicle ambient audio during the operation of the target vehicle, specifically including driving sound data and component operating sound data during the operation of the target vehicle. It should be noted that in this embodiment, the raw operating audio data is obtained by collecting in-vehicle audio during the operation of the target vehicle using a low-power microphone array at a preset sampling frequency, and the preset sampling frequency is automatically set by the user.
[0044] In step S202 of some embodiments, the selected running audio features, also known as non-speech segment features, retain only the audio features of non-speech frames. It should be noted that after the selected running audio features are extracted, the original running audio data is directly erased from the data buffer, and only the selected running audio features are uploaded to the server. Even if an attacker intercepts the selected running audio features during transmission, they will not be able to reconstruct the speech content, thus improving the user's privacy and security within the target vehicle.
[0045] Please see Figure 3 In some embodiments, step S202 may include, but is not limited to, steps S301 to S305: Step S301: Preprocess the original running audio data to obtain the original running audio frame sequence; wherein the original running audio frame sequence includes at least one original running audio frame; Step S302: Extract features from each original running audio frame to obtain the original running audio features; Step S303: Perform speech activity detection on the original running audio features to obtain speech activity detection information; Step S304: Filter out non-speech frames from the original running audio frames based on the speech activity detection information; Step S305: Filter the original running audio features based on non-speech frames to obtain selected running audio features.
[0046] In step S301 of some embodiments, preprocessing includes at least one of the following: pre-emphasis filtering, framing processing, and windowing processing. Pre-emphasis filtering emphasizes the high-frequency components of the input raw audio data to remove the influence of lip radiation and increase the high-frequency resolution of the speech. Framing processing divides the raw audio data into multiple short time segments to facilitate subsequent analysis and processing. Windowing processing uses short-time Fourier transform and other analytical steps to smooth the raw audio data; specifically, it applies a window to each frame to reduce spectral leakage and aliasing.
[0047] Specifically, the original running audio data is defined as x(n). The preprocessing process is as follows: First, the original running audio data is pre-emphasized and filtered to obtain the first running audio data: y(n) = x(n) - α·x(n-1), where α = 0.97. The first running audio data is then segmented into frames to obtain the original running audio frames, and each original running audio frame is windowed: w(n) = 0.54 - 0.46·cos(2πn / N). Then, the windowed original running audio frames are concatenated into a sequence of original running audio frames.
[0048] In step S302 of some embodiments, feature extraction mainly involves extracting spectral features from the original running audio frame. Specifically, a Fast Fourier Transform is first performed on the original running audio frame to obtain the original running audio spectrum, and then the power spectrum is extracted from the original running audio spectrum. The original running audio spectrum is then filtered using a filter and the power spectrum, and finally, a Discrete Cosine Transform is performed on the original running audio spectrum to obtain the original running audio features. It should be noted that the original running audio features are 32-dimensional Mel-frequency cepstral coefficient features, defined as MFCC features.
[0049] In step S303 of some embodiments, speech activity detection is to evaluate whether there is speech content of a speaking object in the original running audio frame. Specifically, it can identify the speech content in the original running audio frame. If there is speech content, it is determined that the speech activity detection information indicates that there is speech activity in the original running audio frame; otherwise, it is determined that the speech activity detection information indicates that there is no speech activity in the original running audio frame.
[0050] In step S304 of some embodiments, the non-speech frame is the original running audio frame that does not have speech activity. When the speech activity detection information indicates that the original running audio frame has speech activity, the original running audio frame is discarded, and only the original running audio frame without speech activity is left as the non-speech frame.
[0051] In step S305 of some embodiments, the original running audio features corresponding to the non-speech frames are selected as running audio features.
[0052] In steps S301 to S305 of this embodiment, when extracting selected operating audio features from the original operating audio data, operating audio frames with voice activity are actively filtered out to ensure that the system cannot collect and process voice content from a technical perspective. Fault prediction is completed only through selected operating audio features without voice content, which improves user privacy and security and ensures the accuracy of fault prediction.
[0053] In step S203 of some embodiments, the selected sensor is a sensor that monitors abnormal components, and the component monitoring data collected by the selected sensor is data related to the operation of the abnormal component. For example... Figure 4 As shown, Figure 4 A schematic diagram of the sensor setup on the target vehicle is shown. The candidate sensors are those located within the target vehicle, specifically including a powertrain sensor group, a rear sensor group, a body interior sensor group, a wheel sensor group, and an acoustic sensor group. The powertrain sensor group includes a vibration sensor, a first acoustic sensor, and a first temperature sensor. The first temperature sensor is used to detect the temperature of the motor and battery in the target vehicle. The rear sensor group includes a trunk vibration accelerometer, a tailgate sensor, and an exhaust system sensor, with the exhaust system sensor being a second acoustic sensor and a second temperature sensor located within the exhaust system. The body interior sensor group includes an array microphone, a directional microphone, and an ultrasonic sensor. The wheel sensor group includes left front wheel accelerometers, right front wheel accelerometers, left rear wheel accelerometers, and right rear wheel accelerometers.
[0054] Please see Figure 5 In some embodiments, step S203 may include, but is not limited to, steps S501 to S505: Step S501: Obtain candidate anomaly categories of preset abnormal audio features; Step S502: Calculate the similarity between the preset abnormal audio features and the selected running audio features to obtain the audio similarity. Step S503: Select the anomaly category from the candidate anomaly categories based on audio similarity; Step S504: Select a trigger signal from the preset trigger signals based on preset voice commands and audio similarity. Step S505: Select the selected sensor from the candidate sensors based on the selected anomaly category and the selected trigger signal.
[0055] In step S501 of some embodiments, each preset abnormal audio feature corresponds to a candidate abnormal category. For example... Figure 6As shown, the candidate anomaly categories include powertrain anomaly, chassis and suspension anomaly, vehicle interior anomaly, and acoustic perception anomaly. The preset abnormal audio features corresponding to the powertrain anomaly category can be at least one of belt squealing, bearing wear sound, and motor noise; the preset abnormal audio features corresponding to the chassis and suspension anomaly category are at least one of wheel bearing noise, suspension damping noise, and brake pad noise; and the preset abnormal audio features corresponding to the vehicle interior anomaly category are at least one of interior material loosening sound, wind noise, and resonance sound. Therefore, based on the preset abnormal audio features, the selected anomaly category of the target vehicle can be determined, and thus the abnormal components can be identified.
[0056] In step S502 of some embodiments, audio similarity represents the feature similarity between preset abnormal audio features and selected running audio features. Specifically, the preset abnormal audio features are converted into preset abnormal audio sequences, and the selected running audio features are converted into selected running audio sequences. The audio similarity between the preset abnormal audio sequences and the selected running audio sequences is calculated using the adaptive weighted DTW matching algorithm. It should be noted that in the process of calculating audio similarity using the adaptive weighted DTW matching algorithm, the importance weight of each audio feature is calculated first, then the matching similarity between audio features is calculated, and finally, the matching similarities are weighted and summed according to the importance weights to obtain the audio similarity.
[0057] In step S503 of some embodiments, candidate anomaly categories with audio similarity greater than a preset similarity are selected as anomaly categories. Specifically, as shown... Figure 6 As shown, if the preset similarity is 85%, and the audio similarity between the selected operating audio feature and the motor abnormality feature is greater than 85%, the powertrain abnormality category will be selected as the abnormality category. If the audio similarity between the selected operating audio feature and the interior loosening sound feature is greater than 85%, the chassis suspension abnormality category will be selected as the abnormality category.
[0058] In steps S504 and S505 of some embodiments, the preset trigger signal is the trigger signal for the main vehicle controller on the target vehicle to execute a response action. Specifically, the trigger level is first determined based on audio similarity and selected operating audio characteristics, and then a selected trigger signal is selected from the preset trigger signals based on the trigger level. Specifically, if the audio similarity is greater than the preset similarity, the trigger level is determined to be Level 1 trigger, and the selected trigger signal is to trigger the selected sensor. The selected sensor needs to be activated, allowing it to complete the data acquisition of the abnormal component according to the preset acquisition configuration parameters to obtain component monitoring data. If the audio similarity is less than the preset similarity, and the voice command indicates activation of all sensors, the selected trigger signal is determined to trigger all candidate sensors, and all candidate sensors are activated to complete data acquisition according to the acquisition configuration parameters. If the audio similarity is less than the preset similarity, and the voice command indicates no voice content is required, the trigger level is determined to be Level 3 trigger, and the original operating audio data continues to be acquired for abnormal noise monitoring.
[0059] It should be noted that abnormal components can also be identified based on raw operating audio data. Specifically, the direction of the abnormal sound source is detected from the raw operating audio data, the abnormal location information is determined based on the abnormal sound source direction, the abnormal component is identified based on the abnormal location information and the component's installation location information, and then the selected sensor is determined based on the abnormal component. Therefore, by identifying the abnormal component first and then selecting the sensor, the number of active sensors is reduced from "a group" to "a single one", further reducing system power consumption by 40% and improving fault location accuracy from "regional level" to "component level".
[0060] In steps S501 to S505 of this embodiment, the audio similarity between the preset abnormal audio and the selected running audio feature is first calculated. Candidate abnormal categories with audio similarity greater than the preset similarity are selected as abnormal categories. Then, the selected trigger signal is further determined. The selected trigger signal and the selected abnormal category are combined to complete the screening of candidate sensors and to activate the selected sensors in a targeted manner. The sensors do not need to be continuously started, thus saving energy consumption of the target vehicle.
[0061] In step S204 of some embodiments, such as Figure 6 As shown, when the selected anomaly category is determined to be a powertrain anomaly, the selected sensors are a vibration sensor, a first acoustic sensor, and a first temperature sensor. When the selected trigger signal activates the selected sensors, the vibration sensor, the first acoustic sensor, and the first temperature sensor are activated. The vibration sensor collects vibration data from the target vehicle, the first acoustic sensor collects sound data from the powertrain, and the first temperature sensor collects temperature data. The vibration data, sound data, and temperature data are then combined to form component monitoring data. It should be noted that the abnormal component of the target vehicle is determined based on the selected anomaly category.
[0062] Taking the fault prediction of "abnormal wear of the left front wheel bearing" as an example, when the left front wheel bearing of the target vehicle begins to show early wear and produces a weak, periodic abnormal noise, the acoustic scene perception chip continuously analyzes the in-vehicle ambient audio and extracts selected operating audio features from the vehicle's ambient audio. If the audio similarity between the selected operating audio feature and the preset abnormal audio feature "bearing wear sound feature" exceeds 85%, the trigger level is determined to be Level 1, the abnormal component is the left front suspension, the vibration sensor of the left front suspension is activated, and the vibration sensor completes vibration data acquisition according to the acquisition configuration parameters.
[0063] In step S205 of some embodiments, the amount of data continuously collected from the selected operating audio features and component monitoring data is large, and uploading it to the cloud subsystem in real time can easily cause network congestion. Therefore, this embodiment extracts key features for fault prediction from the selected operating audio features and component monitoring data as key prediction features, and only uploads these key prediction features to the cloud subsystem to save data transmission volume. It should be noted that the collected key prediction features are extracted by a first feature extraction model on the vehicle subsystem, and the first feature extraction model is obtained through knowledge transfer training of a second feature extraction model on the cloud subsystem. It is a lightweight model capable of extracting key prediction features equivalent to those of the second feature extraction model.
[0064] Please see Figure 7 In some embodiments, extracting predictive key features from selected operating audio features and component monitoring data may include, but is not limited to, steps S701 to S702: Step S701: Extract features from the component monitoring data to obtain component monitoring features; Step S702: Extract the key features for fault prediction from the component monitoring features and the selected operating audio features using a preset first feature extraction model; wherein, the first feature extraction model is obtained by adjusting the model parameters of the second feature extraction model on the server.
[0065] In step S701 of some embodiments, before feature extraction from the component monitoring data, the component monitoring data needs to be preprocessed, including wavelet denoising and standardization. Then, component monitoring features are extracted from the preprocessed component monitoring data. It should be noted that the component monitoring features characterize the key sensor features of the abnormal component. For example, if the abnormal component is the left front suspension, and the component monitoring data is vibration data, then the component monitoring features are the vibration characteristics of the left front suspension.
[0066] In step S702 of some embodiments, as disclosed above, the cloud subsystem receives predicted key features sent by multiple vehicle subsystems. Therefore, the second feature extraction model on the cloud subsystem is trained with a large number of predicted key features and can accurately extract key features for fault prediction. To this end, the model parameters of the second feature extraction model on the cloud subsystem are used to adjust the first feature extraction model, ensuring that the first feature extraction model has the same intermediate layers as the second feature extraction model, enabling it to perform equivalent feature extraction. It should be noted that the second feature extraction model has been trained to collect specific features from component monitoring features and selected operating audio features, and the first feature extraction model can directly collect predicted key features from the component monitoring features and selected operating audio features.
[0067] In some embodiments, the first feature extraction model is a student model trained using knowledge distillation, and the second feature extraction model is a teacher model set up in the cloud subsystem. The network architecture of the second feature extraction model is ResNet-18, while the network architecture of the first feature extraction model is MobileNetV2-Tiny. The second feature extraction model has 11.7M parameters, 1.8G of computation, and 128 dimensions, while the first feature extraction model has 350K parameters, 15M of computation, and 32 dimensions. It can be seen that by setting up a lightweight first feature extraction model in the vehicle subsystem, not only can fault prediction features be accurately extracted, but model space can also be saved and computational load reduced. Furthermore, the 32-dimensional prediction key features only account for 1-5% of the original operating audio data and component monitoring data, significantly reducing transmission bandwidth requirements.
[0068] In steps S701 to S702 of this embodiment, a lightweight first feature extraction model is set to complete the feature extraction of selected running audio features and component monitoring data, ensuring the accuracy of feature extraction, saving model space and computational load, and reducing the bandwidth occupied by feature transmission.
[0069] Please see Figure 8 In some embodiments, step S702 may include, but is not limited to, steps S801 to S805: Step S801: Extract time-domain features from the component monitoring features to obtain the first time-domain features; Step S802: Extract time-domain features from the selected audio features to obtain the second time-domain features; Step S803: Extract frequency domain features from the selected running audio features to obtain the first frequency domain features; Step S804: Concatenate the first time-domain feature, the second time-domain feature, and the first frequency-domain feature to obtain the concatenated feature; Step S805: Perform feature reasoning on the spliced features using the first feature extraction model to obtain the predicted key features.
[0070] In step S801 of some embodiments, the first time-domain feature includes the mean, standard deviation, peak value, peak-to-peak value, waveform factor, impulse factor, margin factor, and kurtosis from the component monitoring features. It should be noted that the first time-domain feature is an 8-dimensional feature.
[0071] In steps S802 and S803 of some embodiments, the second time-domain feature includes wavelet coefficient energy distribution features at eight scales in the selected running audio feature, and the first frequency-domain feature includes the spectral centroid, spectral bandwidth, spectral roll-off point, spectral flatness, and 1st-12th harmonic energy ratio in the selected running audio feature.
[0072] For example, define component monitoring data and selected running audio features as `preprocessed_data`, the first time-domain feature as `time_features=extract_time_domain(preprocessed_data)`, the first frequency-domain feature as `freq_features=extract_frequency_domain(preprocessed_data)`, and the second time-domain feature as `wavelet_features=extract_wavelet_features(preprocessed_data)`. Then, concatenate the first time-domain feature, the first frequency-domain feature, and the second time-domain feature to obtain the concatenated feature as `combined_features=np.concatenate([time_features, freq_features, wavelet_features]`. Then, use the first feature extraction model to infer the predicted key features from the concatenated features as `feature_vector=student_model.inference(combined_features)`.
[0073] In steps S801 to S804 of this embodiment, the first feature extraction model set in this embodiment clarifies the specific features of fault prediction. It extracts the first time-domain feature from the component monitoring feature, and extracts the second time-domain feature and the first frequency-domain feature from the selected operating audio feature. Then, it concatenates the first time-domain feature, the first frequency-domain feature and the second time-domain feature into a concatenated feature. The concatenated feature is then inferred into the prediction key feature through the first feature extraction model. The output of the key features of fault prediction is simple and accurate.
[0074] Please refer to Figure 9 , Figure 9The detailed process of edge feature distillation technology is illustrated. The collected component monitoring data, after wavelet denoising, standardization, and data segmentation, is input into a multi-dimensional feature extraction region. This extracts an 8-dimensional first time-domain feature, a 16-dimensional first frequency-domain feature, and an 8-dimensional second time-domain feature. These features are then concatenated into a 32-dimensional concatenated feature. This concatenated feature is input into the first feature extraction model for inference to output 32-dimensional predictive key features.
[0075] Please see Figure 10 In some embodiments, after step S702, the vehicle fault prediction method further includes, but is not limited to, steps S1001 to S1002: Step S1001: Based on the predicted key features, erase the component monitoring data and the original operating audio data, and perform erase detection on the component monitoring data and the original operating audio data to obtain erase detection information; Step S1002: If the erasure detection information characterization component monitoring data and original operating audio data have been erased, generate erasure proof information.
[0076] In step S1001 of some embodiments, the erasure operation involves overwriting the component monitoring data and original running audio data in the data buffer. This is done by overwriting the component monitoring data and original running audio data corresponding to the generated predicted key features based on the newly received component monitoring data and original running audio data. The writing pattern is: all 0 → all 1 → random → all 0 → all 1 → random → all 0. After the overwrite operation, an erasure detection is performed on the component monitoring data and original running audio data in the data buffer, specifically to determine whether the component monitoring data and original running audio data corresponding to the predicted key features still exist.
[0077] In step S1002 of some embodiments, the erasure proof information is generated by the erasure security model, which can characterize that there is no component monitoring data and original running audio data corresponding to the predicted key features in the data cache area, release the cache space of the data cache area, and facilitate the continued caching of subsequently collected component monitoring data and original running audio data.
[0078] In steps S1001 to S1002 of this embodiment, the monitoring data of the erased component and the original running audio data are generated after predicting key features, and erase proof information is generated, saving the space occupied by the data cache area.
[0079] After extracting the predicted key features, to enhance the security of the transmission of these features, they are encapsulated and encrypted into encrypted data packets. A secure channel is then established between the cloud subsystem and the vehicle subsystem using a pre-defined protocol, and the encrypted data is transmitted to the cloud subsystem through this secure channel.
[0080] Please see Figure 11 In some embodiments, the server obtains fault prediction information by predicting the fault of the target vehicle based on the predicted key features, which may include, but is not limited to, steps S1101 to S1103: Step S1101: The server obtains the vehicle model information of the target vehicle; In step S1102, the server selects a selected anomaly knowledge graph from a preset candidate anomaly knowledge graph based on the vehicle model information; wherein, the selected anomaly knowledge graph is constructed based on the anomaly features of the abnormal components in the target vehicle. In step S1103, the server performs fault prediction on the target vehicle by selecting anomaly knowledge graph and predicting key features, and obtains fault prediction information.
[0081] In step S1101 of some embodiments, the cloud subsystem on the server decrypts the encrypted data packet to obtain predicted key features, and obtains the vehicle model information of the target vehicle based on the predicted key features.
[0082] In step S1102 of some embodiments, the candidate anomaly knowledge graphs are different for different vehicle models, and each candidate anomaly knowledge graph records the abnormal features of abnormal components in the corresponding vehicle model, which can accurately identify the fault information of the target vehicle. Specifically, each candidate anomaly knowledge graph is set with a vehicle model identifier, and the selected anomaly knowledge graph is selected from the candidate anomaly knowledge graphs based on the vehicle model identifier and the vehicle model information of the target vehicle.
[0083] In step S1103 of some embodiments, the cloud subsystem on the server jointly selects anomaly knowledge graphs and predicts key features to complete fault prediction, predicting the fault information of the target vehicle in advance. It should be noted that the fault prediction information includes the fault level, the probability of fault occurrence for the selected anomaly category, and fault indication information. For example, if the selected anomaly category is left front wheel bearing wear, the output fault level is medium, the probability of fault occurrence is 0.93, and the fault indication information is "It is recommended to have it checked within 2000 kilometers." In this embodiment, fault prediction is completed by a fault prediction model, and the fault prediction model is obtained by training a large language model with a large number of vehicle training samples, enabling accurate fault prediction.
[0084] In some embodiments, as shown in Table 1, fault prediction can determine the cause of the fault from the table based on the acoustic features and CAN signals in the prediction key features, and send the cause of the fault to the after-sales service department of the target vehicle, so that the after-sales service department can provide fault maintenance suggestions.
[0085] Table 1
[0086] Furthermore, after completing the fault prediction, the vehicle subsystem provides feedback information based on the fault prediction information, selects reference operating audio features from the selected operating audio features based on the feedback information, and inputs the reference operating audio features into the preset abnormal audio feature feature library to update the feature library, so as to facilitate more accurate audio feature analysis.
[0087] In steps S1101 to S1103 of this embodiment, the fault prediction of the target vehicle is completed by combining the anomaly knowledge graph and the predicted key features through the cloud subsystem on the server. The fault prediction information can be accurately output, realize early fault prediction, and reduce the fault occurrence rate.
[0088] The following is a detailed description and explanation of the solution of this invention, using the specific example of abnormal wear of the left front wheel bearing: The acoustic scene perception chip in the vehicle subsystem continuously analyzes the ambient sound inside the target vehicle and extracts raw operating audio data in non-voice audio segments. When the vehicle speed reaches 80km / h, the chip detects a selected operating audio feature with an anomalous component (related to wheel speed) with a period frequency of approximately 15Hz.
[0089] The vehicle subsystem will match the detected selected operating audio features with the preset abnormal audio features in the feature library. If the matching result is 87% similar to the preset abnormal audio feature "early bearing wear", the trigger level is determined to be Level 1 trigger, and the selected sensor is determined to be the left front suspension vibration sensor group.
[0090] The vehicle subsystem wakes up the vibration sensor in the left front suspension area and activates it. The sampling rate of the vibration sensor is set to 5kHz and the acquisition time is 15 seconds. The amount of vibration data fed back by the vibration sensor is about 450KB.
[0091] The edge computing unit in the vehicle subsystem preprocesses the collected vibration data, including wavelet denoising and standardization. Then, it extracts features from the preprocessed vibration data and the original operating audio data: a 32-dimensional spliced feature consisting of 8-dimensional first time-domain features, 16-dimensional first frequency-domain features, and 8-dimensional second time-frequency features. Then, it uses the first feature extraction model on the vehicle subsystem to infer: compress the spliced features into 32-dimensional predicted key features, and then erase the original operating audio data and vibration data corresponding to the predicted key features, retaining only 128 bytes of predicted key features.
[0092] The predicted key features are uploaded to the cloud subsystem. The cloud subsystem, in conjunction with the anomaly knowledge graph of the target vehicle and the predicted key features, outputs fault prediction information. The fault prediction information is as follows: Fault type: early wear of the left front wheel bearing; Confidence level: 93%; Severity: minor; Recommended measures: It can be used normally at present. It is recommended to check within 5,000 kilometers; Estimated repair cost: 800-1,200 yuan.
[0093] The cloud-based subsystem pushes fault prediction information to users through the vehicle's mobile app and simultaneously synchronizes it to the 4S store's backend, making it convenient for users to schedule repairs.
[0094] This application's embodiments collect operational audio features from non-voice segments to complete fault prediction, avoiding voice content collection at the technical source and improving privacy and security. Simultaneously, intelligent triggering and sensor scheduling are implemented, with on-demand wake-up replacing full-time monitoring, reducing energy consumption by 90%. Key predictive features are extracted at the vehicle end using edge feature distillation technology, achieving 100:1 data compression and reducing network congestion during data transmission. Furthermore, after key predictive feature extraction, operational audio data is physically eliminated, enhancing privacy and security. Finally, tiered verification confirms the fault, implementing a multi-level triggering mechanism to reduce false alarm rates.
[0095] Please see Figure 12 This application also provides a vehicle fault prediction device that can implement the above-described method. The device includes: The data acquisition module 1201 is used to acquire the raw operating audio data of the target vehicle; The feature extraction module 1202 is used to extract selected running audio features from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; The sensor screening module 1203 is used to select a sensor from a preset candidate sensor based on preset abnormal audio characteristics and selected operating audio characteristics; wherein, the selected sensor is used to collect component monitoring data of abnormal component operation; The activation module 1204 is used to perform an activation operation on the selected sensor and receive component monitoring data collected by the selected sensor. The sending module 1205 is used to extract the predicted key features from the selected operating audio features and component monitoring data, send the predicted key features to the server, and receive the fault prediction information sent by the server; wherein, the server performs fault prediction on the target vehicle based on the predicted key features to obtain fault prediction information.
[0096] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0097] This application also provides a vehicle, which includes an on-board terminal and at least one sensor; wherein the on-board terminal is communicatively connected to a server and is used to execute the vehicle fault prediction method described above.
[0098] It is understood that the content of the above method embodiments is applicable to this vehicle embodiment. The specific functions implemented in this vehicle embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0099] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0100] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0101] The vehicle fault prediction method, device, electronic device, storage medium, and program product provided in this application predict faults in a target vehicle in advance. During the fault prediction process, the operating audio characteristics of the target vehicle are used. These operating audio characteristics do not include the speech content characteristics of the speaker, reducing the risk of user privacy leakage. Once the operating audio characteristics are determined, sensors are activated selectively based on preset abnormal audio characteristics and the operating audio characteristics. This eliminates the need to activate all sensors, accurately collecting component monitoring data of abnormal parts while saving vehicle energy consumption. Furthermore, before the server performs fault prediction, the on-board terminal extracts key prediction features from the component monitoring data and operating audio characteristics. Only these key prediction features are uploaded to the server to complete the fault prediction, saving data transmission volume, reducing network congestion probability, and ensuring the accuracy of the target vehicle fault prediction.
[0102] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0103] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0106] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0107] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0109] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A vehicle failure prediction method characterized by, The method, executed by an in-vehicle terminal that is communicatively connected to a server, includes the following steps: Obtain the raw operating audio data of the target vehicle; Selected running audio features are extracted from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; Selected sensors are selected from preset candidate sensors based on preset abnormal audio characteristics and the selected operating audio characteristics; wherein, the selected sensors are used to collect component monitoring data associated with abnormal components; Activate the selected sensor and receive the component monitoring data collected by the selected sensor; Predictive key features are extracted from the selected operating audio features and the component monitoring data, and the predicted key features are sent to the server. The server also receives fault prediction information from the server. The server performs fault prediction on the target vehicle based on the predicted key features to obtain the fault prediction information.
2. The method of claim 1, wherein, The step of extracting selected running audio features from the original running audio data includes: The original running audio data is preprocessed to obtain an original running audio frame sequence; wherein the original running audio frame sequence includes at least one original running audio frame; Feature extraction is performed on each of the original running audio frames to obtain the original running audio features; Speech activity detection is performed on the original running audio features to obtain speech activity detection information; Non-speech frames are filtered out from the original running audio frames based on the speech activity detection information; The original running audio features are filtered based on the non-speech frames to obtain the selected running audio features.
3. The method of claim 1, wherein, The step of selecting a sensor from a preset pool of candidate sensors based on preset abnormal audio features and the selected operating audio features includes: Obtain candidate anomaly categories for the preset abnormal audio features; The similarity between the preset abnormal audio features and the selected running audio features is calculated to obtain the audio similarity. Selected anomaly categories are filtered from the candidate anomaly categories based on the audio similarity. Selected trigger signals are selected from preset trigger signals based on preset voice commands and audio similarity. The selected sensor is selected from the candidate sensors based on the selected anomaly category and the selected trigger signal.
4. The method of claim 1, wherein, The extraction of predictive key features from the selected operating audio features and the component monitoring data includes: Feature extraction is performed on the component monitoring data to obtain component monitoring features; The first feature extraction model is used to extract the key features for fault prediction from the component monitoring features and the selected operating audio features; wherein, the first feature extraction model is adjusted according to the model parameters of the second feature extraction model on the server.
5. The method of claim 4, wherein, The step of extracting key predictive features for fault prediction from the component monitoring features and the selected operating audio features using a preset first feature extraction model includes: The monitoring features of the component are subjected to time-domain feature extraction to obtain the first time-domain feature; Temporal feature extraction is performed on the selected audio features to obtain the second temporal feature; Frequency domain features are extracted from the selected audio features to obtain the first frequency domain features; The first time-domain feature, the second time-domain feature, and the first frequency-domain feature are concatenated to obtain the concatenated feature; The first feature extraction model is used to perform feature inference on the spliced features to obtain the predicted key features.
6. The method of claim 4, wherein, After extracting the key features for fault prediction from the component monitoring features and the selected operating audio features using a preset first feature extraction model, the method further includes: Based on the predicted key features, the component monitoring data and the original operating audio data are erased, and the erase detection is performed on the component monitoring data and the original operating audio data to obtain erase detection information; If the erasure detection information indicates that the component monitoring data and the original operating audio data have been completely erased, erasure proof information is generated.
7. The method according to any one of claims 1 to 6, characterized in that, The server performs fault prediction on the target vehicle based on the predicted key features to obtain the fault prediction information, including: The server obtains the vehicle model information of the target vehicle; The server selects a chosen anomaly knowledge graph from a preset candidate anomaly knowledge graph based on the vehicle model information; wherein, the chosen anomaly knowledge graph is constructed based on the anomaly features of the anomaly components in the target vehicle; The server performs fault prediction on the target vehicle using the selected anomaly knowledge graph and the predicted key features, thereby obtaining the fault prediction information.
8. A vehicle failure prediction device characterized by comprising: The device is installed on a vehicle-mounted terminal, which is communicatively connected to a server, and the device includes: The data acquisition module is used to acquire the raw operating audio data of the target vehicle; The feature extraction module is used to extract selected running audio features from the original running audio data; wherein, the selected running audio features do not include the speech content features of the speaker; The sensor screening module is used to select a sensor from a preset candidate sensor based on preset abnormal audio characteristics and the selected operating audio characteristics; wherein, the selected sensor is used to collect component monitoring data of abnormal component operation; An activation module is used to perform an activation operation on the selected sensor and receive the component monitoring data collected by the selected sensor. The sending module is used to extract predictive key features from the selected operating audio features and the component monitoring data, send the predictive key features to the server, and receive fault prediction information sent by the server; wherein, the server obtains the fault prediction information by performing fault prediction on the target vehicle based on the predictive key features.
9. A vehicle characterized by comprising: The vehicle includes an on-board terminal and at least one sensor; wherein the on-board terminal is communicatively connected to a server, and the on-board terminal is used to execute the vehicle fault prediction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.