Vehicle health determination method and device and electronic equipment

By collecting and fusing vehicle voiceprint and operational status data, and utilizing feature extraction and fusion technologies, the problem of inaccurate assessment in determining vehicle health has been solved, enabling precise quantification of vehicle health levels and improving the scientific nature of used car valuation and the reliability of loan risk management.

CN121789714APending Publication Date: 2026-04-03AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for determining vehicle health cannot fully consider the vehicle's internal operating condition and wear, leading to inaccurate assessments and affecting the scientific granting of vehicle value and loan amounts.

Method used

By acquiring the vehicle's initial voiceprint data and operating status data, using a microphone array to collect voiceprint data, and combining it with OBD data for feature extraction and fusion, and employing technologies such as adaptive filtering, generative adversarial networks, Mel-frequency cepstral coefficients, bidirectional long short-term memory networks, and cross-attention mechanisms, a precise quantitative assessment of the vehicle's health level can be achieved.

Benefits of technology

It enables precise quantitative assessment of vehicle health levels, improves the objectivity and accuracy of the assessment, provides a scientific basis for vehicle value assessment, and reduces the risk of used car loans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789714A_ABST
    Figure CN121789714A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle health determination method and device and electronic equipment. Relates to the field of financial science and technology, and the method comprises the steps: obtaining the initial voiceprint data and initial operation state data of a vehicle, and enabling the initial voiceprint data to be collected through a microphone array in the operation process of the vehicle; performing feature extraction on the initial voiceprint data to obtain target voiceprint features; performing feature extraction on the initial operation state data to obtain target operation features; performing feature fusion on the target voiceprint feature and the target operation feature to obtain a fused feature; and determining a health level of the vehicle based on the fused features. According to the method and the device, the technical problem of inaccurate vehicle health determination caused by difficulty in comprehensively considering the internal running state and wear of the vehicle in the vehicle health determination in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and more specifically, to a method, apparatus, and electronic device for determining vehicle health. Background Technology

[0002] In the field of vehicle health assessment, particularly in risk control for used car loans by financial institutions, the vehicle's internal operating condition and wear are key factors in evaluating its health level and value. However, assessment methods in related technologies often rely on the appraiser's subjective judgment and limited external inspections, such as paintwork, interior wear, and basic engine compartment checks. These methods struggle to comprehensively and accurately reflect the internal health of core vehicle components, such as the engine, transmission, and chassis. Due to the lack of quantitative indicators, the consistency and reliability of assessment results are influenced by the appraiser's personal experience, emotions, and the inspection environment, leading to inaccurate vehicle health assessments and directly impacting the reasonable valuation of the vehicle and the scientific granting of loan amounts. Furthermore, the assessment methods in related technologies are insufficient in detecting minor wear, potential malfunctions, and performance degradation of internal components, making it difficult to comprehensively consider the vehicle's internal operating condition and wear, further contributing to inaccurate vehicle health assessments and affecting vehicle value evaluation.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This invention provides a method, apparatus, and electronic device for determining vehicle health, which at least solves the problem in related technologies that it is difficult to comprehensively consider the internal operating state and wear of a vehicle when determining its health, leading to inaccurate vehicle health determination.

[0005] According to one aspect of the present invention, a vehicle health determination method is provided, comprising: acquiring initial voiceprint data and initial operating state data of a vehicle, wherein the initial voiceprint data is acquired by a microphone array during vehicle operation; performing feature extraction on the initial voiceprint data to obtain target voiceprint features; performing feature extraction on the initial operating state data to obtain target operating features; performing feature fusion on the target voiceprint features and the target operating features to obtain fused features; and determining the vehicle's health level based on the fused features.

[0006] According to another aspect of the present invention, a vehicle health determination device is provided, comprising: a first data acquisition module for acquiring initial voiceprint data and initial operating state data of a vehicle, wherein the initial voiceprint data is acquired by a microphone array during vehicle operation; a voiceprint feature extraction module for extracting features from the initial voiceprint data to obtain target voiceprint features; an operating feature extraction module for extracting features from the initial operating state data to obtain target operating features; a feature fusion module for fusing the target voiceprint features and the target operating features to obtain fused features; and a health level determination module for determining the health level of the vehicle based on the fused features.

[0007] According to another aspect of the present invention, a non-volatile storage medium is also provided, the non-volatile storage medium storing a plurality of instructions adapted for loading by a processor and executing any one of the vehicle health determination methods described herein.

[0008] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the vehicle health determination methods.

[0009] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of any of the vehicle health determination methods described herein.

[0010] In this embodiment of the invention, initial voiceprint data and initial operating status data of a vehicle are acquired, wherein the initial voiceprint data is collected by a microphone array during vehicle operation; feature extraction is performed on the initial voiceprint data to obtain target voiceprint features; feature extraction is performed on the initial operating status data to obtain target operating features; feature fusion is performed on the target voiceprint features and the target operating features to obtain fused features; based on the fused features, the health level of the vehicle is determined. This achieves the goal of accurately quantifying and evaluating the vehicle health level by fusing voiceprint data and operating status data during vehicle operation and utilizing feature extraction and fusion technology. This improves the objectivity and accuracy of vehicle health level assessment, and solves the problem in related technologies where it is difficult to comprehensively consider the vehicle's internal operating status and wear when determining vehicle health, leading to inaccurate vehicle health determination. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0012] Figure 1 This is a flowchart of a vehicle health determination method according to an embodiment of the present invention;

[0013] Figure 2 This is a flowchart of a vehicle valuation method according to an embodiment of the present invention;

[0014] Figure 3 This is a flowchart of an optional vehicle valuation method according to an embodiment of the present invention;

[0015] Figure 4 This is a schematic diagram of a vehicle health determination device according to an embodiment of the present invention;

[0016] Figure 5 This is a schematic diagram of a vehicle valuation device according to an embodiment of the present invention. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0019] First, to facilitate understanding of the embodiments of the present invention, some terms or nouns involved in the present invention will be explained below:

[0020] The Least Mean Squares (LMS) adaptive filtering algorithm is an algorithm that automatically adjusts the filter coefficients based on the minimum mean square error criterion and uses stochastic gradient descent method with instantaneous gradient estimation.

[0021] Speech Enhancement Generative Adversarial Network (SEGAN) is an end-to-end speech enhancement algorithm based on generative adversarial networks.

[0022] Mel-Frequency Cepstral Coefficients (MFCC) are one of the most fundamental feature extraction techniques in the fields of speech signal processing and audio recognition.

[0023] On-Board Diagnostics Parameters (OBD) are data variables used in vehicle on-board diagnostic systems to monitor, record, and provide feedback on the operating status of various vehicle subsystems.

[0024] Mel-frequency cepstral coefficients (MFCCs) are a feature extraction method widely used in speech recognition and audio processing. Their design is inspired by the human ear's perception of different frequency sounds, particularly the perception of important frequency components in speech.

[0025] Bidirectional Long Short-Term Memory (BiLSTM) networks are deep learning models that combine bidirectional recurrent neural networks and long short-term memory networks. They are mainly used to process sequential data (such as text, speech, time series, etc.) and can simultaneously capture the forward and backward dependencies of a sequence, significantly improving the ability to understand contextual information.

[0026] Feature fusion is the effective integration of feature information from different sources, modalities, or levels to improve the model's representation ability and performance.

[0027] Cross-Model Attention is an attention mechanism that achieves information exchange by dynamically calculating the correlation weights between different sequences or feature sets.

[0028] A Transformer is a deep learning model architecture based on a self-attention mechanism.

[0029] According to an embodiment of the present invention, a method for determining vehicle health is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] Figure 1 This is a flowchart of a vehicle health determination method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0031] Step S102: Obtain the initial voiceprint data and initial operating status data of the vehicle. The initial voiceprint data is collected by a microphone array during vehicle operation.

[0032] Optionally, the vehicle can be a used car, and the initial voiceprint data can be obtained by capturing its voiceprint signals using a high-sensitivity microphone array while the vehicle is in operation. Voiceprint data can provide unique information about the vehicle's internal mechanical state, such as the operating noise of the engine and transmission system. The use of a microphone array ensures comprehensive coverage of the captured sound information, including all frequencies within the range of human hearing, thus providing rich raw data for subsequent feature extraction. Initial operating status data can be read in real time through the vehicle's OBD interface, involving key parameters such as engine speed, oil temperature, intake air volume, and torque. This data reflects the state of the vehicle's electronic systems during actual operation, and the combination of both provides basic information for a comprehensive assessment of the vehicle's health.

[0033] Step S104: Extract features from the initial voiceprint data to obtain the target voiceprint features.

[0034] Optionally, feature extraction can transform the raw voiceprint data into a feature representation with diagnostic value. This can be achieved, but is not limited to, through noise reduction, enhancement, framing, Hamming window weighting, and finally, using the MFCC extraction method to convert the spectral features of the voiceprint data into a series of digital features, including static, first-order difference, and second-order difference coefficients. These features can quantitatively reflect the health status of the vehicle's internal mechanical condition, providing a direct basis for subsequent health level classification.

[0035] In one optional embodiment, feature extraction is performed on the initial voiceprint data to obtain target voiceprint features, including: filtering the initial voiceprint data to obtain denoised voiceprint data; performing nonlinear enhancement processing on the denoised voiceprint data to obtain enhanced voiceprint data; performing frame processing on the enhanced voiceprint data according to a preset window length to obtain framed voiceprint data; performing weighted processing on the framed voiceprint data using the Hamming window method to obtain preprocessed voiceprint data; and extracting features from the preprocessed voiceprint data to obtain target voiceprint features.

[0036] Optionally, but not limited to, the LMS adaptive filtering algorithm can be used to preprocess the acquired raw voiceprint data. By dynamically adjusting the filter parameters, environmental noise can be effectively removed, improving the purity of the voiceprint signal and creating favorable conditions for subsequent feature extraction. The SEGAN algorithm can be used to nonlinearly enhance the denoised voiceprint data, further improving the signal-to-noise ratio and ensuring that even weak mechanical noise changes can be clearly captured, enhancing the recognizability of voiceprint features. Based on a preset time window length, the enhanced voiceprint data is divided into a series of continuous frames. This operation is beneficial for capturing short-term acoustic characteristics, facilitating the next step of spectral analysis. The framed voiceprint data is weighted using a Hamming window function, mainly to reduce sidelobe effects during spectral analysis, improve spectral resolution, and ensure the accuracy of the extracted voiceprint features. Further feature extraction is performed on the preprocessed voiceprint data to obtain the corresponding target voiceprint features.

[0037] For example, firstly, the LMS adaptive filtering algorithm is used to dynamically adjust the filter parameters to eliminate environmental noise; secondly, the SEGAN algorithm is used to nonlinearly enhance the denoised signal, increasing the signal-to-noise ratio to over 20 dB; finally, the signal is framed into frames with a window length of T and an overlap of t, and a Hamming window is used to reduce spectral leakage.

[0038] Through the above steps, the raw voiceprint data is transformed into target voiceprint features that reflect the mechanical health status of the vehicle's interior, providing high-quality data support for subsequent feature fusion and health level determination. The entire feature extraction process reduces the impact of external interference while preserving key information in the voiceprint data, ensuring the authenticity and accuracy of the voiceprint features, thereby enabling a more accurate assessment of the vehicle's health status.

[0039] In one optional embodiment, feature extraction is performed on the initial voiceprint data to obtain target voiceprint features, including: performing frame-segmentation processing based on the initial voiceprint data to obtain multiple frames of voiceprint data; performing Fourier transform on each of the multiple frames of voiceprint data to obtain multiple voiceprint spectra, wherein the multiple voiceprint spectra correspond one-to-one with the multiple frames of voiceprint data; using a Mel-scale filter to map the multiple voiceprint spectra to a Mel scale to obtain the mapping results corresponding to each of the multiple voiceprint spectra; performing a discrete cosine transform on the mapping results corresponding to each of the multiple voiceprint spectra to obtain a first voiceprint feature; performing first-order difference processing on the first voiceprint feature to obtain a second voiceprint feature; and performing second-order difference processing on the first voiceprint feature to obtain a third voiceprint feature, wherein the target voiceprint features include the first voiceprint feature, the second voiceprint feature, and the third voiceprint feature.

[0040] Optionally, the voiceprint data obtained during vehicle operation can be segmented into a series of frames within a fixed time window, with each frame representing the acoustic characteristics over a certain duration. Framing facilitates subsequent spectrum analysis and feature extraction. A Fast Fourier Transform (FFT) is performed on each frame of voiceprint data to convert the time-domain signal into a frequency-domain signal, obtaining the corresponding sound spectrum. Spectrum analysis reveals the frequency composition of the sound signal, which is crucial for identifying the source and type of mechanical noise. The obtained voiceprint spectrum is processed using a Mel-scale filter bank, mapping the spectrum to a Mel frequency scale, which better matches the human ear's perception of different frequency components. Mel-scale mapping is a key step in MFCC feature extraction, helping to extract voiceprint features closely related to human auditory perception. A Discrete Cosine Transform (DCT) is performed on the mapped result to generate a set of coefficients, i.e., the first voiceprint feature (static MFCC feature). DCT can convert the power distribution of the voiceprint spectrum into a set of easily analyzable features; static MFCC reflects the steady-state properties of the voiceprint at a certain point in time. Further processing of the first voiceprint feature using first-order difference yields the second voiceprint feature, which reflects the instantaneous rate of change of the voiceprint. The first voiceprint feature is then processed using second-order difference to obtain the third voiceprint feature, which reflects the change in voiceprint acceleration. Extraction of dynamic features aims to capture the behavior of the voiceprint over time, which is very useful for detecting abnormal vibrations or mechanical faults.

[0041] Optionally, the preprocessed audio signal (i.e., voiceprint data) is first divided into frames of duration t. Windowing is applied to each frame to reduce edge effects, and the spectrum is obtained using Fast Fourier Transform (FFT). Subsequently, the spectrum is mapped to the Mel scale using a Mel-scale filter bank, and after calculating the logarithmic energy, 13-dimensional static MFCC features are extracted using Discrete Cosine Transform (DCT). To enhance dynamic features, first-order and second-order differences are further calculated for these 13-dimensional static MFCC features to characterize the instantaneous rate of change and acceleration of the voiceprint, respectively. Finally, a 39×L-dimensional vector (13 static + 13 first-order + 13 second-order) is generated for each frame, where L is the number of MFCC frames within a given time period. Static features reflect the steady-state operating mode of the equipment, while dynamic features capture microscopic vibration anomalies (such as a sudden increase in high-frequency harmonics caused by bearing cracks).

[0042] The target voiceprint features obtained through the above methods include three types of features: static, first-order difference, and second-order difference, namely, first voiceprint features, second voiceprint features, and third voiceprint features. These features together constitute a voiceprint representation that comprehensively reflects the mechanical health status of the vehicle's interior, providing crucial information for subsequent feature fusion and health level determination. This method fully utilizes the frequency and time domain characteristics of voiceprint signals, improving the richness and diagnostic value of the features, and plays an important role in refined vehicle health assessment (such as used cars).

[0043] Optionally, the initial voiceprint data can be preprocessed to obtain preprocessed voiceprint data; the preprocessed voiceprint data can be segmented into frames to obtain multi-frame voiceprint data. The process of preprocessing the initial voiceprint features to obtain preprocessed voiceprint features is the same as in the aforementioned embodiments and will not be repeated here.

[0044] Step S106: Extract features from the initial running state data to obtain the target running features.

[0045] Optionally, feature extraction of operational status data can be achieved using a BiLSTM network. This network can capture long-term dependencies in time-series data and transform the data into a time-series feature vector reflecting the vehicle's operational health status, i.e., the target operational feature. This feature contains dynamic information about the operational status of the vehicle's electronic systems, which helps to comprehensively understand the vehicle's operational health level.

[0046] In one optional embodiment, feature extraction is performed on the initial operating state data to obtain target operating features, including: aligning the initial operating state data with the voiceprint data by sampling frequency to obtain aligned operating state data; normalizing the aligned operating state data to obtain preprocessed operating feature data; and extracting features from the preprocessed operating feature data to obtain target operating features.

[0047] Optionally, firstly, the sampling frequency of the OBD data (i.e., the initial operating status data of the vehicle) is aligned with the sampling frequency of the voiceprint data. Since voiceprint data and OBD data may originate from different sensors, their sampling frequencies may differ. Frequency alignment ensures consistency of the two types of data on the time axis, facilitating subsequent feature fusion operations and enhancing information synchronization and contrast. The aligned operating status data is then normalized to eliminate dimensional differences between different physical quantities (such as temperature, speed, and torque), ensuring all parameters are compared and analyzed on the same scale. This prevents any single data point from having a dominant influence on feature extraction due to its numerical value, ensuring the fairness of data processing and the stability of model training. Further feature extraction is performed on the preprocessed operating status data to obtain the corresponding target operating features. For example, OBD data (i.e., the initial operating status data of a used car) reads parameters such as engine speed, oil temperature, intake air volume, and torque in real time through the OBD-II interface. The sampling frequency is synchronized with the voiceprint signal and normalized to eliminate dimensional differences, ensuring the accuracy of subsequent feature extraction. Through alignment, normalization, and deep learning extraction, the resulting features can more accurately reflect the vehicle's operating status, providing strong support for feature fusion and health level assessment, and ensuring the accuracy and scientific nature of the entire vehicle health determination method.

[0048] In one optional embodiment, feature extraction is performed on the initial running state data to obtain target running features, including: using a bidirectional long short-term memory network to extract the temporal dependencies in the initial running state data to obtain bidirectional temporal features; and performing dimensionality reduction processing on the bidirectional temporal features through a first fully connected layer to match the dimension of the target voiceprint features to obtain the target running features.

[0049] Optionally, OBD data (i.e., operational status data) can be input into a bidirectional long short-term memory (BiLSTM) network, which can simultaneously analyze the forward and backward dependencies of the data. BiLSTM can capture long-term dependencies in vehicle operational status data, such as the patterns of changes in parameters like engine speed, oil temperature, and intake air volume over time. Through BiLSTM, the dynamic changes of parameters over time and potential abnormal behaviors, such as parameter abrupt changes during rapid acceleration or braking, can be learned. Dimensionality reduction of the bidirectional temporal features extracted by BiLSTM through a first fully connected layer can match the dimensionality of the obtained features with that of the voiceprint features, facilitating subsequent feature fusion. Dimensionality reduction helps reduce computational cost and the dimensionality of the feature space while retaining the most relevant information, which positively impacts model efficiency and performance. The target operational features obtained through the above process contain temporal information from the OBD data and content consistent with the dimensionality of the voiceprint features after dimensionality reduction, facilitating cross-modal information fusion. The target operational features are an effective representation of OBD data in the temporal dimension, reflecting the operational status of the vehicle's internal systems over time, and are an important component in assessing vehicle health levels and evaluating the value of used cars.

[0050] For example, a bidirectional long short-term memory (BiLSTM) network is used to extract long-term dependencies from OBD data. Specifically, OBD data (such as engine speed and temperature) stores the vehicle's operating state information. After normalization, this data forms an input vector, which is then fed into a two-layer BiLSTM structure. Each layer contains m memory units, where the forward LSTM captures the temporal dependencies from the current moment to future states, and the backward LSTM captures the temporal dependencies from the current moment to historical states. The bidirectional outputs are concatenated through fully connected layers and mapped to an n-dimensional temporal feature vector. A gating mechanism is used to mitigate the gradient vanishing problem, and bidirectional modeling is used to analyze both historical and future data simultaneously, improving the ability to extract features from sudden changes in vehicle state (such as rapid acceleration).

[0051] This embodiment uses a BiLSTM network to mine time-series dependencies from OBD data and performs dimensionality reduction to ensure these features are dimensionally consistent with voiceprint features, thereby obtaining target operational features that can be fused across modalities. This feature extraction method can more comprehensively reflect the vehicle's operational status and has a significant effect on improving the accuracy and efficiency of vehicle (such as used cars) health level assessment and valuation.

[0052] Optionally, the initial running state data can be preprocessed to obtain preprocessed running state data; a bidirectional long short-term memory network can then be used to extract the temporal dependencies in the preprocessed running state data to obtain bidirectional temporal features. The process of preprocessing the initial running state data to obtain the preprocessed running state data is the same as in the aforementioned embodiments and will not be repeated here.

[0053] Step S108: Perform feature fusion on the target voiceprint features and target operation features to obtain fused features.

[0054] Optionally, feature fusion combines voiceprint features and operational features within a unified feature space to enhance the model's comprehensive ability to identify vehicle health status. The two types of features complement each other and share information, resulting in fused features that contain comprehensive information needed to determine vehicle health, thus improving the model's representational ability and classification accuracy.

[0055] In one optional embodiment, feature fusion is performed on the target voiceprint features and the target running features to obtain fused features, including: mapping the target voiceprint features and the target running features to the same feature space through a second fully connected layer to obtain mapped voiceprint features and mapped running features; using the mapped running features as a query vector and the mapped voiceprint features as a key-value vector, a cross-attention mechanism is used to calculate attention weights; the mapped voiceprint features are weighted according to the attention weights to obtain weighted voiceprint features; and the weighted voiceprint features and the mapped running features are concatenated to obtain fused features.

[0056] Optionally, firstly, a second fully connected layer is used to map the target voiceprint features and target operational features to a common feature space. This operation aims to make features from different modalities comparable, facilitating subsequent fusion processing. The mapping process can involve dimensionality reduction or expansion of the features to ensure that the voiceprint features and operational features match in dimensionality. The mapped operational features are used as the query vector, and the mapped voiceprint features as the key vector, preparing for the use of a cross-attention mechanism. In the attention mechanism, the query vector represents the information that needs attention during fusion, while the key vector helps determine which voiceprint information is most relevant to the operational features. The cross-attention mechanism dynamically calculates the correlation weights between the mapped voiceprint features and operational features, i.e., attention weights. This mechanism automatically identifies the correlations between features from different modalities through the interaction between the query vector and the key vector, thereby assigning different weights to the voiceprint features, enabling them to be processed more intelligently and specifically during fusion. Based on the calculated attention weights, the mapped voiceprint features are weighted to obtain weighted voiceprint features. The weighting process highlights voiceprint features that are highly correlated with operational characteristics or more important to the overall information, further enhancing the representational power of the fused features. Finally, the weighted voiceprint features are concatenated with the mapped operational features to generate the fused features. This fusion method can fully utilize the complementary information between voiceprint and operational data to form a more comprehensive and accurate representation of vehicle health status.

[0057] For example, cross-modal fusion can be achieved through dynamic feature alignment and attention mechanisms. First, multi-scale MFCC features (i.e., target voiceprint features) and OBD temporal features (i.e., target running features) are mapped to a unified dimension through a fully connected layer. .in ,in It is the result of mapping the OBD temporal features after mapping. It is the input of OBD features. The OBD features are biased, and the MFCC features are mapped using the same method. Finally, the two are concatenated and input into the multi-head attention mechanism layer. A cross-model attention mechanism is employed to fuse the concatenated features to achieve cross-modal complementarity, capture the correlation between different feature subspaces, and enhance the model's expressive power. The query vector Q is obtained by linearly projecting the OBD features (such as rotation speed and temperature), representing the current context to be focused on. The key vector K is obtained by linearly projecting the MFCC audio features, and the value vector V represents the retrieveable information of the MFCC audio features, obtained by linearly projecting the MFCC audio features, representing the actual information content of the audio features. Attention weights are also included. Depend on It is calculated, among which It is the projection dimension of the query vector and the key vector, and the output is a fused feature vector obtained by weighting. This enables cross-modal complementarity while addressing the issue of differences in feature dimensions.

[0058] This embodiment employs a cross-attention mechanism and a fully connected layer to integrate voiceprint features and operational characteristics, achieving intelligent fusion of multimodal information. This process generates comprehensive features that integrate vehicle operational status and voiceprint information, playing a crucial role in improving the accuracy of used car health level assessments and the reliability of subsequent used car value assessments. This approach not only fully utilizes multimodal data but also ensures the intelligence of the fusion process through the attention mechanism.

[0059] Step S110: Determine the health level of the vehicle based on the fused features.

[0060] Optionally, the fused features can be input into a classifier, which can output a quantified health level based on the vehicle's health status as represented. This level classification can guide vehicle maintenance and repair decisions, and is also an important basis for assessing vehicle value and controlling risks in used car loans. Through this step, an objective and quantitative assessment of the vehicle's health status can be achieved.

[0061] In one optional embodiment, determining the health level of a vehicle based on fused features includes: inputting the fused features into a multimodal health classifier based on a Transformer architecture to obtain the vehicle's health level; wherein the multimodal health classifier is used to capture the global dependencies of the fused features through a multi-head self-attention mechanism and learn the interaction between different features; the multimodal health classifier is obtained by machine learning based on the fused features and health levels corresponding to multiple vehicles.

[0062] Optionally, the fused features generated in the above steps can be used as input to a multimodal health classifier based on the Transformer architecture. The fused features contain a comprehensive representation of voiceprint information and operational status data, capable of fully reflecting the vehicle's health status. The multi-head self-attention mechanism in the Transformer architecture allows the model to learn the interdependencies of features simultaneously from multiple different representation subspaces. This means that it can not only capture global dependencies in the fused features but also handle complex relationships between different features, such as subtle interactions between OBD data and voiceprint features, significantly improving the classifier's sensitivity to information and its ability to handle complexity. Through the multimodal health classifier, the method in this embodiment can learn the interactions between different information in the fused features. For example, changes in engine speed may affect voiceprint features at specific frequencies; this interaction can be captured and learned by the classifier to more accurately assess the vehicle's health status. The multimodal health classifier is trained on a large amount of vehicle data, including fused features of various vehicles and known health levels. Through machine learning, the classifier can identify the correlation between different feature combinations and health levels, thereby predicting the health level of new vehicles.

[0063] For example, the fusion feature input is based on a Transformer-based multimodal fusion classifier, capturing the global dependencies of the fusion features. Multi-head attention is used to learn the dependencies of multiple subspaces in parallel. Simultaneously, considering the gradual wear and tear caused by long-term high-load operation of a car engine, its audio features and historical OBD data exhibit long-range dependencies. A self-attention mechanism is used to automatically associate the relationship between the current MFCC frame and historical OBD data. The classification output includes five categories: high risk, medium-high risk, medium risk, medium-low risk, and low risk, representing a health level classification.

[0064] This embodiment employs a multimodal health classifier based on the Transformer architecture. It uses a multi-head self-attention mechanism to learn global dependencies and interactions between features in the fused features to determine the vehicle's health level. This process fully leverages the powerful learning capabilities of deep learning models, enabling the extraction of patterns from historical data and accurate assessment of vehicle health status, thus providing a scientific basis for used car valuation.

[0065] Through the above steps S102 to S110, by fusing voiceprint data and operating status data during vehicle operation and utilizing feature extraction and fusion technology, the goal of accurately quantifying and assessing the vehicle health level can be achieved. This improves the objectivity and accuracy of vehicle health level assessment and solves the problem in related technologies where it is difficult to comprehensively consider the vehicle's internal operating status and wear when determining vehicle health, leading to inaccurate vehicle health determination.

[0066] According to an embodiment of the present invention, a method for vehicle valuation is provided. Figure 2 This is a flowchart of a vehicle valuation method according to an embodiment of the present invention, such as... Figure 2 As shown, the method includes the following steps:

[0067] Step S202: Obtain the vehicle's health level, wherein the health level is obtained based on any of the above-mentioned vehicle health determination methods;

[0068] Step S204: Obtain the vehicle's basic information, mileage, and market data;

[0069] Step S206: Based on the health level, basic vehicle information, vehicle mileage, and market data, a vehicle valuation model is used to obtain the vehicle's valuation result.

[0070] Optionally, firstly, the vehicle's health rating is obtained by implementing any of the aforementioned vehicle health determination methods. This rating is a quantitative evaluation of the vehicle's intrinsic quality and operational condition, directly affecting its residual value and potential risks. Basic vehicle information (such as brand, model, and year), mileage, and relevant market data (such as market price and supply and demand for the same model) are collected. A comprehensive vehicle valuation model is then used, taking the health rating, basic vehicle information, mileage, and market data as inputs, to calculate the vehicle's valuation result. This model fully considers the influence of multi-source information, incorporating the vehicle health rating into the evaluation framework, thereby providing a more comprehensive vehicle value estimate.

[0071] In steps S202 to S206 above, a multi-level, multi-dimensional evaluation system is formed by combining internal health levels with external information. The introduction of health levels means that the evaluation results are no longer simple inferences based on surface information and statistical regularities, but incorporate consideration of the vehicle's actual operating status, significantly enhancing the accuracy and credibility of the evaluation. Furthermore, the integration with market data ensures that the evaluation results are consistent with current market demand, reflecting the vehicle's value in actual transactions. This method provides financial institutions with a more scientific and reasonable valuation basis when handling used car loan business, helping to reduce loan risk and optimize the loan process. In other words, the vehicle valuation method in this embodiment, by integrating vehicle health levels and comprehensive vehicle information, achieves a more accurate and comprehensive assessment of vehicle value, providing strong support for risk management and service improvement for banks and other financial institutions in the used car loan business.

[0072] In the process of reviewing consumer loans for vehicles (such as used cars) by financial institutions, the valuation of used cars is often closely related to the loan credit limit. However, used car valuations in related technologies mainly rely on the appraiser's experience and visual observation (such as checking the paint, interior wear, and engine compartment). This is greatly influenced by factors such as the appraiser's personal experience, emotions, and environment, resulting in inconsistent and unobjective valuations. Currently, financial institutions still primarily rely on manual assessments for used car valuations. This involves account managers conducting on-site investigations and collaborating with appraisers' experience-based judgments and visual inspections (such as checking the paint, interior wear, and engine compartment). This approach makes it difficult to comprehensively assess the internal operating condition and potential wear of core vehicle components (such as the engine, transmission, and chassis), especially hidden faults. The lack of accuracy and quantifiable indicators introduces uncertainty and risk into the granting of used car consumer loans.

[0073] To address the aforementioned problems, and in conjunction with the above embodiments and optional embodiments, the present invention proposes an optional implementation method. Figure 3 This is a flowchart of an optional vehicle valuation method according to an embodiment of the present invention, such as... Figure 3 As shown, the method includes:

[0074] S1, time series information acquisition and preprocessing, specifically includes:

[0075] Voiceprint information (i.e., initial voiceprint data of used cars) is collected using a high-sensitivity microphone array to cover the range of human hearing. The preprocessing process includes: first, using the LMS adaptive filtering algorithm to dynamically adjust filter parameters and eliminate environmental noise; second, using the SEGAN algorithm to nonlinearly enhance the denoised signal, increasing the signal-to-noise ratio to over 20dB; and finally, dividing the signal into frames with overlapping window lengths T and t, and using a Hamming window to reduce spectral leakage.

[0076] Meanwhile, OBD data (i.e., the initial operating status data of the used car) is read in real time through the OBD-II interface, including parameters such as engine speed, oil temperature, intake air volume, and torque. The sampling frequency is synchronized with the voiceprint signal and is normalized to eliminate dimensional differences and ensure the accuracy of subsequent feature extraction.

[0077] S2, OBD-LSTM feature extraction, specifically includes:

[0078] A bidirectional long short-term memory (BiLSTM) network is employed to extract long-term dependencies from OBD time-series data. Specifically, OBD data (such as engine speed and temperature) stores the vehicle's operating state information. After normalization, this data forms an input vector, which is then fed into a two-layer BiLSTM structure. Each layer contains m memory units, where the forward LSTM captures the temporal dependencies from the current moment to future states, and the backward LSTM captures the temporal dependencies from the current moment to historical states. The bidirectional outputs are concatenated through fully connected layers and mapped to an n-dimensional temporal feature vector. A gating mechanism is used to mitigate the vanishing gradient problem, and bidirectional modeling is used to simultaneously analyze historical and future data, improving the feature extraction capability for sudden changes in vehicle state (such as rapid acceleration).

[0079] S3, audio multi-scale MFCC feature extraction, specifically includes:

[0080] First, the preprocessed audio signal (i.e., voiceprint data) is divided into frames of duration t. Windowing is applied to each frame to reduce edge effects, and the spectrum is obtained through Fast Fourier Transform (FFT). Then, the spectrum is mapped to the Mel scale using a Mel-scale filter bank. After calculating the logarithmic energy, 13-dimensional static MFCC features are extracted using Discrete Cosine Transform (DCT). To enhance dynamic features, first-order and second-order differences are further calculated to characterize the instantaneous rate of change and acceleration of the voiceprint, respectively. Finally, a 39×L-dimensional vector (13 static + 13 first-order + 13 second-order) is generated, where L is the number of MFCC frames within a given time period. Static features reflect the steady-state operating mode of the equipment, while dynamic features capture microscopic vibration anomalies (such as a sudden increase in high-frequency harmonics caused by bearing cracks).

[0081] S4, cross-modal feature fusion, specifically includes:

[0082] Cross-modal fusion is achieved through dynamic feature alignment and attention mechanisms. First, multi-scale MFCC features and OBD temporal features are mapped to a unified dimension through a fully connected layer. .in ,in It is the result of mapping the OBD temporal features after mapping. It is the input of OBD features. It is the bias of OBD features, and MFCC features are mapped using the same method. Finally, the two are concatenated and input into the multi-head attention mechanism layer.

[0083] A cross-model attention mechanism is employed to fuse concatenated features to achieve cross-modal complementarity, capture the correlation between different feature subspaces, and enhance the model's expressive power. The query vector Q is obtained by linearly projecting OBD features (such as rotation speed and temperature), representing the current context requiring attention. The key vector K is obtained by linearly projecting MFCC audio features, and the value vector V represents the retrieveable information of the MFCC audio features, obtained by linearly projecting MFCC audio features, representing the actual information content of the audio features. Attention weights are also included. Depend on It is calculated, among which It is the projection dimension of the query vector and the key vector, and the output is a fused feature vector obtained by weighting. This enables cross-modal complementarity while addressing the issue of differences in feature dimensions.

[0084] S5, Health Status Classification - Regression Output, specifically includes:

[0085] The fusion feature input is based on a Transformer-based multimodal fusion classifier, capturing the global dependencies of the fusion features. Multi-head attention is used to learn the dependencies of multiple subspaces in parallel. For example, considering the gradual wear and tear of a car engine due to long-term high-load operation, its audio features and historical OBD data have long-range dependencies. A self-attention mechanism is used to automatically associate the relationship between the current MFCC frame and historical OBD data. The classification output includes five categories: high risk, medium-high risk, medium risk, medium-low risk, and low risk, representing a health level classification.

[0086] The used car mortgage valuation model uses a health rating (high risk, medium-high risk, medium risk, medium-low risk, low risk) as its core, providing specific loan-to-value (LTV) parameters and information prompts. For example, if the health rating is low risk, the system will suggest a LTV of no more than 80%; if the health rating is medium risk, it will suggest a LTV of no more than 60%, and a third-party inspection report is required. This step, through end-to-end process automation, significantly improves the effectiveness of used car mortgage valuation and the risk control capabilities of banks' used car consumer loans.

[0087] It should be noted that when used cars are used as loan collateral, the wear and tear of internal components (such as wear, fatigue, misalignment, and bearing eccentricity) can affect their lifespan and thus their value. Currently, used car valuation cannot account for this wear and tear, relying solely on manual inspection or listening, which is costly. Traditional depreciation methods are inflexible, failing to account for internal component damage, affecting valuation accuracy, and manual measurement is costly, often requiring disassembly, which is time-consuming and labor-intensive. Therefore, this embodiment proposes a used car valuation method based on multimodal temporal feature fusion. This method acquires voiceprint data from mechanical equipment using a microphone array, combines filtering and speech enhancement techniques to reduce noise and improve voiceprint quality, extracts voiceprint features, fuses them with OBD data features, and classifies and outputs a health level. This method assesses the collateral value of used cars based on their health level, significantly improving loan review efficiency and valuation accuracy, providing strong support for financial institutions' risk management.

[0088] This embodiment can achieve at least one of the following effects: 1) This embodiment fully utilizes multimodal information (voiceprint data and OBD data) and the relationships between them. Through feature extraction and fusion, a health level classification model is constructed, which can comprehensively and accurately quantify and assess the overall health status of used cars, significantly improving the reliability and accuracy of financial institutions' valuation of used cars. 2) This embodiment proposes a used car valuation method based on multimodal temporal feature fusion, effectively solving the problems of valuation methods in related technologies that cannot consider internal component losses, high labor costs, and time and labor consumption. It can significantly improve review efficiency and provide strong support for banks' risk management.

[0089] This embodiment also provides a vehicle health determination device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0090] According to an embodiment of the present invention, an apparatus embodiment for implementing the above-described vehicle health determination method is also provided. Figure 4 This is a schematic diagram of a vehicle health determination device according to an embodiment of the present invention, as shown below. Figure 4 As shown, the aforementioned vehicle health determination device includes: a first data acquisition module 400, a voiceprint feature extraction module 402, a running feature extraction module 404, a feature fusion module 406, and a health level determination module 408, wherein:

[0091] The first data acquisition module 400 is used to acquire the vehicle's initial voiceprint data and initial operating status data. The initial voiceprint data is collected by a microphone array during vehicle operation.

[0092] The voiceprint feature extraction module 402 is connected to the first data acquisition module 400 and is used to extract features from the initial voiceprint data to obtain the target voiceprint features.

[0093] The feature extraction module 404 is connected to the voiceprint feature extraction module 402 and is used to extract features from the initial running state data to obtain the target running features.

[0094] The feature fusion module 406 is connected to the running feature extraction module 404 and is used to fuse the target voiceprint features and the target running features to obtain fused features.

[0095] The health level determination module 408 is connected to the feature fusion module 406 and is used to determine the health level of the vehicle based on the fused features.

[0096] According to an embodiment of the present invention, an apparatus embodiment for implementing the above-described vehicle valuation method is also provided. Figure 5 This is a schematic diagram of the structure of a vehicle valuation device according to an embodiment of the present invention, such as... Figure 5 As shown, the aforementioned vehicle health determination device includes: a health level acquisition module 500, a second data acquisition module 502, and a vehicle value assessment module 504, wherein:

[0097] The health level acquisition module 500 is used to acquire the health level of the vehicle, wherein the health level is obtained based on any of the above-mentioned vehicle health determination methods;

[0098] The second data acquisition module 502 is connected to the health level acquisition module 500 and is used to acquire the vehicle's basic information, vehicle mileage, and market data.

[0099] The vehicle valuation module 504 is connected to the second data acquisition module 502. It is used to obtain the vehicle valuation result based on the health level, basic vehicle information, vehicle mileage and market data, using a vehicle valuation model.

[0100] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0101] It should be noted that the first data acquisition module 400, voiceprint feature extraction module 402, runtime feature extraction module 404, feature fusion module 406, and health level determination module 408 mentioned above correspond to steps S102 to S110 in the embodiments. Similarly, the health level acquisition module 500, second data acquisition module 502, and vehicle value assessment module 504 correspond to steps S202 to S206 in the embodiments. The instances and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content disclosed in the above embodiments. It should also be noted that these modules, as part of the device, can run on a computer terminal.

[0102] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.

[0103] The aforementioned vehicle health determination device may further include a processor and a memory. The first data acquisition module 400, voiceprint feature extraction module 402, operation feature extraction module 404, feature fusion module 406, health level determination module 408, health level acquisition module 500, second data acquisition module 502, vehicle value assessment module 504, etc., are all stored in the memory as program modules, and the processor executes the aforementioned program modules stored in the memory to realize the corresponding functions.

[0104] The processor contains a core that retrieves the corresponding program modules from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0105] According to an embodiment of this application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein, when the program runs, the program controls the device where the non-volatile storage medium is located to execute any of the steps of the vehicle health determination method, or execute any of the vehicle valuation methods.

[0106] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals, and the non-volatile storage medium includes stored programs.

[0107] Optionally, during program execution, a program may be used to control the device containing the non-volatile storage medium to execute any of the above-described vehicle health determination method steps, or to execute any of the above-described vehicle valuation methods.

[0108] According to an embodiment of this application, an embodiment of a processor is also provided. Optionally, in this embodiment, the processor is used to run a program, wherein the program executes the steps of any of the above-described vehicle health determination methods, or executes the steps of any of the above-described vehicle value assessment methods.

[0109] According to an embodiment of this application, an embodiment of a computer program product is also provided, which, when executed on a data processing device, is adapted to execute a program that initializes the vehicle health determination method steps described above, or a program that performs any of the vehicle value assessment methods described above.

[0110] This invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the above-described vehicle health determination methods, or implements any of the above-described vehicle value assessment methods.

[0111] The order of the above embodiments of the present invention is merely for description and does not represent the superiority or inferiority of the embodiments.

[0112] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0113] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of modules described above can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between modules, and may be electrical or other forms.

[0114] The modules described above as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0115] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0116] If the aforementioned integrated modules are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0117] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for determining vehicle health, characterized in that, include: Acquire initial voiceprint data and initial operating status data of the vehicle, wherein the initial voiceprint data is collected by a microphone array during the operation of the vehicle; Feature extraction is performed on the initial voiceprint data to obtain the target voiceprint features; Feature extraction is performed on the initial running state data to obtain the target running features; The target voiceprint features and the target running features are fused to obtain fused features; Based on the fusion features, the health level of the vehicle is determined.

2. The method according to claim 1, characterized in that, The step of extracting features from the initial voiceprint data to obtain target voiceprint features includes: The initial voiceprint data is filtered to obtain the noise-reduced voiceprint data; The denoised voiceprint data is subjected to nonlinear enhancement processing to obtain enhanced voiceprint data; The enhanced voiceprint data is divided into frames according to a preset window length to obtain framed voiceprint data. The Hamming window method is used to perform weighted processing on the framed voiceprint data to obtain preprocessed voiceprint data. Feature extraction is performed on the preprocessed voiceprint data to obtain the target voiceprint features.

3. The method according to claim 1, characterized in that, The step of extracting features from the initial running state data to obtain target running features includes: The initial operating status data and the voiceprint data are aligned by sampling frequency to obtain aligned operating status data; The aligned running status data is normalized to obtain preprocessed running feature data; Feature extraction is performed on the preprocessed operational feature data to obtain the target operational features.

4. The method according to claim 1, characterized in that, The step of extracting features from the initial voiceprint data to obtain target voiceprint features includes: Based on the initial voiceprint data, frame segmentation processing is performed to obtain multiple frames of voiceprint data; Fourier transform is performed on the multiple frames of voiceprint data to obtain multiple voiceprint spectra, wherein the multiple voiceprint spectra correspond one-to-one with the multiple frames of voiceprint data; The multiple acoustic signature spectra are mapped to the Mel scale using a Mel scale filter to obtain the mapping results corresponding to each of the multiple acoustic signature spectra. The first voiceprint feature is obtained by performing a discrete cosine transform on the mapping results corresponding to the multiple voiceprint spectra. The first voiceprint feature is subjected to first-order difference processing to obtain the second voiceprint feature; and the first voiceprint feature is subjected to second-order difference processing to obtain the third voiceprint feature, wherein the target voiceprint feature includes the first voiceprint feature, the second voiceprint feature and the third voiceprint feature.

5. The method according to claim 1, characterized in that, The step of extracting features from the initial running state data to obtain target running features includes: A bidirectional long short-term memory network is used to extract the temporal dependencies in the initial running state data to obtain bidirectional temporal features; The bidirectional temporal features are dimensionality-reduced by passing them through a first fully connected layer to match the dimension of the target voiceprint features, thus obtaining the target operational features.

6. The method according to claim 1, characterized in that, The feature fusion of the target voiceprint features and the target operational features to obtain fused features includes: The target voiceprint features and the target operation features are mapped to the same feature space through a second fully connected layer to obtain the mapped voiceprint features and the mapped operation features. Using the mapped running features as the query vector and the mapped voiceprint features as the key vector, an attention weight is calculated using a cross-attention mechanism. The mapped voiceprint features are weighted according to the attention weights to obtain weighted voiceprint features; The weighted voiceprint features and the mapped running features are concatenated to obtain the fused features.

7. The method according to any one of claims 1 to 6, characterized in that, Determining the health level of the vehicle based on the fused features includes: The fused features are input into a multimodal health classifier based on the Transformer architecture to obtain the health level of the vehicle; The multimodal health classifier is used to capture the global dependencies of fused features through a multi-head self-attention mechanism and learn the interaction between different features. The multimodal health classifier is obtained by machine learning based on the fused features and health levels of multiple vehicles.

8. A vehicle health determination device, characterized in that, include: The first data acquisition module is used to acquire the vehicle's initial voiceprint data and initial operating status data, wherein the initial voiceprint data is acquired by a microphone array during the vehicle's operation. The voiceprint feature extraction module is used to extract features from the initial voiceprint data to obtain target voiceprint features; The feature extraction module is used to extract features from the initial running state data to obtain target running features; The feature fusion module is used to fuse the target voiceprint features and the target running features to obtain fused features; A health level determination module is used to determine the health level of the vehicle based on the fused features.

9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions adapted for loading by a processor and executing the vehicle health determination method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the vehicle health determination method according to any one of claims 1 to 7.