Tail gas emission prediction and fault early warning method fusing knowledge graph and machine learning
By integrating knowledge graphs and machine learning, high-precision prediction of vehicle exhaust emissions and rapid fault location are achieved, solving the problems of low prediction accuracy and poor location efficiency in existing technologies, and enabling rapid adaptation to new vehicle models.
Patent Information
- Application Number
- CN202511614526.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-06
AI Technical Summary
Existing technologies for monitoring and diagnosing vehicle exhaust emissions have low accuracy in emission prediction, cannot provide early warnings of excessive emission risks, have poor fault location efficiency, are prone to misdiagnosis or missed diagnosis of multiple faults, and have poor compatibility with new vehicle models.
By employing a method that integrates knowledge graphs and machine learning, and through modules for data collection, processing, knowledge graph construction, fault determination and localization, combined with a hybrid machine learning model of LSTM and XGBoost, multi-dimensional data analysis and fault early warning are achieved.
It achieves high-precision exhaust emission prediction, provides accurate early warning of excessive emissions 10 minutes in advance, shortens fault diagnosis time, reduces training time for adapting to new models by 50%, and improves the accuracy and efficiency of fault location.
Smart Images

Figure CN121071684A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of motor vehicle exhaust monitoring and fault diagnosis, and more particularly to an exhaust emission prediction and fault early warning method fusing a knowledge graph and machine learning. BACKGROUND
[0002] Under the background of the continuous growth of motor vehicle population and the increasingly stringent emission regulations of national standard VI B and above, the real-time monitoring demand for exhaust emission exceeding the standard and vehicle equipment failure is increasingly urgent, especially for fuel / hybrid motor vehicles, the dynamic change of exhaust composition is directly related to environmental quality and vehicle operation safety, and accurate prediction of emission trend and rapid positioning of faulty equipment become the core demand of the industry.
[0003] Currently, in the field of motor vehicle exhaust monitoring and fault diagnosis, the existing technologies mainly include an emission prediction method based on single time series data, a fault diagnosis method of manually matching fault codes, and a single machine learning model trained for a specific vehicle type, wherein the single time series data prediction only relies on historical exhaust concentration without combining operating parameters such as speed and throttle opening and environmental factors such as temperature and altitude; manual fault diagnosis relies on the experience of maintenance personnel, and misjudgment and missed judgment may occur due to feature cross interference when there are double / multiple faults; the single model needs to be trained from zero for new vehicle types, and has poor adaptability.
[0004] However, these existing technologies have obvious defects in actual application: first, the emission prediction accuracy is low, the use of historical exhaust data leads to high error, and the risk of exceeding the standard cannot be warned in advance, for example, the correlation between speed and NOx deviation rate cannot be used to predict the emission exceeding the standard under the condition of sudden acceleration; Second, the fault positioning efficiency is poor, and it takes more than 2 hours to check single fault, and double / multiple faults often appear due to the lack of weight distribution and region grouping logic, such as misjudging CO deviation + NOx deviation as a single three-way catalyst fault, ignoring the synergistic effect of throttle sticking. SUMMARY
[0005] Therefore, the embodiments of the present application provide an exhaust emission prediction and fault early warning method fusing a knowledge graph and machine learning, and the present application provides the following technical solutions: S1, taking fuel / hybrid motor vehicles with displacement greater than or equal to 0.8L and a unique vehicle identification code as monitoring objects, and collecting exhaust composition, vehicle operation, fault association and environmental association data in real time through a data acquisition module; S2, a data processing module processes the real-time collected data, calculates historical exhaust deviation rate and real-time exhaust deviation rate, and generates exhaust operation data set stored in a database; S3, according to the tail gas operation data set, combine the displacement vehicle historical fault data, build and update the knowledge graph association rule base through the knowledge graph construction module; S4, the difference type and the gap number are extracted from the tail gas component data, the fault type is judged through the fault judgment module, if it is a single fault, the feature word is generated, if it is a double / multiple fault, the feature word matrix is constructed, and the fault type and the equipment in the knowledge graph association rule base are matched; S5, according to the fault type and the equipment, the initial weight is distributed according to the function area through the fault positioning module, the fault area is positioned according to the weight order, and the fault positioning result and the priority repair suggestion are output; S6, the historical tail gas deviation rate and the vehicle operation data are combined to form input characteristics, and a pre-trained hybrid machine learning model is input to predict tail gas emission, and the emission deviation rate in the future period is output; S7, the warning module formulates multi-dimensional warning standards according to the fault positioning result and the tail gas emission deviation rate in the future period, and carries out visual display; S8, the difference between the emission deviation rate in the future period and the real-time tail gas deviation rate is monitored in real time through the model dynamic optimization module, and the hybrid machine learning model is corrected combined with the fault data.
[0006] The technical effects and advantages of the present application are: The present application collects multi-dimensional data such as tail gas components, vehicle operation, fault association and environmental association, combines isolated forest algorithm preprocessing and knowledge graph association rule base, adopts LSTM and XGBoost two-stage hybrid machine learning model, controls the emission prediction error at a low level, solves the problem of low prediction accuracy caused by relying on single time series data in the prior art, and realizes 10-minute early warning of emission over-standard risk; The present application processes single / double / multiple faults by layering, constructs a feature word matrix for double / multiple faults, and distributes the initial weight-dynamic correction-grouping superposition logic to locate the fault area, combines the knowledge graph feature word-equipment co-occurrence frequency to output the fault equipment, shortens the fault troubleshooting time, and solves the problem of misjudgment and missed judgment of multiple fault positioning in the prior art, and the problem of low efficiency; The present application provides initial parameters for new models by transfer learning, which can be adapted by fine-tuning a small amount of data, shortening the training time, and at the same time combining a real-time correction mechanism to solve the problem of weak model adaptability and slow convergence of new model training in the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 The present application is a whole structure schematic diagram.
[0008] Figure 2 The present application is a whole structure flow chart. DETAILED DESCRIPTION
[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0010] As attached Figure 1 The exhaust emission prediction and fault early warning method integrating knowledge graph and machine learning is characterized by including a data acquisition module, a data processing module, a database, a knowledge graph construction module, a fault determination module, a fault location module, an early warning module, and a model dynamic optimization module. The data acquisition module is used to collect exhaust gas composition, vehicle operation, fault correlation and environmental correlation data; The data processing module is used to process multi-dimensional datasets and generate exhaust gas operation datasets; The database is used to store exhaust gas operation datasets and other process data; The knowledge graph construction module is used to build and update the knowledge graph association library; The fault determination module is used to determine faults in a hierarchical manner and match fault types with devices in the knowledge graph; The fault location module is used to locate the fault area and output the fault location result and priority repair suggestions. The early warning module is used to formulate multi-dimensional early warning standards and visualize them. The model dynamic optimization module is used to make corrections by mixing machine learning models; As attached Figure 2 The exhaust emission prediction and fault early warning method integrating knowledge graphs and machine learning, as shown, includes the following steps in its specific implementation: S1. The monitoring targets are fuel / hybrid vehicles with an engine displacement of ≥0.8L and a unique vehicle identification number. The data acquisition module collects exhaust gas composition, vehicle operation, fault correlation and environmental correlation data in real time. It should be explained that the exhaust gas composition data includes carbon monoxide (CO), hydrocarbons (HC), nitrogen oxides (NOx), and particulate matter (PM). 2.5 / PM 10The NDIR infrared type and laser scattering type vehicle exhaust sensor and the portable emission test system PEMS are used for collecting, wherein the PEMS is used for backup collection when the vehicle sensor fails, so as to ensure uninterrupted data collection; the vehicle operation data including the rotation speed, the vehicle speed and the throttle opening degree are collected through the OBD system and the CAN bus; the fault correlation data including the fault code, the oxygen sensor response time and the three-way catalyst temperature are collected through the special diagnostic instrument; the environmental correlation data including the environmental temperature, the humidity and the altitude are collected through the vehicle environmental sensor and the GPS; the collection frequency is set to 30 seconds once during normal vehicle operation, and is adjusted to 10 seconds once during idling / rapid acceleration, so as to adapt to the high-frequency monitoring demand under the high-emission risk working condition.
[0011] S2, the data processing module processes the real-time collected data, calculates the historical exhaust deviation rate and the real-time exhaust deviation rate, and generates an exhaust operation data set stored in a database; It should be explained that the historical exhaust deviation rate calculation formula is: Xdev=(Xr-Xs) / Xs*100%, wherein Xdev is the exhaust deviation rate, Xr is the real-time value of the exhaust component collected, and Xs is the exhaust component threshold value corresponding to the national sixth b emission standard. It should be further explained that the real-time exhaust deviation rate is calculated using the same formula.
[0012] It should be further explained that the historical exhaust deviation rate refers to the average value or time series data of the exhaust component deviation rate calculated by the formula in the past collection period; the real-time exhaust deviation rate refers to the single deviation rate calculated based on the latest collected real-time value of the exhaust component in the current collection period.
[0013] It should be further explained that the exhaust operation data set includes: the historical exhaust deviation rate; the cleaned vehicle operation parameters: the rotation speed, the vehicle speed and the throttle opening degree; the basic correlation data: the unique vehicle identification code VIN, the GPS time stamp, the environmental temperature, the humidity and the altitude.
[0014] The processing process includes using the isolation forest algorithm to identify and eliminate the extreme abnormal values in the exhaust component, the vehicle operation parameter and the fault correlation data, so as to ensure the data validity; the missing data are differentiated according to the missing time length, the missing data with a missing time length <30 seconds are filled by using the forward filling method, and the missing data with a missing time length ≥30 seconds are filled by using the interpolation method of the same displacement and the same working condition vehicle type, so as to avoid the error caused by a single filling method; the time base of the historical exhaust deviation rate, the vehicle operation parameter and the environmental correlation data is unified based on the GPS time stamp, so as to ensure that the data at the same time node can be correlated and analyzed, and to provide a time sequence consistent data source for subsequent model prediction and fault judgment.
[0015] S3, according to the exhaust operation data set, the historical fault data of the same displacement vehicle type are combined, and a knowledge graph correlation rule library is constructed and updated through a knowledge graph construction module; It needs to be explained that the ontology design of knowledge graph contains 7 types of core entities and 10 types of association relationships. The 7 types of entities are motor vehicles, tail gas components, operating parameters, fault types, vehicle-mounted devices, fault areas, and environmental factors. The 10 types of association relationships include: motor vehicle - [produces] - tail gas component, tail gas component - [represents] - fault type, fault type - [corresponds to] - vehicle-mounted device, vehicle-mounted device - [belongs to] - fault area, motor vehicle - [runs and produces] - operating parameter, environmental factor - [influences] - tail gas component, fault area - [contains] - vehicle-mounted device, operating parameter - [correlates with] - tail gas component, fault type - [corresponds to] - repair scheme, and motor vehicle - [adapts to] - emission standard. Among them, fault area - [contains] - vehicle-mounted device is a reverse explicit association relationship that a certain fault area contains all vehicle-mounted devices, which is used to support the regional investigation in multi-fault scenarios.
[0016] Updating the knowledge graph association rule base needs to follow the process from the initial rule base construction to the real-time incremental update, ensuring the coherence of the rule base from the foundation to the dynamic optimization. The specific operations are as follows: Based on the historical fault data of the same displacement vehicle model, the initial association rule base is constructed: To provide basic association logic for the knowledge graph, it is necessary to integrate the historical fault data of the same displacement vehicle model that meets the standards. The data needs to meet two core requirements: one is that the data volume is ≥5000, and the other is that it covers fault cases of different vehicle ages of the same brand and displacement vehicle model to cover multiple vehicle conditions and ensure the universality of the rules. Through the association analysis of the tail gas deviation characteristics, device state parameters, and fault results in the historical data, standardized initial rules are generated and recorded in the knowledge graph, forming a basic association system. For example, when the CO deviation rate +20% and the oxygen sensor response time >0.5s, the associated fault type = oxygen sensor failure, and the associated device = oxygen sensor. For example, when the NOx deviation rate +30% and the throttle opening <8%, the associated fault type = throttle sticking, and the associated device = throttle.
[0017] Relying on real-time data collection and processing to realize incremental update of the rule base: With the technical support of Cypher statements of Neo4j database, an incremental update mechanism is established to synchronize with the data collection and processing rhythm: after completing a round of data collection and preprocessing, the standardized historical tail gas deviation rate, device state, and other data are written into the knowledge graph based on the vehicle unique identification code as the index. If the new data meets the existing rules, the corresponding entity attributes are supplemented, and the rule details are improved. If the new data presents an uncovered tail gas deviation-fault association mode, new association relationships are added to expand the dimension of the rule base. Subsequently, the rule base is iteratively updated based on new collected data, gradually improving the accuracy and comprehensiveness of fault matching.
[0018] S4, extract the difference type and the number of gaps from the tail gas composition data, determine the fault type through the fault determination module, if it is a single fault, generate a feature word, if it is a double / multiple fault, construct a feature word matrix, and match the fault type and the device in the knowledge graph association rule base; It needs to be explained that the fault determination module will first extract the difference type from the tail gas composition data, and determine the two core judging bases of difference type and gap number; among them, the difference type refers to the tail gas composition type and the corresponding deviation state exceeding the threshold of Guo Liu B standard, such as CO deviation rate +20% and NOx deviation rate +33%; the gap number is the total number of tail gas composition types exceeding the threshold in the same collection period, for example, only CO deviation gap number is 1, CO and NOx deviation gap number is 2.
[0019] For single fault, generate feature word, the determination process is divided into three steps: First, extract the difference type a and the gap number b in the current period, b=1 at this time; Second, generate feature word according to the format of difference type・gap number, that is, PM 2.5 deviation rate +15%・1; Third, match the feature word with the knowledge graph association rule base, if there is PM 2.5 deviation rate +15% and the temperature of the particulate filter <400℃→ the fault type is particulate filter blockage→ the associated device is the particulate filter preset rule, and the real-time collected fault associated data meets the conditions, which can directly output the corresponding fault type and associated device.
[0020] For double fault scene, the determination logic is further extended on the basis of single fault: First, extract two types of difference types a1, a2 and corresponding gap numbers b1, b2, b1=b2=1 at this time; Second, construct a 2×2 feature word matrix with difference type-gap number as row and column dimension; Third, assign initial weight according to the attribution relationship of tail gas composition-fault area in knowledge graph, for example, a1 corresponds to the initial weight of emission after treatment system 0.4, a2 corresponds to the initial weight of fuel supply system 0.5, and then modify the initial weight through the formula wi'=wi×(1+0.2×bi), where wi is the initial weight, bi is the gap number, and 0.2 is the influence coefficient of gap number on weight, which is determined by multivariate linear regression analysis of 5000 historical fault data of the same displacement vehicle, and the modified weight is 0.48 and 0.6 respectively; Fourth step, sort by modified weight, such as fuel supply system 0.6> exhaust aftertreatment system 0.48, preferentially match the associated rules of the area with higher weight, and combine the cross effects of the two types of faults, finally output the double fault result of fuel nozzle blockage + three-way catalyst activity decline.
[0021] Multi-fault scenario, the judgment of difference category n≥3 is based on the extension of double fault logic, the core is the superposition of grouping weight and priority sorting: First step, extract n types of difference categories a 1-n And the number of gaps b 1-n , construct an n×2 feature word matrix; Second step, according to the classification of fault area-function attribute in the knowledge graph, divide the n types of difference categories into the corresponding fault area group, for example, a1=NOx deviation rate+30%, a3=throttle opening abnormality into the intake system group, a2=CO deviation rate+20% into the exhaust system group; Third step, use the weight modification formula of double fault, calculate the modified weight of each difference category in each group and superimpose to get the total weight of each group, such as intake system group total weight=0.5×(1+0.2×1)+0.3×(1+0.2×1)=0.96, exhaust system group total weight=0.4×(1+0.2×1)=0.48; Fourth step, determine the main fault area according to the group total weight, and then combine the co-occurrence frequency of tail gas abnormal feature word-equipment, which is based on the effective fault case statistics of the same brand and same displacement vehicle in the past year, to prioritize the fault equipment in the area, for example, in the intake system group, the co-occurrence frequency of throttle sticking and NOx deviation rate+30% reaches 85%, the co-occurrence frequency of intake pressure sensor fault is 60%, finally output the multi-fault result of throttle sticking (0.85)→intake pressure sensor fault (0.6)→three-way catalyst aging (0.48), where the value in the bracket is the association degree, the higher the association degree, the higher the priority, which is convenient for subsequent maintenance.
[0022] S5, according to the fault type and equipment, assign an initial weight to each fault type and equipment through the fault positioning module, prioritize the fault area according to the weight, and output the fault positioning result and priority repair suggestion; The fault positioning module will first assign an initial weight to the fault region corresponding to each type of difference based on the double / multiple fault feature word matrix and the correlation between the exhaust composition-fault region knowledge graph. The assignment of the initial weight is based on the association frequency of the exhaust composition anomaly and the corresponding region fault in the historical fault data. For example, in a double fault scenario, if the historical association frequency of the CO deviation rate anomaly and the exhaust aftertreatment system fault reaches 70%, the initial weight is assigned as 0.7. If the association frequency of the NOx deviation rate anomaly and the intake system fault reaches 80%, the initial weight is assigned as 0.8, ensuring that the initial weight can reflect the historical association degree of the fault region.
[0023] In the weight correction stage, the correction formula wi' = wi x (1 + 0.2 x bi) is used, where wi is the initial weight, bi is the number of differences, and 0.2 is the influence coefficient of the number of differences on the weight. This coefficient is determined through multivariate linear regression analysis of 5000 historical fault data of the same displacement vehicle, and the initial weight is dynamically adjusted according to the number of differences in the current fault. Taking a multiple fault scenario as an example, if the intake system group contains NOx deviation rate +30%, b1 = 1, initial weight 0.8, and throttle opening anomaly, b2 = 1, initial weight 0.6, the corrected weights are 0.8 x (1 + 0.2 x 1) = 0.96 and 0.6 x (1 + 0.2 x 1) = 0.72, respectively. The corrected weights within the group are superimposed to obtain the total weight of the intake system group, which is 1.68. If the fuel system group contains HC deviation rate +25%, b3 = 1, initial weight 0.5, and corrected weight 0.6, the total weight of the group is 0.6. According to the total weight from high to low, the intake system (1.68) > fuel system (0.6), and the intake system with higher total weight is prioritized as the main fault region, and the secondary fault region is sorted according to the weight.
[0024] The output fault positioning result process is as follows: the fault positioning module combines the fault region-vehicle-mounted device attribution relationship in the knowledge graph to refine the located fault region into specific fault devices: for example, after positioning the intake system as the main fault region, according to the association relationship in the graph intake system-[contains]-throttle, intake pressure sensor, and the co-occurrence frequency of exhaust anomaly feature words-equipment, the throttle sticking is further determined as the main fault device, and the intake pressure sensor fault is the secondary fault device.
[0025] The priority repair suggestion is generated by comprehensively considering the relevance of the faulty equipment, the influence of the fault on the tail gas emission, and the repair difficulty; the higher the relevance, the greater the influence on the emission, and the lower the repair difficulty, the higher the repair priority; taking a double fault scenario as an example, if the throttle sticking has a relevance of 90%, causes a NOx deviation rate of +33%, and only needs to be replaced for repair, and the three-way catalyst activity decline has a relevance of 70%, causes a CO deviation rate of +20%, and needs to be disassembled for repair, then the priority repair suggestion is: 1. repair the throttle sticking, which is expected to take 30 minutes; 2. repair the three-way catalyst activity decline, which is expected to take 2 hours; at the same time, the suggestion also includes verification requirements after repair, such as continuously monitoring for 24 hours after repair to ensure that the tail gas deviation rate returns to within the threshold of the national VIb standard, and if the deviation rate is still more than 5%, the fault needs to be re-evaluated, forming a data closed loop to ensure the repair effect.
[0026] S6, the historical tail gas deviation rate and vehicle operation data are combined into input features, which are input into a pre-trained hybrid machine learning model for tail gas emission prediction, and the emission deviation rate in the future period is output; It needs to be explained that in addition to the tail gas deviation rate and vehicle operation data, two types of core features also need to be extracted to improve prediction accuracy, including: Trend features: based on the historical tail gas deviation rate data in the past 5 minutes, trend features are extracted through a sliding window method, including the mean, variance and slope of the deviation rate, which are used to reflect the dynamic change law of the tail gas composition; Correlation features: calculate the Pearson correlation coefficient p of the historical tail gas deviation rate and the vehicle operation data, and introduce the environmental correction coefficient of the operating data to form a multi-dimensional correlation feature set, which eliminates the influence of environmental interference on prediction.
[0027] The pre-trained hybrid machine learning model realizes high-precision emission prediction through a two-stage training architecture of LSTM and XGBoost, and the specific steps of outputting the emission deviation rate in the future period are as follows: A1, divide the extracted input features into training set and validation set in the ratio of 7:3, and process the data using Min-Max standardization; A2, build a 3-layer LSTM network, set the input layer dimension to the number of input features, the number of hidden layer neurons to 64, 32, and 16 respectively, use ReLU as the activation function, and set the dropout coefficient to 0.2 to prevent overfitting; A3, take the tail gas basic deviation rate in the next 10 minutes as the target, use mean square error as the loss function, and use Adam optimizer for training, iterate until the validation set loss is stable, and output the LSTM basic prediction result; A4, concatenate the LSTM output basic prediction result with the correlation features and real-time environmental data to form the final input feature matrix; A5, adopt XGBoost regression model, set the number of decision trees to 100, the maximum tree depth to 5, and the learning rate to 0.05, and optimize the hyperparameters through 5-fold cross-validation; A6, take the exhaust emission deviation rate of the next 10 minutes as the prediction target, and output the final prediction value.
[0028] It should be noted that the calculation formula of the final prediction value is: Xdev'(t+10)=fXGBoost(fLSTM(Xdev(t-5,t)),p(Xdev,Xrun),kenv), Where Xdev' is the prediction value of the exhaust deviation rate in the next 10 minutes, t+10 represents 10 minutes after the current time point, fXGBoost is the XGBoost model function, which is used to fuse multi-dimensional features and output accurate prediction deviation rate, fLSTM is the LSTM model function, which is a machine learning model used in the first stage of training and is specifically used to learn time series features, the input parameters in the parentheses are used to output the basic deviation rate in the next 10 minutes based on the exhaust deviation rate in the historical time period, providing time series dimension data for the XGBoost model, Xdev(t-5,t) is the time series of the past 5 minutes of exhaust deviation rate, p(Xdev,Xrun) is the Pearson correlation coefficient of exhaust deviation rate and running parameters, which is used to quantify the linear correlation degree of exhaust components and vehicle running parameters Xrun, and kenv is the environmental correction coefficient, which is used to correct the interference of environmental factors on exhaust deviation rate, and the value is calculated from the deviation of environmental data.
[0029] It should be further explained that the model outputs the emission deviation rate prediction result of the next 10 minutes every 10 minutes, covering CO, HC, NOx, PM 2.5 / PM 10 four types of core exhaust components, and also outputs the confidence of each component prediction result, based on the validation set accuracy, ≥90% for high reliability, 80%-90% for medium reliability, and <80% for triggering recalibration; the precision control is realized through a double mechanism, the 5-fold cross-validation in the training stage ensures that MAE≤5% and R²≥0.92; the real-time correction mechanism automatically calls the latest 100 data to fine-tune the XGBoost weight parameters and correct the deviation when the difference between the prediction value and the actual exhaust deviation rate is >8%.
[0030] S7, the warning module formulates multi-dimensional warning standards based on the fault location result and the exhaust emission deviation rate in the future period and visualizes the display; It should be explained that the warning module will combine the fault emergency level and the risk of exceeding the exhaust emission deviation rate in the future period to construct a two-dimensional warning standard, and then present the information through a visualization platform; the warning is divided into four levels: Red early warning for core equipment failure and future 10 minutes emission deviation rate ≥ 50% and so on, 15 minutes to push to more parties and require 24 hours repair; Orange early warning corresponds to key equipment failure and deviation rate 30%-50%, 30 minutes notice to the owner and repair station, suggest 48 hours repair; Yellow early warning for auxiliary equipment failure and deviation rate 10%-30%, 1 hour remind the owner to check within 7 days; Blue early warning is deviation rate 5%-10% or low correlation fault, 2 hours prompt maintenance.
[0031] At the same time, the old car early warning threshold is lowered by 10%, and the response time of the operating car is compressed by 50%, which is suitable for different car models and scenes; In addition, the early warning standard will further adapt to the difference between car models. For vehicles with age ≥ 5 years, the early warning trigger threshold is lowered by 10%. Because the old vehicle parts are aging, the emission baseline is higher, and a more sensitive early warning mechanism is needed. For example, the original yellow early warning deviation rate is 10%-30%, and the old car is adjusted to 8%-28%. For operating vehicles, the response time is compressed by 50%. For example, the original orange early warning response time is 30 minutes, and the operating car is adjusted to 15 minutes, to ensure that the early warning fits the actual use scene of different vehicles.
[0032] The visualization platform supports three types of roles to view, and the overview displays the vehicle early warning level with a map superimposed color icon, and the VIN code, over-standard components can be seen floating; The detail page locates the fault on the left side, and the right side displays the emission prediction curve and maintenance suggestion; The traceability module allows professionals to view 24-hour raw data, fault judgment records, etc., and review the early warning logic; In addition, the owner interface simplifies the information, the maintenance interface adds parts models and other practical content, and the supervision interface focuses on statistical analysis. Early warning information will also be pushed through multiple channels such as SMS and APP, forming a response loop.
[0033] S8, through the model dynamic optimization module, real-time monitoring of the difference between the emission deviation rate in the future period and the real-time tail gas deviation rate, combined with fault data to correct the hybrid machine learning model.
[0034] The core of the model dynamic optimization module is to adjust the hybrid machine learning model in real time according to the prediction deviation feedback and fault data supplement in two dimensions, to ensure that the model adapts to the change of vehicle emission characteristics for a long time.
[0035] First of all, the judgment condition of triggering optimization: the module will continuously compare the prediction value of the future 10 minutes emission deviation rate with the real-time tail gas deviation rate, when the difference between the two is > 8%, or after the single fault judgment is completed, the optimization process is started immediately: the former is for the case of prediction accuracy decline, and the latter is to supplement the model's learning ability for abnormal working conditions through fault data.
[0036] For the optimization triggered by the prediction deviation, the process is divided into two steps: first, extract the latest 100 data, regenerate the time series features and correlation features, and use them as the training data set for model fine-tuning; second, fix the layer structure and core parameters of the LSTM network, only retrain the weight parameters of the XGBoost model, minimize the deviation between the predicted value and the actual value through gradient descent method, and at the same time, keep the original 5-fold cross-validation mechanism to ensure that the MAE after fine-tuning is less than or equal to 5%, the whole process takes less than 5 minutes, avoiding affecting real-time prediction.
[0037] For the optimization triggered by the fault data: first, label the verified fault data as fault samples and add them to the model's training sample library; second, fine-tune the LSTM model, only update the connection weights of the input layer and the first hidden layer, let the model learn the trend of exhaust composition changes before the fault occurs, and update the feature importance weight of the XGBoost model to improve the influence of fault correlation features on the prediction results, so that the model can capture the early warning signs of faults.
[0038] In addition, the module also provides migration optimization support for new vehicle models: when a new vehicle model completes 3 fault determinations for the first time, the module uses the model parameters of the mature vehicle with the same displacement as the initial value, and performs transfer learning with 300-500 measured data of the new vehicle model, freezes the bottom feature extraction network of the model, and only trains the top output layer parameters, so that the convergence time of the new vehicle model is shortened by more than 50%, and the adaptation accuracy is improved to more than 85%.
[0039] After each optimization is completed, the module automatically records the optimization log and synchronously updates it to the knowledge graph association rule library, supplements the association relationship, and provides a reference for subsequent similar optimization.
[0040] Secondly, in the drawings of the disclosed embodiments, only the structures related to the disclosed embodiments are involved, other structures can be referred to the usual design, and in the case of no conflict, the same embodiment and different embodiments of the present application can be combined with each other; Finally, the above-mentioned is only the preferred embodiment of the present application, and is not used to limit the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for exhaust emission prediction and fault early warning by fusing knowledge graph and machine learning, characterized in that, The method comprises a data acquisition module, a data processing module, a database, a knowledge graph construction module, a fault determination module, a fault positioning module, an early warning module, and a model dynamic optimization module. S1. A motor vehicle with a displacement of ≥0.8L and a unique vehicle identification code is taken as a monitoring object, and real-time collection of tail gas composition, vehicle operation, fault correlation, and environmental correlation data is performed through the data acquisition module; S2. The data processing module processes the real-time collected data, calculates the historical tail gas deviation rate and the real-time tail gas deviation rate, and generates a tail gas operation data set stored in the database; S3. According to the tail gas operation data set, the historical fault data of the same displacement vehicle type is combined, and the knowledge graph construction module is used to construct and update the knowledge graph association rule base; The tail gas operation data set includes: historical tail gas deviation rate; cleaned vehicle operation parameters: speed, vehicle speed, throttle opening; basic correlation data: unique vehicle identification code VIN, GPS timestamp, environmental temperature, humidity, and altitude; S4. The difference types and gap quantities are extracted from the tail gas composition data, the fault type is determined through the fault determination module, if it is a single fault, the feature words are generated, if it is a double / multiple fault, the feature word matrix is constructed, and the fault type and equipment in the knowledge graph association rule base are matched; The knowledge graph association rule base specifically includes: 7 types of core entities and 10 types of association relationships: 7 types of entities are motor vehicles, tail gas composition, operation parameters, fault type, vehicle-mounted equipment, fault area, and environmental factors; 10 types of association relationships include: motor vehicle-[produces]-tail gas composition, tail gas composition-[represents]-fault type, fault type-[corresponds to]-vehicle-mounted equipment, vehicle-mounted equipment-[belongs to]-fault area, motor vehicle-[produces]-operation parameter, environmental factor-[influences]-tail gas composition, fault area-[contains]-vehicle-mounted equipment, operation parameter-[correlates with]-tail gas composition, fault type-[corresponds to]-repair scheme, and motor vehicle-[adapts to]-emission standard; S5. According to the fault type and equipment, the fault positioning module assigns an initial weight according to the functional area, locates the fault area according to the weight order, and outputs the fault positioning result and the priority repair suggestion; S6. The historical tail gas deviation rate and the vehicle operation data are combined to form input features, which are input into a pre-trained hybrid machine learning model to predict tail gas emissions, and the future period emission deviation rate is output; The specific steps of outputting the future period emission deviation rate are as follows: A1. The extracted input features are divided into a training set and a validation set in a ratio of 7:3, and the data is standardized; A2. A 3-layer LSTM network is built, the input layer dimension is set to the number of input features, the number of hidden layer neurons is 64, 32, and 16 respectively, the activation function is ReLU, and the dropout coefficient is set to 0.2 to prevent overfitting; A3. The tail gas basic deviation rate predicted for the next 10 minutes is taken as the target, the mean square error is used as the loss function, the Adam optimizer is used for training, and the LSTM basic prediction result is output after the validation set loss is stable. A4. The basic prediction results output by LSTM are concatenated with the associated features and real-time environmental data to form the final input feature matrix; A5. Using the XGBoost regression model, the number of decision trees is set to 100, the maximum tree depth is 5, and the learning rate is 0.
05. The hyperparameters are optimized through 5-fold cross-validation. A6. Using the exhaust emission deviation rate of the next 10 minutes as the prediction target, output the final predicted value; S7. The early warning module formulates multi-dimensional early warning standards and displays them visually based on the fault location results and the exhaust emission deviation rate in the future period. S8. The model dynamic optimization module monitors the difference between the emission deviation rate in future periods and the real-time exhaust gas deviation rate in real time, and corrects the hybrid machine learning model by combining fault data. 2.The method of claim 1, wherein the method further comprises: The specific formula for calculating the historical exhaust gas deviation rate is as follows: Xdev = (Xr - Xs) / Xs × 100%, Where Xdev is the historical exhaust gas deviation rate, Xr is the real-time value of the collected exhaust gas components, and Xs is the exhaust gas component threshold corresponding to the China VI b emission standard. 3.The method of claim 1, wherein the method further comprises: The process of generating feature words consists of three steps: The first step is to extract the difference type 'a' and the difference quantity 'b' within the current period, where b=1. The second step is to generate feature words according to the format of difference type and number of differences; The third step is to match the feature word with the knowledge graph association rule base. If the feature word exists in the rule base and the fault association data collected in real time meets the conditions, the corresponding fault type and associated device can be directly output. 4.The method of claim 1, wherein the method further comprises: If the problem involves a dual fault, the process for constructing the feature word matrix is further extended based on the single fault scenario: The first step is to extract the two types of differences a1 and a2 and the corresponding difference quantities b1 and b2, at which point b1=b2=1; The second step is to construct a 2×2 feature word matrix with the difference type and the number of differences as the row and column dimensions; The third step is to assign initial weights based on the attribution relationship between exhaust gas components and fault areas in the knowledge graph. The fourth step is to sort the data according to the corrected weights, prioritize matching the association rules of regions with higher weights, and combine the cross-influence of the two types of faults to finally output the double fault result.
5. The method of claim 1, wherein the method further comprises: If there are multiple faults, the determination of the feature word matrix is based on the extension of the dual-fault logic: First, extract n kinds of difference categories a 1-n With the number of gaps b 1-n , build an n x 2 feature word matrix; The second step is to classify the n types of differences into the corresponding fault region groups according to the classification of fault region-functional attributes in the knowledge graph. The third step is to use the weight correction formula for dual faults to calculate the correction weights for each type of difference in each group and then sum them up to obtain the total weight for each group. The fourth step is to determine the main fault areas by sorting them according to the total weight of the group, and then prioritize the faulty equipment in the area by combining the co-occurrence frequency of exhaust gas abnormality feature words and equipment. 6.The method of claim 1, wherein the method further comprises: The formula for calculating the final predicted value is as follows: Xdev'(t + 10) = f XGBoost (f LSTM (Xdev(t - 5, t)), p(Xdev, Xrun), kenv), where Xdev' is the prediction of exhaust deviation rate in the next 10 minutes, t + 10 represents 10 minutes after the current time point, f XGBoost is the XGBoost model function, which is used to fuse multi-dimensional features and output the predicted deviation rate, f LSTM is the LSTM model function, which is used to learn the time series features, Xdev(t-5, t) is the time series of exhaust deviation rate in the past 5 minutes, p(Xdev, Xrun) is the Pearson correlation coefficient of historical exhaust deviation rate and vehicle running data, which is used to quantify the linear correlation degree of exhaust components and vehicle running data Xrun, and Kenv is the environmental correction coefficient, the value of which is calculated from the deviation of environmental data.
Citation Information
Patent Citations
Road motor vehicle tail gas high emission early warning method based on PA-LSTM network
CN112949930A
Mixed gas deviation self-learning method and system, readable storage medium and electronic equipment
CN113239966A
Machine learning model iteration method and system of intelligent operation and maintenance system
CN117371553A
Steam turbine vibration fault diagnosis system fused with deep learning
CN120180040A
Charger fault diagnosis method based on knowledge graph
CN120494796A