Data stream intelligent diagnosis method, system and device based on LSTM algorithm and medium
Through the intelligent data flow diagnosis method based on the LSTM algorithm and combined with the cluster analysis of the KNN algorithm, the problem of failure to detect new types of faults in the existing technology is solved, and efficient and accurate fault diagnosis is achieved.
Patent Information
- Application Number
- CN202510011382.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Existing data flow intelligently diagnoses vehicle failures and cannot detect new types of failures in a timely manner.
Using the intelligent data flow diagnosis method based on the LSTM algorithm, the LSTM algorithm is trained to collect vehicle operation data in real time and output the fault probability based on the vehicle operation data of the preset historical time period. When the fault probability is in the preset fuzzy probability range, further judgment is made by calculating the distance between the current vehicle operating data and the fault clustering center and the normal clustering center to determine whether there is a new fault.
It realizes the timely detection and processing of new fault types, improves the timeliness of fault diagnosis, and performs secondary confirmation of faults through cluster analysis of KNN algorithm, improving the accuracy of diagnosis.
Smart Images

Figure CN119939386A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data stream intelligent diagnosis, and in particular to a data stream intelligent diagnosis method, system, device and medium based on LSTM algorithm. Background Art
[0002] With the widespread use of heavy trucks in logistics and transportation, vehicle fault diagnosis and maintenance have become increasingly important. Traditional diagnostic methods often rely on manual inspection and static data analysis, resulting in low efficiency and untimely response. Signal flow analysis technology based on big data can improve the accuracy and timeliness of fault diagnosis by real-time monitoring and analysis of vehicle operating status.
[0003] The existing technologies for diagnosing vehicles using data streams mainly include the following: 1. The OBD-II (On-Board Diagnostics II, the second generation of onboard diagnostic systems) diagnostic system is widely used in modern cars and can monitor various sensors and systems of the vehicle in real time and provide fault codes. However, it can only perform customized fault detection based on the predefined fault codes provided, and cannot detect new types of faults in a timely manner. 2. Rule-based fault diagnosis uses preset rules to make fault judgments and can quickly identify common problems. The disadvantage is that it has poor adaptability to new problems or unknown faults and is prone to missing complex faults or special situations. 3. Mining real-time automobile fault data through data stream algorithms, extracting useful information and storing it in a temporary fault case library for updating the automobile fault knowledge base. Then, the knowledge base is updated using the temporary case library, and similarity matching is performed to solve the diagnostic problem. The disadvantage is that it cannot identify new problems or unknown faults that do not exist in the temporary fault case library. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present application provides a data stream intelligent diagnosis method, system and medium based on the LSTM algorithm to solve the problem that the existing data stream intelligent diagnosis of vehicle faults cannot detect new types of faults in a timely manner.
[0005] In a first aspect, the present application provides a data stream intelligent diagnosis method based on LSTM algorithm, the method comprising: Based on the vehicle operation data of a preset historical time period, an LSTM (Long Short-Term Memory, Long Short-Term Memory Network) algorithm is trained to obtain a trained LSTM algorithm; wherein a sigmoid activation function is used to output the fault probability; vehicle operation data of heavy trucks are collected in real time through on-board sensors and an OBD-II interface; the real-time collected vehicle operation data is used as input to the trained LSTM algorithm to obtain the fault probability; when the fault probability is within a preset fuzzy probability range, it is determined that there is a possibility of a new fault; the current vehicle operation data is input into a preset KNN algorithm to obtain a first distance between the current vehicle operation data and a fault cluster center and a second distance between the normal cluster center; when the difference between the first distance and the first distance is less than a preset difference threshold, it is determined that a new fault has occurred.
[0006] In one implementation of the present application, based on the vehicle operation data of a preset historical time period, the LSTM algorithm is trained to obtain a trained LSTM algorithm, specifically including: Process the vehicle operation data of a preset historical time period into time series data; set the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; set several layers of bidirectional LSTM layers to capture long-term dependencies in the time series; set the number of units and activation function in the fully connected layer of the LSTM algorithm; receive time series data through the LSTM algorithm input layer; wherein the time series data includes at least: the number of samples in each training session, the length of the time series, and the number of features in each time step; pass the data in the input layer to the bidirectional LSTM layer for processing; connect the output of the bidirectional LSTM layer to the fully connected layer for feature fusion, and then use the sigmoid activation function to output the fault probability.
[0007] In one implementation of the present application, when the difference between the first distance and the first distance is less than a preset difference threshold, after determining that a new fault has occurred, the method also includes: displaying the new fault on a user diagnosis interface; after receiving fault information marked on the new fault, training the LSTM algorithm again through the marked fault information and specific vehicle operation data to obtain a trained LSTM algorithm; and updating the fault cluster center and the normal cluster center using the marked fault information and specific vehicle operation data.
[0008] In one implementation of the present application, before inputting the current vehicle operation data into the preset KNN algorithm, the method further includes: training the preset KNN algorithm through the vehicle operation data of a preset historical time period to obtain a cluster center set of fault data and normal data.
[0009] In the second aspect, the present application provides a data stream intelligent diagnosis system based on the LSTM algorithm, the system comprising: A training module is used to train an LSTM algorithm based on vehicle operation data of a preset historical time period to obtain a trained LSTM algorithm; wherein a sigmoid activation function is used to output the fault probability; an acquisition module is used to collect vehicle operation data of heavy trucks in real time through on-board sensors and an OBD-II interface; a determination module is used to use the real-time collected vehicle operation data as input to a trained LSTM algorithm to obtain a fault probability; when the fault probability is within a preset fuzzy probability range, it is determined that there is a possibility of a new fault; the current vehicle operation data is input into a preset KNN algorithm to obtain a first distance between the current vehicle operation data and a fault cluster center and a second distance between the normal cluster center; when the difference between the first distance and the first distance is less than a preset difference threshold, it is determined that a new fault has occurred.
[0010] In one implementation of the present application, the training module includes a training unit for processing vehicle operation data of a preset historical time period into time series data; setting the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; setting several layers of bidirectional LSTM layers to capture long-term dependencies in the time series; setting the number of units and activation function of the fully connected layer of the LSTM algorithm; receiving time series data through the LSTM algorithm input layer; wherein the time series data includes at least: the number of samples in each training, the length of the time series, and the number of features in each time step; passing the data of the input layer to the bidirectional LSTM layer for processing; connecting the output of the bidirectional LSTM layer to the fully connected layer for feature fusion, and then using the sigmoid activation function to output the fault probability.
[0011] In one implementation of the present application, the system also includes an update module for displaying newly added faults on the user diagnosis interface; after receiving the fault information marked on the newly added fault, the LSTM algorithm is trained again through the marked fault information and specific vehicle operation data to obtain a trained LSTM algorithm; and the fault cluster center and the normal cluster center are updated using the marked fault information and specific vehicle operation data.
[0012] In one implementation of the present application, the system also includes a clustering module for training a preset KNN algorithm using vehicle operation data of a preset historical time period to obtain a cluster center set of fault data and normal data.
[0013] In the third aspect, the present application provides a data flow intelligent diagnosis device based on the LSTM algorithm, the device comprising: a processor; and a memory on which executable code is stored, and when the executable code is executed, the processor executes a data flow intelligent diagnosis method based on the LSTM algorithm such as any one of the above.
[0014] In a fourth aspect, the present application provides a non-volatile computer storage medium on which computer instructions are stored. When the computer instructions are executed, they implement a data flow intelligent diagnosis method based on an LSTM algorithm as described above.
[0015] Those skilled in the art can understand that the present application has at least the following beneficial effects: The present application provides a data stream intelligent diagnosis method, system and medium based on LSTM algorithm, which combines the time series data processing capability of LSTM algorithm and the cluster analysis advantage of KNN algorithm; the vehicle operation data of preset historical time period is used as training set. Design LSTM neural network including input layer, bidirectional LSTM layer and fully connected layer. Use sigmoid activation function to output fault probability in fully connected layer so that the prediction result is between 0 and 1. Set a fuzzy probability range (such as [0.4, 0.6]). When the fault probability falls within this range, it is considered that there is a possibility of new fault. When the fault probability is in the fuzzy range, introduce KNN algorithm for further judgment. Calculate the distance (first distance and second distance) between the current vehicle operation data and the known fault cluster center and the normal cluster center. If the distance between the current data and the fault cluster center is less than the distance between the current data and the normal cluster center, and the difference between the two is less than the preset difference threshold, it is determined that a new fault has occurred. In summary, the present application can timely discover and handle new fault types through real-time prediction of LSTM algorithm, thereby improving the timeliness of fault diagnosis. Combined with the cluster analysis of the KNN algorithm, the fault is reconfirmed to improve the accuracy of diagnosis. This solves the problem that the existing data stream intelligent diagnosis of vehicle faults cannot detect new types of faults in a timely manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0017] Figure 1 This is a flow chart of a data stream intelligent diagnosis method based on LSTM algorithm provided in an embodiment of the present application.
[0018] Figure 2 It is a schematic diagram of the internal structure of a data stream intelligent diagnosis system based on the LSTM algorithm provided in an embodiment of the present application.
[0019] Figure 3 This is a schematic diagram of the internal structure of a data stream intelligent diagnosis device based on the LSTM algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] It should be understood by those skilled in the art that the embodiments described below are only preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through the preferred embodiments. The preferred embodiments are only used to explain the technical principles of the present disclosure, and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work should still fall within the protection scope of the present disclosure.
[0021] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0022] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0023] The embodiment provides a data stream intelligent diagnosis method based on LSTM algorithm, such as Figure 1 As shown, the method provided in the embodiment of the present application mainly includes the following steps: Step 110: Based on the vehicle operation data of a preset historical time period, train the LSTM algorithm to obtain a trained LSTM algorithm.
[0024] Among them, the sigmoid activation function is used to output the failure probability.
[0025] In some embodiments, the LSTM algorithm is trained by: Process the vehicle operation data of a preset historical time period into time series data; set the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; set several layers of bidirectional LSTM layers to capture long-term dependencies in the time series; set the number of units and activation function in the fully connected layer of the LSTM algorithm; receive time series data through the LSTM algorithm input layer; wherein the time series data includes at least: the number of samples in each training session, the length of the time series, and the number of features in each time step; pass the data in the input layer to the bidirectional LSTM layer for processing; connect the output of the bidirectional LSTM layer to the fully connected layer for feature fusion, and then use the sigmoid activation function to output the fault probability.
[0026] It should be noted that this application also includes preset processing procedures, which can be specifically: Data cleaning: remove missing values and outliers to ensure the integrity and accuracy of the data. Feature engineering: extract relevant features from the original data, such as moving average, differential value, etc., to capture the time dependency of the data. Normalization: scale the data to the same range to improve the stability and efficiency of LSTM training.
[0027] Step 120: Collect the vehicle operation data of the heavy truck in real time through the vehicle-mounted sensor and the OBD-II interface.
[0028] Step 130: Use the real-time collected vehicle operation data as input to the trained LSTM algorithm to obtain the fault probability; when the fault probability is within a preset fuzzy probability range, determine that there is a possibility of a new fault.
[0029] It should be noted that the preset fuzzy probability range may be [0.4, 0.6].
[0030] Step 140: Input the current vehicle operation data into a preset KNN algorithm to obtain a first distance between the current vehicle operation data and the fault cluster center and a second distance between the current vehicle operation data and the normal cluster center; when the difference between the first distance and the first distance is less than a preset difference threshold, it is determined that a new fault has occurred.
[0031] In addition, the present application can determine the fault cluster center and the normal cluster center with historical data. When there is more than one first distance, the minimum value is taken; when there is more than one second distance, the minimum value is taken.
[0032] Before inputting the current vehicle operation data into the preset KNN algorithm, the method may further include: training the preset KNN algorithm through the vehicle operation data of a preset historical time period to obtain a cluster center set of fault data and normal data.
[0033] In some embodiments, when the difference between the first distance and the second distance is less than a preset difference threshold, after determining that a new fault has occurred, the method also includes: displaying the new fault on a user diagnostic interface; after receiving fault information marked on the new fault, training the LSTM algorithm again through the marked fault information and specific vehicle operation data to obtain a trained LSTM algorithm; and updating the fault cluster center and the normal cluster center using the marked fault information and specific vehicle operation data.
[0034] Based on the foregoing description, this embodiment collects the vehicle operation data of heavy trucks in a preset historical time period, including but not limited to engine parameters, vehicle speed, fuel consumption, temperature, etc., and constructs an LSTM neural network, whose output layer uses a sigmoid activation function to output the probability of a fault.
[0035] Use the prepared historical data to train the LSTM model until the model reaches satisfactory accuracy and stability. Save the trained LSTM model for subsequent use.
[0036] The vehicle operation data of heavy trucks is collected in real time through on-board sensors and OBD-II interfaces. The collected data is preprocessed in necessary steps, such as denoising and normalization, to ensure the quality and consistency of the data. The real-time collected vehicle operation data is used as the input of the trained LSTM model. The LSTM model outputs a probability value of a fault, which is between 0 and 1.
[0037] A preset fuzzy probability range is set, such as [0.4, 0.6], indicating that when the fault probability falls within this range, there is a possibility of a new fault.
[0038] If the fault probability output by LSTM is within the preset fuzzy probability range, the subsequent KNN algorithm is triggered for further judgment.
[0039] Prepare a cluster center set containing known fault data and normal data, which can be obtained by clustering historical data. Calculate the distance between the current vehicle operation data and the fault cluster center and the normal cluster center, which are recorded as the first distance and the second distance respectively. If the first distance is less than the second distance, and their difference is less than a preset difference threshold (for example, 0.1), it is determined that a new fault has occurred.
[0040] When a new fault is identified, the system can trigger an alarm mechanism to notify the driver or maintenance personnel. Combining the output of LSTM and the results of KNN, the fault can be further diagnosed and analyzed to determine the specific type and location of the fault.
[0041] In addition, this application Figure 2 A data stream intelligent diagnosis system based on LSTM algorithm is provided in the embodiment of the present application. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes: The training module 210 is used to train the LSTM algorithm based on the vehicle operation data of a preset historical time period to obtain a trained LSTM algorithm; wherein a sigmoid activation function is used to output the fault probability.
[0042] The training module 210 includes a training unit, which is used to process the vehicle operation data of a preset historical time period into time series data; set the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; set several layers of bidirectional LSTM layers to capture long-term dependencies in the time series; set the number of units and activation function of the LSTM algorithm fully connected layer; receive time series data through the LSTM algorithm input layer; wherein the time series data includes at least: the number of samples in each training, the length of the time series, and the number of features in each time step; pass the data of the input layer to the bidirectional LSTM layer for processing; connect the output of the bidirectional LSTM layer to the fully connected layer for feature fusion, and then use the sigmoid activation function to output the fault probability.
[0043] The acquisition module 220 is used to collect the vehicle operation data of the heavy truck in real time through the vehicle-mounted sensor and the OBD-II interface.
[0044] The determination module 230 is used to use the real-time collected vehicle operation data as the input of the trained LSTM algorithm to obtain the fault probability; when the fault probability is within the preset fuzzy probability range, determine the possibility of a new fault; input the current vehicle operation data into the preset KNN algorithm to obtain the first distance between the current vehicle operation data and the fault cluster center and the second distance between the normal cluster center; when the difference between the first distance and the first distance is less than the preset difference threshold, determine that a new fault has occurred.
[0045] The system also includes a clustering module for training a preset KNN algorithm through vehicle operation data in a preset historical time period to obtain a cluster center set of fault data and normal data.
[0046] The system also includes an update module, which is used to display newly added faults on the user diagnosis interface; after receiving the fault information marked on the newly added fault, the LSTM algorithm is trained again through the marked fault information and specific vehicle operation data to obtain a trained LSTM algorithm; and the fault cluster center and the normal cluster center are updated using the marked fault information and specific vehicle operation data.
[0047] The above is a method embodiment of the present application. Based on the same inventive concept, the present application embodiment also provides a data stream intelligent diagnosis device based on the LSTM algorithm. Figure 3 As shown, the device includes: a processor; and a memory on which executable codes are stored. When the executable codes are executed, the processor executes a data flow intelligent diagnosis method based on an LSTM algorithm as described in the above embodiment.
[0048] Specifically, the server side trains the LSTM algorithm based on the vehicle operation data of a preset historical time period to obtain a trained LSTM algorithm; wherein the sigmoid activation function is used to output the fault probability; the vehicle operation data of the heavy truck is collected in real time through the on-board sensor and the OBD-II interface; the real-time collected vehicle operation data is used as the input of the trained LSTM algorithm to obtain the fault probability; when the fault probability is within a preset fuzzy probability range, it is determined that there is a possibility of a new fault; the current vehicle operation data is input into a preset KNN algorithm to obtain a first distance between the current vehicle operation data and the fault cluster center and a second distance between the normal cluster center; when the difference between the first distance and the first distance is less than a preset difference threshold, it is determined that a new fault has occurred.
[0049] In addition, an embodiment of the present application further provides a non-volatile computer storage medium on which executable instructions are stored. When the executable instructions are executed, a data stream intelligent diagnosis method based on the LSTM algorithm as described above is implemented.
[0050] So far, the technical solutions of the present disclosure have been described in combination with the above multiple embodiments, but it is easy for those skilled in the art to understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principles of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or replacements to the relevant technical features. Any changes, equivalent replacements, improvements, etc. made within the technical concept and / or technical principle of the present disclosure will fall within the protection scope of the present disclosure.
Claims
1. A data stream intelligent diagnosis method based on LSTM algorithm, characterized in that: The method comprises: Based on the vehicle operation data of the preset historical time period, the LSTM algorithm is trained to obtain a trained LSTM algorithm; wherein the sigmoid activation function is used to output the fault probability; Real-time collection of vehicle operation data of heavy trucks through on-board sensors and OBD-II interface; The real-time collected vehicle operation data is used as the input of the trained LSTM algorithm to obtain the fault probability; when the fault probability is within the preset fuzzy probability range, it is determined that there is a possibility of a new fault; The current vehicle operation data is input into the preset KNN algorithm to obtain the first distance between the current vehicle operation data and the fault cluster center and the second distance between the normal cluster center; when the difference between the first distance and the first distance is less than the preset difference threshold, it is determined that a new fault has occurred.
2. The data stream intelligent diagnosis method based on LSTM algorithm according to claim 1 is characterized in that: Based on the vehicle operation data of the preset historical time period, the LSTM algorithm is trained to obtain the trained LSTM algorithm, which specifically includes: Processing vehicle operation data of a preset historical time period into time series data; Set the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; set several layers of bidirectional LSTM layers to capture long-term dependencies in time series; set the number of units and activation function of the fully connected layer of the LSTM algorithm; Receive time series data through the LSTM algorithm input layer; wherein the time series data at least includes: the number of samples in each training, the length of the time series, and the number of features in each time step; Pass the data of the input layer to the bidirectional LSTM layer for processing; The output of the bidirectional LSTM layer is connected to the fully connected layer for feature fusion, and then the sigmoid activation function is used to output the fault probability.
3. The data stream intelligent diagnosis method based on LSTM algorithm according to claim 1 is characterized in that: When the difference between the first distance and the second distance is less than the preset difference threshold, after determining that a new fault occurs, the method further includes: Display new faults on the user diagnosis interface; When receiving the fault information of the newly added fault label, the LSTM algorithm is trained again through the labeled fault information and the specific vehicle operation data to obtain the trained LSTM algorithm; And the fault cluster center and the normal cluster center are updated by using the marked fault information and specific vehicle operation data.
4. The data stream intelligent diagnosis method based on LSTM algorithm according to claim 1 is characterized in that: Before inputting the current vehicle operation data into the preset KNN algorithm, the method further includes: The vehicle operation data of the preset historical time period is used to train the preset KNN algorithm to obtain the cluster center set of fault data and normal data.
5. A data stream intelligent diagnosis system based on LSTM algorithm, characterized in that: The system comprises: A training module is used to train the LSTM algorithm based on the vehicle operation data of a preset historical time period to obtain a trained LSTM algorithm; wherein a sigmoid activation function is used to output the fault probability; The acquisition module is used to collect the vehicle operation data of heavy trucks in real time through the vehicle sensors and OBD-II interface; The determination module is used to use the real-time collected vehicle operation data as the input of the trained LSTM algorithm to obtain the fault probability; when the fault probability is within the preset fuzzy probability range, it is determined that there is a possibility of a new fault; the current vehicle operation data is input into the preset KNN algorithm to obtain the first distance between the current vehicle operation data and the fault cluster center and the second distance between the normal cluster center; when the difference between the first distance and the first distance is less than the preset difference threshold, it is determined that a new fault has occurred.
6. The data stream intelligent diagnosis system based on LSTM algorithm according to claim 5 is characterized in that: The training module includes training units, Used to process vehicle operation data of a preset historical time period into time series data; Set the number of units in the LSTM layer, whether to use dropout, and whether to use batch normalization; set several bidirectional LSTM layers to capture long-term dependencies in the time series; Set the number of units and activation function of the fully connected layer of the LSTM algorithm; Receive time series data through the LSTM algorithm input layer; wherein the time series data at least includes: the number of samples in each training, the length of the time series, and the number of features in each time step; Pass the data of the input layer to the bidirectional LSTM layer for processing; The output of the bidirectional LSTM layer is connected to the fully connected layer for feature fusion, and then the sigmoid activation function is used to output the fault probability.
7. The data stream intelligent diagnosis system based on LSTM algorithm according to claim 5 is characterized in that: The system further comprises an update module, Used to display newly added faults on the user diagnosis interface; When receiving the fault information of the newly added fault label, the LSTM algorithm is trained again through the labeled fault information and the specific vehicle operation data to obtain a trained LSTM algorithm; And the fault cluster center and the normal cluster center are updated by using the marked fault information and specific vehicle operation data.
8. The data stream intelligent diagnosis system based on LSTM algorithm according to claim 5 is characterized in that: The system also includes a clustering module, It is used to train the preset KNN algorithm through the vehicle operation data of the preset historical time period to obtain the cluster center set of fault data and normal data.
9. A data stream intelligent diagnosis device based on LSTM algorithm, characterized in that: The device comprises: processor; and a memory having executable codes stored thereon, which, when the executable codes are executed, enable the processor to execute a data stream intelligent diagnosis method based on an LSTM algorithm as described in any one of claims 1 to 4.
10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, the method for intelligent diagnosis of data stream based on LSTM algorithm as described in any one of claims 1 to 4 is implemented.