A Reinforcement Learning-Based Intelligent Lifespan Prediction Method and System for Heavy-Duty Vehicle Electronic Components
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-08-11
AI Technical Summary
现有电子元器件寿命预测方法多依赖固定阈值、经验模型或单一传感参数进行判断,难以充分反映重型汽车实际运行过程中多源时序数据之间的耦合关系,也难以准确刻画元器件由正常、轻度退化、严重退化至失效的动态演化过程
[0012] Compared to existing technologies, the beneficial effects of this invention include: employing a reinforcement learning-based intelligent lifespan prediction method and system for heavy-duty automotive electronic components, this invention acquires a data stream of electronic component operation monitoring data, which includes identifiers, a set of time-series load condition parameters, a set of time-series environmental stress parameters, and a set of time-series electrical response parameters; it performs degradation trajectory reinforcement modeling on the data stream, generating a set of state degradation trajectory vectors and a lifespan stage transition probability array; based on these, it constructs a lifespan prediction strategy network and trains it to convergence using the lifespan stage transition probability array to obtain a set of weight parameters; it calls the converged network to process the state degradation trajectory vector set and outputs the remaining effective lifespan prediction distribution; based on the prediction distribution, it generates a set of component replacement scheduling instructions, each containing a sequence of identifiers to be replaced and a corresponding time window. This invention transforms multi-dimensional time-series monitoring data into degradation state features and stage transition probabilities, captures the uncertainty of dynamic degradation transitions through a reinforcement learning strategy network, outputs probabilistic lifespan predictions, and generates optimized predictive replacement scheduling, significantly improving lifespan prediction accuracy and maintenance efficiency, and reducing the risk of unplanned downtime.
Smart Images

Figure CN122548149A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for intelligent prediction of the lifespan of heavy-duty automotive electronic components based on reinforcement learning. Background Technology
[0002] Heavy-duty vehicles operate under conditions of high load, strong vibration, high temperature, low temperature, humidity, dust, and complex electromagnetic fields. Their electronic components are susceptible to the combined effects of load conditions, environmental stress, and electrical response fluctuations, leading to gradual performance degradation. Existing methods for predicting the lifespan of electronic components mostly rely on fixed thresholds, empirical models, or single sensor parameters. These methods fail to adequately reflect the coupling relationships between multi-source time-series data during the actual operation of heavy-duty vehicles, and also struggle to accurately depict the dynamic evolution of components from normal operation, mild degradation, severe degradation, to failure. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for intelligent prediction of the lifespan of heavy-duty automotive electronic components based on reinforcement learning.
[0004] In a first aspect, embodiments of the present invention provide a method for intelligent prediction of the lifespan of heavy-duty automotive electronic components based on reinforcement learning, the method comprising:
[0005] The operation monitoring data stream of heavy-duty vehicle electronic components is acquired. The operation monitoring data stream includes electronic component identifiers, a set of time-series load condition parameters associated with the electronic component identifiers, a set of time-series environmental stress parameters associated with the electronic component identifiers, and a set of time-series electrical response parameters associated with the electronic component identifiers.
[0006] The operation monitoring data stream is subjected to degradation trajectory enhancement modeling processing. Based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier are generated.
[0007] A lifetime prediction policy network is constructed based on the state degradation trajectory vector set and the lifetime stage transition probability array, and the lifetime prediction policy network is trained using the lifetime stage transition probability array to obtain a set of weight parameters for the lifetime prediction policy network that has been trained and converged.
[0008] The lifetime prediction strategy network, which has been trained and converged, is invoked to process the set of state degradation trajectory vectors to generate the remaining effective lifetime prediction distribution corresponding to the electronic component identifier.
[0009] A set of component replacement scheduling instructions is generated based on the predicted distribution of remaining effective lifetime. The set of component replacement scheduling instructions includes a sequence of identifiers of electronic components to be replaced and corresponding time window information.
[0010] Secondly, embodiments of the present invention provide a heavy-duty vehicle electronic component life prediction system based on reinforcement learning, including at least one service node;
[0011] The service node includes a storage unit and a computing unit; the storage unit is used to store program code; the computing unit is used to run the program code to execute the reinforcement learning-based intelligent prediction method for the lifespan of heavy-duty vehicle electronic components as described in the first aspect.
[0012] Compared to existing technologies, the beneficial effects of this invention include: employing a reinforcement learning-based intelligent lifespan prediction method and system for heavy-duty automotive electronic components, this invention acquires a data stream of electronic component operation monitoring data, which includes identifiers, a set of time-series load condition parameters, a set of time-series environmental stress parameters, and a set of time-series electrical response parameters; it performs degradation trajectory reinforcement modeling on the data stream, generating a set of state degradation trajectory vectors and a lifespan stage transition probability array; based on these, it constructs a lifespan prediction strategy network and trains it to convergence using the lifespan stage transition probability array to obtain a set of weight parameters; it calls the converged network to process the state degradation trajectory vector set and outputs the remaining effective lifespan prediction distribution; based on the prediction distribution, it generates a set of component replacement scheduling instructions, each containing a sequence of identifiers to be replaced and a corresponding time window. This invention transforms multi-dimensional time-series monitoring data into degradation state features and stage transition probabilities, captures the uncertainty of dynamic degradation transitions through a reinforcement learning strategy network, outputs probabilistic lifespan predictions, and generates optimized predictive replacement scheduling, significantly improving lifespan prediction accuracy and maintenance efficiency, and reducing the risk of unplanned downtime. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating the steps of the intelligent prediction method for the lifespan of heavy-duty automotive electronic components based on reinforcement learning provided in this embodiment of the invention.
[0015] Figure 2 This is a schematic diagram of the lifetime prediction strategy network construction and training process provided in an embodiment of the present invention;
[0016] Figure 3 A schematic diagram of the degradation trajectory enhancement modeling process provided for the implementation of this invention;
[0017] Figure 4 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0020] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the intelligent prediction method for the lifespan of heavy-duty vehicle electronic components based on reinforcement learning provided in this embodiment. The following is a detailed description of this intelligent prediction method for the lifespan of heavy-duty vehicle electronic components based on reinforcement learning.
[0021] Step S201: Obtain the operation monitoring data stream of heavy-duty vehicle electronic components. The operation monitoring data stream includes electronic component identifiers, a set of time-series load condition parameters associated with the electronic component identifiers, a set of time-series environmental stress parameters associated with the electronic component identifiers, and a set of time-series electrical response parameters associated with the electronic component identifiers.
[0022] The server, acting as the execution entity, is deployed within the remote operation and maintenance platform of the heavy-duty truck fleet, communicating with the vehicle's onboard gateway, electronic control unit, sensor acquisition module, and maintenance dispatch system. When heavy-duty trucks operate in environments involving heavy-load transportation in mining areas, long-distance trunk line transportation, port traction, long downhill braking, and alternating high and low temperature conditions, the engine controller, power management module, brake controller, transmission controller, power relays, sensor interface boards, and solenoid valve drive modules on the vehicles continuously upload operational data. The operational monitoring data stream received by the server uses electronic component identifiers as indexes; for example, "HV-TRUCK-023-ECU-PWR-MOS-05" indicates the 5th power switching device in the 23rd heavy-duty truck's engine control unit. The server synchronously acquires time-series load condition parameter sets, time-series environmental stress parameter sets, and time-series electrical response parameter sets based on this identifier. The time-series load condition parameter set includes traction load, engine torque, braking trigger frequency, peak supply current, number of start-stop impacts, and transient power fluctuations; the time-series environmental stress parameter set includes ambient temperature, humidity, vibration acceleration, altitude air pressure, controller cavity temperature rise, and heat dissipation duct temperature difference; the time-series electrical response parameter set includes output voltage ripple, leakage current, on-state voltage drop, insulation impedance, switching delay, sampling bias, and equivalent series resistance. After receiving the data, the server performs multi-source heterogeneous sampling frequency alignment processing, mapping millisecond-level electrical response data, second-level load condition data, and minute-level environmental stress data to the same time axis, forming a time-aligned operation monitoring data stream, and retaining the original timestamp, interpolation marker, outlier marker, and data credibility marker, providing a unified input for subsequent degradation modeling and reinforcement learning decision-making.
[0023] Step S202: Perform degradation trajectory enhancement modeling processing on the operation monitoring data stream, and generate a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set and the time-series electrical response parameter set.
[0024] The server first performs load impact event detection processing on the time-series load condition parameter set. For example, when a power relay starts under full load, the traction current increases from 80A to 210A within 2 seconds. The server records this time point as the load mutation time point and 130A as the load mutation amplitude value, generating a load impact event sequence containing time and amplitude labels. The server then performs environmental stress fluctuation segmentation processing on the time-series environmental stress parameter set. The process of the engine compartment temperature continuously rising from 45℃ to 92℃, the process of the controller housing vibration changing from a low-frequency stable state to a high-frequency impact state, and the process of rapid increase in humidity under rain and snow conditions are respectively divided into environmental stress fluctuation segments. Each segment includes the start time, end time, and environmental stress change gradient sequence. The server also performs response anomaly offset identification processing on the time-series electrical response parameter set. For example, if it detects that the steady-state value of the output voltage of a power management module has shifted from 24.1V to 23.6V for three consecutive days, accompanied by increased ripple and increased on-state voltage drop, the server records the electrical response offset start timestamp and the cumulative offset, generating an electrical response offset feature sequence.
[0025] In one optional implementation, the server uses a sliding window approach to extract degraded input features. For load parameters, the server calculates the change in load parameters within a window of length T1. When the change exceeds a preset load mutation threshold and the duration exceeds a preset hold time, the corresponding time point is determined as the load mutation time point, and the difference between the mean values of the stable windows before and after the mutation is used as the load mutation amplitude value. For environmental stress parameters, the server calculates the gradient of changes at adjacent time points. When the gradient continuously exceeds the environmental gradient threshold, the start time of the environmental stress fluctuation segment is determined, and the end time is determined when the gradient recovers to below the recovery threshold. For electrical response parameters, the server uses the initial stable operating data of the components as a reference interval. When the mean value of the current window continuously deviates from the reference interval and the deviation direction conforms to the degradation direction, the starting point of the electrical response steady-state offset is determined, and the offset of each window is accumulated to obtain the cumulative electrical response steady-state offset.
[0026] The server inputs the load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence into the degradation trajectory encoder. This degradation trajectory encoder includes a multi-layer temporal convolutional operation layer and a channel cross-attention operation layer. In one optional embodiment, the degradation trajectory encoder includes an input embedding layer, a multi-layer temporal convolutional operation layer, a channel cross-attention operation layer, and a feature output layer. The input embedding layer maps the load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence into input feature vectors of the same dimension, which can be set to 32, 64, or 128 dimensions. The multi-layer temporal convolutional operation layer includes 2 to 6 one-dimensional temporal convolutional layers, with kernel lengths of 3, 5, or 7, used to extract temporal features such as short-term impact, sustained high temperature, vibration accumulation, and electrical drift. The channel cross-attention operation layer treats the load impact temporal feature map, the environmental stress temporal feature map, and the electrical response temporal feature map as different channels, calculates the similarity between channels, generates channel interaction weights based on the similarity, and performs weighted fusion of the features of each channel to obtain a multi-source stress coupling feature representation set. Multi-layer temporal convolutional computation layers extract multi-scale features such as short-term current surges, transient overvoltages, vibration spikes, continuous high temperatures, periodic heavy loads, and long-term electrical drift, forming load impact time-series feature maps, environmental stress time-series feature maps, and electrical response time-series feature maps, respectively. Channel cross-attention computation layers further calculate the coupling relationships between multi-source features, such as identifying the correlation between sustained high-temperature segments and increased conduction voltage drop, and the correlation between vibration shocks and sudden increases in connector contact resistance, assigning higher interaction weights to strongly correlated channels to obtain a set of multi-source stress coupling feature representations. The server also constructs a degradation property constraint layer, injecting material property parameters, process property parameters, and operating limit property parameters into the degradation trajectory encoder. For example, the upper limit of power device junction temperature, solder joint thermal fatigue characteristics, capacitor withstand voltage range, and relay contact wear patterns are used as physical constraints. A degradation property constraint gating unit performs physical feasibility judgment on the time-series features and generates stress damage accumulation path curves.
[0027] The server performs degradation stage clustering analysis based on multi-source stress coupling characteristics and stress damage accumulation path curves. For a power control board, the server classifies stages with stable electrical response and normal temperature rise into the first degradation stage cluster, stages with frequent temperature rise and increasing voltage ripple into the second degradation stage cluster, stages with continuously increasing leakage current and decreasing insulation impedance into the third degradation stage cluster, and stages with significantly increased switching delay and approaching the operating limit into the fourth degradation stage cluster. The server counts the transition frequency between different degradation stage clusters within adjacent time windows and normalizes it to obtain a lifetime stage transition probability array. The row and column indices of this array correspond to different degradation stage clusters, and each element represents the probability of transitioning from one degradation stage to another. The server also concatenates the central feature vector of each degradation stage cluster with the start and end time labels of that stage to form a state degradation trajectory vector set. The server can also perform cross-component degradation mode migration on the transfer probability arrays of different electronic components, extract shared degradation structures, form a set of general degradation mode representation vector bases, and inject them into the channel cross-attention operation layer to enhance the degradation stage recognition capability in the cold start scenario of new electronic components.
[0028] In one optional implementation, the server normalizes each temporal feature vector in the multi-source stress coupling feature representation set and calculates the feature distance between any two temporal feature vectors. The feature distance can be Euclidean distance or cosine distance. When the feature distance between two temporal feature vectors is less than a preset stage similarity threshold, and the time interval between them does not exceed a preset stage continuous time threshold, the server assigns them to the same degradation stage cluster. The central feature vector of each degradation stage cluster is obtained by averaging all feature vectors within the cluster. The server counts the transfer frequency between degradation stage clusters within adjacent time windows, and denotes the frequency of transfer from the i-th degradation stage cluster to the j-th degradation stage cluster as N. ij And calculate the transition probability as follows: P ij =(N ij +α) / (Σ j N ij +Kα), where α is the smoothing value and K is the number of degradation stage clusters, thus obtaining the lifetime stage transition probability array.
[0029] Step S203: Construct a lifetime prediction strategy network based on the state degradation trajectory vector set and the lifetime stage transition probability array, and train the lifetime prediction strategy network using the lifetime stage transition probability array to obtain a set of weight parameters for the lifetime prediction strategy network that has been trained and converged.
[0030] The server performs temporal reorganization on the set of state degradation trajectory vectors, forming a temporal sequence of degradation trajectories according to the order of degradation stages, and further generates an input state representation sequence. This input state representation sequence includes the current degradation stage center representation vector, historical load impact density, cumulative environmental stress intensity, electrical response offset trend, stage duration, physical property constraint damage index, and stage transition probability prior. The lifetime prediction strategy network constructed by the server includes a state encoding module, a strategy inference module, and a lifetime prediction output module. The state encoding module is used to extract the temporal dependencies of the degradation trajectory, the strategy inference module is used to learn the value of actions such as continuing monitoring and waiting, shortening the monitoring interval, entering early warning observation, and immediately executing replacement, and the lifetime prediction output module is used to output the remaining effective lifetime prediction probability distribution.
[0031] In one optional implementation, the server constructs the lifetime prediction process as a reinforcement learning environment. The states in the reinforcement learning environment include the current degradation stage center representation vector, stage duration, number of load impacts, cumulative environmental stress intensity, cumulative electrical response offset, data credibility weight, and the stage transition probability vector corresponding to the current degradation stage. The action space includes at least the actions of continuing monitoring and waiting, shortening the monitoring interval, entering early warning observation, and immediately executing replacement. The server calculates rewards based on the action results: a positive waiting reward is given when continuing monitoring is chosen in a low-risk stage; a positive safety reward is given when early warning observation or immediate replacement is chosen in a high-risk stage; a resource waste penalty is given for premature replacement when the remaining lifetime is long; and a failure risk penalty is given when waiting continues beyond the risk warning time point. The server combines the state, action, reward, and next state into training experience data tuples and iteratively updates the lifetime prediction strategy network based on experience replay until the prediction error and cumulative reward changes are both less than a preset threshold.
[0032] During training, the server determines the stage transition information corresponding to the input state representation sequence based on the lifetime stage transition probability array, and uses this information as the state transition dynamic mechanism of the reinforcement learning environment. The server defines an action space including the actions of continuing monitoring and waiting, and immediately executing replacement actions, and constructs a replacement decision reward function based on the remaining effective lifetime prediction distribution, the reliable remaining operating time interval, and the risk warning time point. For example, if a power relay is in the early degradation stage and subsequently operates stably, the policy network chooses to continue monitoring, and the server gives a positive reward; if a controller power device has entered the high-risk degradation stage, the policy network chooses to replace it within the planned maintenance window, and the server gives a higher positive reward; if premature replacement with a long remaining lifetime causes waste of spare parts, the server gives a negative reward; if waiting continues after the failure risk increases, leading to an increased failure risk, the server gives a stronger negative reward. The server updates the weight parameters of the lifetime prediction policy network according to the reward and introduces a degradation property constraint regularization term to make the prediction output conform to the physical degradation law corresponding to the operating limit of the component. When the training loss is stable, the reward converges, the prediction deviation is below the threshold, and the property constraints meet the requirements, the server obtains the set of weight parameters of the lifetime prediction policy network that has been trained and converged. For multiple electronic components, the server can also perform a weighted average of the gradients of the loss functions of their respective strategies to form a meta-policy gradient, and simultaneously update multiple lifetime prediction policy networks to achieve cross-component collaborative training.
[0033] Step S204: The lifetime prediction strategy network that has been trained and converged is invoked to process the state degradation trajectory vector set to generate the remaining effective lifetime prediction distribution corresponding to the electronic component identifier.
[0034] The server extracts a current state observation segment from the set of state degradation trajectory vectors. For example, it extracts the central characterization vector of the degradation stage, cumulative temperature rise path, voltage ripple change, number of load impacts, and environmental vibration distribution of a power management module over the past 30 days. The server inputs this current state observation segment into a converged lifetime prediction strategy network to obtain an initial remaining effective lifetime prediction distribution. This distribution expresses the probability of different remaining operating hours in probabilistic form; for example, the probability is highest in the 200-260 hour range, moderate in the 260-340 hour range, and low in the range above 340 hours. The server further extracts the stage transition probability vector corresponding to the final degradation stage of the current state observation segment based on the lifetime stage transition probability array, and uses this vector to correct the initial prediction distribution for state transitions. When the probability of transitioning from the current final degradation stage to a higher-risk stage is high, the server corrects the prediction distribution towards a shorter remaining lifetime range; when the current degradation stage remains stable and the probability of transitioning to a higher-risk stage is low, the server retains the probability weights for a longer remaining lifetime range. The server then performs time-series smoothing and normalization reshaping on the corrected remaining effective lifetime prediction distribution to eliminate abnormal fluctuations caused by single sensor noise, and finally outputs a remaining effective lifetime prediction distribution that is consistent with historical degradation trajectories, stage transition patterns and physical degradation constraints.
[0035] In one optional implementation, the server divides the remaining working time into multiple consecutive time intervals, and the lifetime prediction strategy network outputs the initial probability values for each time interval, forming an initial remaining effective lifetime prediction distribution. The server observes the degradation stage at the end of the current state segment and reads the stage transition probability vector from the lifetime stage transition probability array. When the probability of transitioning to a high-risk degradation stage is high, the probability weight of shorter remaining lifetime time intervals is increased; when the probability of maintaining the current stable degradation stage is high, the probability weight of longer remaining lifetime time intervals is increased. After correction, the server smooths and normalizes the probability values of adjacent time intervals, making the sum of all probability values equal to 1. The server uses the time point when the cumulative failure probability first reaches the preset risk warning probability as the risk warning time point, the center point of the time interval with the highest probability value as the expected failure time point, and, considering spare parts cost, downtime loss, maintenance resource conflicts, and failure risk cost, selects the candidate replacement time point with the lowest overall cost as the optimal replacement time.
[0036] Step S205: Generate a component replacement scheduling instruction set based on the remaining effective lifetime prediction distribution. The component replacement scheduling instruction set includes a sequence of electronic component identifiers to be replaced and corresponding time window information.
[0037] The server locates the failure risk time points based on the predicted distribution of remaining effective lifespan, obtaining risk warning time points and expected failure time points. For example, if the failure risk of a power device in an engine control unit increases rapidly in the next 180 hours, the server determines this moment as the risk warning time point; the peak of the predicted distribution corresponds to the next 240 hours, and the server determines this moment as the expected failure time point. The server also extracts the reliable remaining operating time interval and constructs a replacement time window selection strategy by combining vehicle transportation tasks, planned maintenance cycles, repair station workstations, spare parts inventory, maintenance personnel scheduling, and vehicle downtime costs. Within the replacement time window, the server generates multiple candidate replacement time points and evaluates the spare parts occupancy cost, downtime loss, failure risk, and maintenance resource conflicts for each candidate time point, selecting the optimal replacement time with the lowest overall cost. Finally, the server generates a set of component replacement scheduling instructions based on the electronic component identifier and the optimal replacement time. This set contains the sequence of electronic component identifiers to be replaced and the corresponding time window information, and is sent to the maintenance scheduling system. The maintenance scheduling system then arranges for vehicles to enter designated repair stations to complete the replacement. After the replacement is completed, the server continues to receive new operational monitoring data streams and updates the state degradation trajectory vector set, life stage transition probability array, remaining effective life prediction distribution, and replacement time window selection strategy, thereby forming a closed-loop intelligent operation and maintenance process for life prediction and replacement decision-making of heavy-duty vehicle electronic components.
[0038] In this embodiment of the invention, the degradation trajectory enhancement modeling process of the operation monitoring data stream, which generates a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, can be implemented through the following example.
[0039] The time-series load condition parameter set is subjected to load shock event detection processing. The load change time point and load change amplitude value are extracted from the time-series load condition parameter set to generate a load shock event sequence. Each load shock event in the load shock event sequence includes a load change time point label and a load change amplitude value label.
[0040] The time-series environmental stress parameter set is segmented into environmental stress fluctuation segments to obtain a set of environmental stress fluctuation segments. Each environmental stress fluctuation segment in the set of environmental stress fluctuation segments includes the start time of environmental stress change, the end time of environmental stress change, and the gradient sequence of environmental stress change.
[0041] The time-series electrical response parameter set is subjected to response anomaly offset identification processing. The starting point of the steady-state offset of the electrical response and the cumulative amount of the steady-state offset of the electrical response are extracted from the time-series electrical response parameter set to generate an electrical response offset feature sequence. The electrical response offset feature sequence includes the electrical response offset starting timestamp and the electrical response offset cumulative amount value.
[0042] The load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence are input into a pre-constructed degenerate trajectory encoder. The degenerate trajectory encoder includes multiple layers of temporal convolution operation layers and channel cross-attention operation layers. The multiple layers of temporal convolution operation layers are used to extract multi-dimensional temporal features from the load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence to obtain load impact temporal feature maps, environmental stress temporal feature maps, and electrical response temporal feature maps. The channel cross-attention operation layers are used to perform cross-channel interactive weight allocation processing on the load impact temporal feature maps, environmental stress temporal feature maps, and electrical response temporal feature maps to generate a multi-source stress coupling feature representation set.
[0043] The degradation stage clustering analysis is performed on the multi-source stress coupling feature representation set. The feature distribution distance between any two time-domain feature vectors in the degradation stage space is calculated, and the time-domain feature vectors that satisfy the feature distribution distance constraint are assigned to the same degradation stage cluster. The degradation stage division result corresponding to the electronic component identifier is obtained. The degradation stage division result includes several degradation stage clusters and the time-domain coverage of each degradation stage cluster.
[0044] The degradation stage division results are processed by degradation stage transition relationship mining, the stage transition frequency between different degradation stage clusters within adjacent time windows is counted, and the stage transition frequency is normalized to obtain the lifetime stage transition probability array corresponding to the electronic component identifier. The row index and column index of the lifetime stage transition probability array correspond to different degradation stage clusters. Each element in the lifetime stage transition probability array represents the transition probability value from the degradation stage cluster corresponding to the row index to the degradation stage cluster corresponding to the column index.
[0045] The degradation stage division results are processed by degradation trajectory vector mapping. The central feature vector of each degradation stage cluster is concatenated with the start and end time information of its time domain coverage to generate the state degradation trajectory vector set corresponding to the electronic component identifier. Each degradation trajectory element in the state degradation trajectory vector set includes the degradation stage central representation vector, the degradation stage start time label, and the degradation stage end time label.
[0046] In an exemplary embodiment of the invention, the server performs degradation trajectory enhancement modeling processing on the power switching device corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05". The server first reads the time-series load condition parameter set of the component during heavy-load start-up, full-load hill climbing, and long downhill braking processes of the vehicle, continuously detecting the traction current, power supply, number of switching actions, and transient current peak value. When the server detects that the traction current increases from 80A to 210A within 2 seconds and the power supply simultaneously spikes, the server extracts this moment as the load mutation time point and uses 130A as the load mutation amplitude value to generate a load impact event sequence. Each event in the sequence carries a load mutation time point label and a load mutation amplitude value label.
[0047] The server then processes the time-series environmental stress parameter set associated with the same identifier. The server segments the continuous temperature rise in the engine compartment from 45°C to 92°C, the change in controller housing vibration from low-frequency stability to high-frequency impact, and the rapid increase in humidity under rain and snow conditions into environmental stress fluctuation segments. For each segment, the server records the start and end times of environmental stress changes, as well as the gradient sequence formed by the changes in temperature, humidity, and vibration acceleration over time. Next, the server identifies response anomaly offsets in the time-series electrical response parameter set. After identifying a continuous increase in the on-state voltage drop of the power switching device over several days, an increase in output voltage ripple, and a gradual increase in leakage current, the server extracts the steady-state offset start point and the cumulative offset amount, generating an electrical response offset feature sequence containing the offset start timestamp and the cumulative offset value.
[0048] The server inputs the load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence into a pre-constructed degradation trajectory encoder. The multi-layer temporal convolutional operation layer in the degradation trajectory encoder extracts multi-dimensional temporal features such as short-term current impact, sustained high temperature, vibration accumulation, and electrical drift, forming load impact temporal feature maps, environmental stress temporal feature maps, and electrical response temporal feature maps. The channel cross-attention operation layer further calculates the correlation weights between different feature channels; for example, it assigns higher weights to the correlation between "continuously rising high temperature" and "increased conduction voltage drop," and higher weights to the correlation between "vibration impact" and "abnormal contact resistance," thereby generating a set of multi-source stress coupling feature representations.
[0049] The server performs degradation stage clustering analysis on the multi-source stress coupling feature representation set, calculates the feature distribution distance between any two time-domain feature vectors in the degradation stage space, and groups vectors whose distances satisfy the constraints into the same degradation stage cluster. Thus, the server classifies periods of stable operation and normal temperature rise into the first degradation stage cluster, periods of frequent temperature rise and increased ripple into the second degradation stage cluster, and periods of increased leakage current and decreased insulation impedance into the third degradation stage cluster, recording the time-domain coverage of each degradation stage cluster. The server further calculates the transition frequency between degradation stage clusters within adjacent time windows and normalizes it to generate a lifetime stage transition probability array, where each element represents the probability value of transitioning from a row-indexed degradation stage cluster to a column-indexed degradation stage cluster. Finally, the server concatenates the central feature vector of each degradation stage cluster with the start and end times of that stage to generate a state degradation trajectory vector set, ensuring that each degradation trajectory element simultaneously contains the degradation stage central representation vector, the degradation stage start time label, and the degradation stage end time label.
[0050] Please refer to the following: Figure 2 , Figure 2 This is a schematic diagram of the lifetime prediction strategy network construction and training process provided in an embodiment of the present invention. In this embodiment, the step of constructing a lifetime prediction strategy network based on the state degradation trajectory vector set and the lifetime stage transition probability array, and training the lifetime prediction strategy network using the lifetime stage transition probability array to obtain a weight parameter set of the converged lifetime prediction strategy network can be implemented through the following example.
[0051] The state degradation trajectory vector set is subjected to temporal sequence recombination processing. The degradation trajectory elements in the state degradation trajectory vector set are arranged in the order of the starting time labels of the degradation stages to generate a degradation trajectory temporal sequence. The degradation trajectory temporal sequence contains a degradation stage center representation vector chain organized in ascending time order.
[0052] The degenerate trajectory time series is processed by sliding window slicing. A time observation window of a preset length slides along the time axis of the degenerate trajectory time series. Each slide extracts a degenerate stage center representation vector subsequence within a time observation window as the actual input state representation of the policy network, resulting in an input state representation sequence. The input state representation sequence contains the actual input state representations corresponding to multiple time observation windows.
[0053] A lifespan prediction strategy network main structure is constructed, which includes a state encoder layer, a policy inference fully connected layer, and a lifespan prediction output layer. The state encoder layer maps the actual input state representation to a hidden state encoding vector. The policy inference fully connected layer performs nonlinear transformation on the hidden state encoding vector to obtain the remaining lifespan prediction mean parameter. The lifespan prediction output layer generates the remaining effective lifespan prediction probability distribution based on the remaining lifespan prediction mean parameter.
[0054] Read the future true transition path information corresponding to each actual input state representation in the input state representation sequence from the lifetime stage transition probability array. The future true transition path information includes a true stage transition sequence that continuously transitions from the current actual input state representation's degradation stage to subsequent degradation stages.
[0055] The real stage transition sequence and the remaining effective lifetime prediction probability distribution are input into a predefined reinforcement learning reward evaluation function. The reinforcement learning reward evaluation function calculates the immediate reward value based on the path overlap between the predicted transition path corresponding to the highest probability in the remaining effective lifetime prediction probability distribution and the real stage transition sequence, and calculates the cumulative reward expectation value based on the immediate reward value and the prediction confidence of the remaining effective lifetime prediction probability distribution.
[0056] A policy optimization loss function is constructed based on the expected value of the cumulative reward and the probability distribution of the remaining effective lifetime prediction. The policy optimization loss function includes a policy gradient update term and a value function estimation bias penalty term. The policy gradient update term is used to guide the lifetime prediction policy network to update the weight parameters of the state encoder layer, the policy inference fully connected layer and the lifetime prediction output layer in the direction of increasing the expected value of the cumulative reward.
[0057] The lifetime prediction strategy network is trained using a training iteration method based on a store-and-playback mechanism. The actual input state representation, the remaining effective lifetime prediction probability distribution, and the cumulative reward expectation value generated in each training session are combined into training experience data tuples and stored in the experience playback storage area. In subsequent training iterations, historical training experience data tuples are randomly selected from the experience playback storage area to train and update the lifetime prediction strategy network.
[0058] When the preset training convergence condition is met, the training process is terminated. The weight parameters of the state encoder layer, the weight parameters of the policy inference fully connected layer, and the weight parameters of the lifetime prediction output layer of the lifetime prediction policy network at this time are extracted. The extracted weight parameters of the state encoder layer, the weight parameters of the policy inference fully connected layer, and the weight parameters of the lifetime prediction output layer are combined into the weight parameter set of the lifetime prediction policy network that has achieved training convergence.
[0059] In an embodiment of the invention, for example, after obtaining the state degradation trajectory vector set and lifetime stage transition probability array corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05", the server constructs and trains a lifetime prediction strategy network. The server first performs temporal sequence reorganization processing on the state degradation trajectory vector set, arranging each degradation trajectory element from earliest to latest according to the degradation stage start time label, forming a degradation trajectory temporal sequence. This sequence sequentially includes degradation stage center representation vectors corresponding to the stable operation stage, the temperature rise ripple increase stage, the leakage current rise stage, and the switching delay deterioration stage, forming a degradation stage center representation vector chain organized in ascending time order.
[0060] The server then performs sliding window slicing on the degradation trajectory time series. The server slides along the time axis using a preset length observation window, such as the most recent 7 days, 14 days, or 30 days. Each time, it extracts a sequence of consecutive degradation stage center representations within the window, which serves as the actual input state representation for the lifetime prediction strategy network. The server thus obtains multiple input state representations, each reflecting the continuous evolution of the component's degradation state over a certain period. For example, one window might represent the gradual increase in voltage ripple after frequent high temperatures, while another window might represent the decrease in insulation resistance after increased leakage current.
[0061] The server constructs the main structure of the lifetime prediction policy network, which includes a state encoder layer, a policy inference fully connected layer, and a lifetime prediction output layer. The state encoder layer maps each actual input state representation to a hidden state encoding vector, representing the temporal relationship between the current degradation stage, historical stress accumulation, and electrical response offset. The policy inference fully connected layer performs a nonlinear transformation on the hidden state encoding vector to obtain the remaining lifetime prediction mean parameter. The lifetime prediction output layer generates the remaining effective lifetime prediction probability distribution based on this mean parameter, enabling the server to obtain the probability corresponding to different remaining operating time intervals.
[0062] The server reads the future true transition path information corresponding to each actual input state representation from the lifetime stage transition probability array. For example, if the current window ends in the "temperature rise ripple increase stage", the server reads the true stage transition sequence that subsequently transitions to the "leakage current rise stage" and the "switching delay deterioration stage". The server inputs this true stage transition sequence and the remaining effective lifetime prediction probability distribution generated by the lifetime prediction output layer into the reinforcement learning reward evaluation function. This function calculates the path overlap between the predicted transition path corresponding to the highest probability in the prediction distribution and the true stage transition sequence. The higher the path overlap, the larger the immediate reward value; the more stable the prediction confidence, the higher the cumulative reward expectation value. Thus, the server incorporates lifetime prediction accuracy and stage transition consistency into the training objective.
[0063] The server constructs a policy optimization loss function based on the expected cumulative reward and the predicted remaining useful life probability distribution. This loss function includes a policy gradient update term and a value function estimation bias penalty term. The policy gradient update term guides the state encoder layer, the fully connected policy inference layer, and the useful life prediction output layer to update the weight parameters in the direction of increasing expected cumulative reward; the value function estimation bias penalty term suppresses the deviation between the predicted probability distribution and the actual degradation transition trend. The server uses a store-and-play mechanism for training iterations, combining the actual input state representation, the predicted remaining useful life probability distribution, and the expected cumulative reward generated in each training iteration into a training experience data tuple, which is stored in the experience replay storage area. In subsequent training, the server randomly selects historical experience data tuples to participate in parameter updates, avoiding the model only remembering the most recent operating conditions and improving the generalization ability to various scenarios such as heavy load, vibration, high temperature, and electrical drift.
[0064] When the training loss stabilizes, the expected cumulative reward no longer increases significantly, and the overlap between the predicted path and the actual stage transition sequence reaches a preset requirement, the server terminates the training process. The server extracts the weight parameters of the state encoder layer, the weight parameters of the policy inference fully connected layer, and the weight parameters of the lifetime prediction output layer at this point, and combines these three types of parameters into a set of weight parameters for the converged lifetime prediction policy network, which is then used for subsequent prediction of the remaining effective lifetime.
[0065] In this embodiment of the invention, the process of calling the converged lifetime prediction strategy network to process the state degradation trajectory vector set and generate the remaining effective lifetime prediction distribution corresponding to the electronic component identifier can be implemented through the following example.
[0066] The state degradation trajectory vector set is processed by current observation segment extraction. A subset of state degradation trajectory vectors within the most recent time window is extracted from the state degradation trajectory vector set as the current state observation segment. The current state observation segment includes the degradation stage center representation vector sequence within the coverage of the current time window and the corresponding degradation stage start and end time labels.
[0067] The current state observation segment is input into the state encoder layer of the lifetime prediction policy network that has been trained and converged. The weight parameters of the trained and converged state encoder layer are used to encode and map the degeneration stage center representation vector sequence in the current state observation segment to obtain the current hidden state encoding vector.
[0068] The current hidden state encoding vector is input into the policy inference fully connected layer of the lifetime prediction policy network that has been trained and converged. The weight parameters of the trained and converged policy inference fully connected layer are used to perform nonlinear transformation on the current hidden state encoding vector to generate the current remaining lifetime prediction mean parameter and the current remaining lifetime prediction variance parameter.
[0069] The mean parameter and variance parameter of the current remaining lifetime prediction are input into the lifetime prediction output layer of the lifetime prediction strategy network that has been trained and converged. The weight parameters of the lifetime prediction output layer that has been trained and converged are used to construct the probability distribution of the mean parameter and variance parameter of the current remaining lifetime prediction to generate the initial remaining effective lifetime prediction distribution corresponding to the electronic component identifier.
[0070] Extract the stage transition probability vector corresponding to the terminal degradation stage of the current state observation segment from the lifetime stage transition probability array. The stage transition probability vector contains the transition probability values from the terminal degradation stage of the current state observation segment to each possible subsequent degradation stage.
[0071] The initial remaining effective lifetime prediction distribution is state transition corrected using the stage transition probability vector. The stage transition probability vector is then used as the distribution correction weight and the initial remaining effective lifetime prediction distribution is weighted and fused point by point to obtain the state transition corrected remaining effective lifetime prediction distribution.
[0072] The remaining effective lifetime prediction distribution after state transition correction is subjected to time-series smoothing. The probability value sequence of the remaining effective lifetime prediction distribution after state transition correction is subjected to neighborhood weighted averaging along the time prediction dimension to obtain the time-series smoothed remaining effective lifetime prediction distribution.
[0073] The time-smoothed remaining effective lifetime prediction distribution is subjected to distribution normalization and renormalization. The probability values corresponding to all prediction time points in the time-smoothed remaining effective lifetime prediction distribution are summed to obtain the total probability value. The probability value corresponding to each prediction time point is divided by the total probability value to obtain the remaining effective lifetime prediction distribution corresponding to the electronic component identifier.
[0074] In an embodiment of the invention, for example, the server invokes a trained and converged lifetime prediction strategy network to predict the remaining effective lifetime of the state degradation trajectory vector set corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05". The server first performs a current observation segment extraction process, extracting a subset of state degradation trajectory vectors within the most recent time window from the state degradation trajectory vector set. For example, it extracts the degradation trajectory elements corresponding to the power switch device within the last 30 days to form the current state observation segment. This current state observation segment contains a sequence of central characterization vectors representing the degradation stages within the current time window, while retaining the start and end time labels of each degradation stage, enabling the server to identify that the component has recently experienced a continuous degradation process, such as frequent temperature rises, increased ripple, and increased on-state voltage drop.
[0075] The server inputs the current state observation segment into the state encoder layer of the converged lifetime prediction strategy network. The state encoder layer uses the converged weight parameters to encode and map the sequence of degradation stage center representation vectors in the current state observation segment, compressing multiple recent degradation states into a current hidden state encoding vector. This current hidden state encoding vector simultaneously expresses the current degradation stage of the component, the rate of degradation, the intensity of accumulated stress, and the electrical response shift trend. For example, the server's encoding result shows that the power switching device has gradually moved from the "temperature rise ripple increase stage" to the "leakage current increase stage."
[0076] The server then inputs the current hidden state encoding vector into the fully connected policy inference layer that has converged during training. The policy inference layer uses its weight parameters to perform a nonlinear transformation on the current hidden state encoding vector, generating the current remaining lifetime prediction mean parameter and the current remaining lifetime prediction variance parameter. The current remaining lifetime prediction mean parameter represents the central estimate of the remaining effective operating time of the component, while the current remaining lifetime prediction variance parameter represents the uncertainty of the prediction result. For example, if the server obtains a remaining lifetime mean of approximately 240 hours and a large fluctuation range in the variance, it indicates that the component still has some uncertainty in its failure time under alternating high temperature and heavy load conditions.
[0077] The server inputs the mean and variance parameters of the current remaining lifetime prediction into the lifetime prediction output layer. The lifetime prediction output layer uses the converged weight parameters from training to construct a probability distribution, generating an initial remaining effective lifetime prediction distribution. This distribution provides probability values according to the prediction time point or prediction time interval; for example, the interval between 180 and 220 hours has a higher probability of failure, the interval between 220 and 280 hours has the highest probability of remaining lifetime, and the probability gradually decreases above 280 hours.
[0078] The server further extracts the stage transition probability vector corresponding to the degradation stage at the end of the current state observation segment from the lifetime stage transition probability array. This vector contains the probability values of transitioning from the current end degradation stage to each possible subsequent degradation stage. For example, if the current end stage is the "temperature rise ripple increase stage", the transition probability vector shows that the probability of continuing to maintain the current stage is 0.35, the probability of transitioning to the "leakage current increase stage" is 0.50, and the probability of transitioning to the "switching delay deterioration stage" is 0.15. The server uses this stage transition probability vector as a distribution correction weight and performs point-by-point weighted fusion processing with the initial remaining effective lifetime prediction distribution to ensure that the prediction results fully reflect the actual degradation stage transition trend. After state transition correction, the probability weight of shorter remaining lifetime intervals is appropriately increased, and the probability weight of longer remaining lifetime intervals is correspondingly decreased.
[0079] The server performs temporal smoothing on the remaining effective lifetime prediction distribution after state transition correction. It applies a neighborhood-weighted average to the probability values of adjacent prediction time points along the time prediction dimension to eliminate probability spikes caused by short-term sensor fluctuations or single abnormal shocks. Finally, the server normalizes and reshapes the temporally smoothed distribution, summing the probability values corresponding to all prediction time points to obtain a total probability value. Then, it divides the probability value of each prediction time point by this total probability value, ensuring that the sum of all probability values equals 1. Thus, the server obtains the final remaining effective lifetime prediction distribution corresponding to the electronic component identifier. This distribution reflects both the learning results of the lifetime prediction strategy network and incorporates the degradation transition patterns in the lifetime stage transition probability array, and can be directly used for subsequent risk warning and replacement scheduling.
[0080] In this embodiment of the invention, the generation of a set of component replacement scheduling instructions based on the remaining effective lifetime prediction distribution can be implemented through the following example.
[0081] The remaining effective lifetime prediction distribution is processed to locate the failure risk time point. The probability density value is detected point by point along the time dimension of the remaining effective lifetime prediction distribution. The time point when the probability density value first exceeds the preset probability density limit value is extracted as the risk warning time point, and the time point when the probability density value reaches the peak of the distribution curve is extracted as the expected failure time point.
[0082] The remaining effective lifetime prediction distribution is processed by reliable remaining time interval extraction. The time interval during which the cumulative probability value grows from the preset reliable working probability lower bound to the preset failure probability upper bound is extracted from the remaining effective lifetime prediction distribution as the reliable remaining working time interval. The reliable remaining working time interval includes the reliable working start time point and the reliable working end time point.
[0083] A replacement time window selection strategy is constructed based on the expected failure time point and the reliable remaining working time interval. The reliable working start time point is taken as the earliest allowed replacement time of the replacement time window, and the risk warning time point is taken as the latest mandatory replacement time of the replacement time window. The replacement time window range corresponding to the electronic component identifier is generated.
[0084] The replacement time window range is discretized at the time granularity. According to the preset scheduling time step, the replacement time window range is divided into multiple discrete candidate replacement time grids to generate a candidate replacement time grid set. The candidate replacement time grid set contains multiple candidate replacement time points arranged in chronological order.
[0085] For each candidate replacement time grid in the candidate replacement time grid set, a replacement cost evaluation process is performed. Based on the time interval between the candidate replacement time grid and the expected failure time point, the vehicle downtime loss parameter corresponding to the time period of the candidate replacement time grid, and the spare parts preparation cost parameter required for the candidate replacement time grid, a comprehensive replacement cost evaluation value for each candidate replacement time grid is generated.
[0086] The comprehensive replacement cost evaluation values of all candidate replacement time grid points in the candidate replacement time grid point set are compared and sorted, and the candidate replacement time grid point with the smallest comprehensive replacement cost evaluation value is selected as the optimal replacement time corresponding to the electronic component identifier.
[0087] The electronic component identifier and the optimal replacement time are combined into a single component replacement scheduling instruction. The single component replacement scheduling instructions corresponding to all electronic components to be replaced on the same vehicle are merged and arranged according to the order of the optimal replacement time to generate the component replacement scheduling instruction set. The component replacement scheduling instruction set contains the sequence of electronic component identifiers to be replaced and the corresponding time window information.
[0088] In this embodiment of the invention, for example, the server generates a set of component replacement scheduling instructions based on the remaining effective lifetime prediction distribution corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05". The server first checks the probability density value point-by-point along the time dimension of the remaining effective lifetime prediction distribution. For example, if the failure probability density of the power switch device is low before the next 160 hours, and exceeds a preset probability density threshold for the first time at the next 180 hours, the server determines the time point corresponding to the next 180 hours as the risk warning time point; simultaneously, if the prediction distribution reaches its peak at the next 240 hours, the server determines the time point corresponding to the next 240 hours as the expected failure time point. Thus, the server clarifies the time relationship between the significant increase in risk and the most likely failure of the component.
[0089] The server then performs reliable remaining time interval extraction on the remaining effective lifetime prediction distribution. The server calculates the cumulative probability value from the distribution curve and extracts the time interval during which the cumulative probability value increases from a preset lower bound of reliable operation probability to a preset upper bound of failure probability, as the reliable remaining operation time interval. For example, if the server identifies that the component is still in a high reliable operation probability range between the current time and the next 170 hours, and that the failure risk enters a rapidly increasing phase after 170 hours, the server uses the current time as the reliable operation start point and the corresponding time in the next 170 hours as the reliable operation end point. This interval is used to constrain replacement scheduling to avoid resource waste due to premature replacements and to prevent entering a high-risk operating state too late.
[0090] The server constructs a replacement time window selection strategy based on the expected failure time and the remaining reliable operating time interval. The server uses the reliable operating start time as the earliest permissible replacement time within the replacement time window and the risk warning time as the latest mandatory replacement time, generating the replacement time window range corresponding to the electronic component identifier. For example, the server determines the 24th hour after the current time as the earliest permissible replacement time and the next 180 hours as the latest mandatory replacement time, ensuring that maintenance is scheduled within the safe operating range of the component and avoiding continued operation close to the expected failure time.
[0091] The server discretizes the replacement time window range at the time granularity, dividing it into multiple discrete candidate replacement time grids according to a preset scheduling time step. For example, with a scheduling time step of 6 hours, the server divides the replacement window into candidate replacement time points such as the 24th hour, 30th hour, 36th hour, and up to the 180th hour, generating a set of candidate replacement time grids arranged chronologically. The server performs a replacement cost assessment for each candidate replacement time grid, incorporating the time interval between the candidate replacement time point and the expected failure time point, the vehicle downtime loss parameters corresponding to that time period, the occupancy status of repair station workstations, parts preparation cost parameters, and spare parts inventory turnover pressure into a comprehensive calculation. For example, while replacement at the 48th hour carries a low risk, it is far from the expected failure time point, resulting in high early spare parts occupancy and high residual value loss of components; replacement at the 168th hour is closer to the risk warning time point, with higher spare parts utilization, but higher costs due to vehicle transportation task conflicts and failure risk; the 120th hour falls within the vehicle's planned maintenance window, with available workstations and spare parts already in stock, resulting in the lowest overall replacement cost.
[0092] The server compares and ranks the comprehensive replacement cost evaluation values of all candidate replacement time grids, selecting the candidate replacement time grid with the lowest cost as the optimal replacement time corresponding to the electronic component identifier. Subsequently, the server combines the electronic component identifier and the optimal replacement time into a single component replacement scheduling instruction. For multiple components to be replaced simultaneously on the same vehicle, such as engine control unit power devices, brake controller capacitor modules, and sensor interface board connectors, the server merges and arranges them according to the order of their respective optimal replacement times. Instructions with similar time windows, adjacent repair locations, and the same required workstation are merged into the same repair batch, generating a component replacement scheduling instruction set. This set contains the sequence of electronic component identifiers to be replaced and their corresponding time window information, and is sent to the maintenance scheduling system, ensuring that vehicle replacement is completed during planned downtime, reducing the risk of sudden failures and unplanned downtime.
[0093] In this embodiment of the invention, after acquiring the operation monitoring data stream of heavy-duty vehicle electronic components, the method further includes:
[0094] The operation monitoring data stream is subjected to multi-source heterogeneous sampling frequency alignment processing to construct a time-aligned operation monitoring data stream containing a unified time base index. The time-aligned operation monitoring data stream retains the time synchronization mapping relationship between the electronic component identifier, the aligned time-series load condition parameter set, the aligned time-series environmental stress parameter set, and the aligned time-series electrical response parameter set.
[0095] Based on the time-aligned operation monitoring data stream, a state space representation structure for reinforcement learning environment is constructed. The time-series load condition parameter set is mapped to load state dimension components, the time-series environmental stress parameter set is mapped to environmental state dimension components, and the time-series electrical response parameter set is mapped to electrical state dimension components. The load state dimension components, the environmental state dimension components, and the electrical state dimension components are concatenated to generate a multi-dimensional state space representation vector sequence.
[0096] The lifetime stage transition probability array is used to construct a state transition dynamics mechanism for reinforcement learning environment. Each row of probability vector in the lifetime stage transition probability array is used as the action transition probability distribution of the corresponding degradation stage state. The corresponding action transition probability distribution is read from the lifetime stage transition probability array according to the degradation stage where the current multidimensional state space representation vector is located, and the degradation stage to which the multidimensional state space representation vector belongs at the next moment is determined according to the action transition probability distribution.
[0097] Define a reinforcement learning action space for component replacement decision-making. The reinforcement learning action space includes a continue monitoring and waiting action and an immediate replacement action. The continue monitoring and waiting action triggers the system to continue to acquire the multi-dimensional state space representation vector of the next time step and delay the replacement decision. The immediate replacement action triggers the system to terminate the current electronic component lifetime prediction process and generate a replacement execution signal.
[0098] A replacement decision reward function is constructed based on the multidimensional state space representation vector corresponding to the execution of the immediate replacement action and the remaining effective lifetime prediction distribution. If the execution time of the immediate replacement action falls within the reliable remaining working time interval, the replacement decision reward function outputs a positive replacement reward value. If the execution time of the immediate replacement action exceeds the risk warning time point, the replacement decision reward function outputs a negative replacement penalty value.
[0099] In the training process of the lifespan prediction strategy network, a replacement decision collaborative training mechanism is introduced. The remaining effective lifespan prediction distribution output by the lifespan prediction strategy network is used as the basis for action selection in the reinforcement learning action space. When the cumulative probability value of the remaining effective lifespan prediction distribution exceeding the risk warning time point exceeds the preset emergency replacement initiation threshold, the reinforcement learning action space selects the immediate replacement action with a preset inertial probability bias.
[0100] The monitoring time interval corresponding to the continued monitoring waiting action is dynamically and adaptively adjusted. The length of the monitoring time interval is adjusted according to the dispersion parameter of the remaining effective lifespan prediction distribution at the current time. When the dispersion parameter indicates that the prediction uncertainty is increased, the monitoring time interval is shortened; when the dispersion parameter indicates that the prediction uncertainty is reduced, the monitoring time interval is extended.
[0101] The newly acquired running monitoring data stream within the monitoring time interval after the dynamic adaptive adjustment process is appended to the time-aligned running monitoring data stream, and the incremental update process of the multidimensional state space representation vector sequence is triggered to generate the incrementally updated multidimensional state space representation vector sequence.
[0102] The incrementally updated multidimensional state space representation vector sequence is re-input into the converged lifetime prediction policy network for iterative update of the remaining effective lifetime prediction distribution to obtain the updated remaining effective lifetime prediction distribution. The replacement time window selection policy is then re-executed based on the updated remaining effective lifetime prediction distribution.
[0103] In this embodiment of the invention, for example, after acquiring the operational monitoring data stream of heavy-duty vehicle electronic components, the server first performs multi-source heterogeneous sampling frequency alignment processing on data from different sources and with different sampling frequencies. Using a unified time base index as the main thread, the server maps millisecond-level collected output voltage ripple, leakage current, and switching delay data, second-level collected traction load, engine torque, and power supply current data, and minute-level collected ambient temperature, humidity, and vibration intensity data onto the same time axis. For the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05", the server retains the synchronization relationship between the identifier, the aligned time-series load condition parameter set, the aligned time-series environmental stress parameter set, and the aligned time-series electrical response parameter set at each alignment time point, enabling simultaneous correlation analysis of heavy-load impacts, high-temperature increases, and abnormal conduction voltage drops occurring at a certain moment.
[0104] The server constructs a state-space representation structure for the reinforcement learning environment based on time-aligned monitoring data streams. The server maps traction current peak values, load surge amplitude, and number of start-stop impacts to load state-dimensional components; controller cavity temperature, humidity, vibration acceleration, and heat dissipation temperature difference to environmental state-dimensional components; and output voltage ripple, leakage current, on-state voltage drop, insulation impedance, and switching delay to electrical state-dimensional components. The server concatenates these three types of components to generate a multi-dimensional state-space representation vector sequence, ensuring that each vector represents the comprehensive operating state of the component at the corresponding moment. For example, a vector might simultaneously express a degradation state characterized by "increased full-load ramp current, engine compartment temperature approaching 90°C, and continuously increasing on-state voltage drop."
[0105] The server utilizes a lifetime stage transition probability array to construct the state transition dynamics mechanism of the reinforcement learning environment. Each row of probability vectors in the array serves as the action transition probability distribution for the corresponding degradation stage, and the server reads the corresponding row based on the degradation stage to which the current multidimensional state space representation vector belongs. For example, if the current component is in the "temperature rise ripple increase stage," the server reads the probabilities of maintaining the current stage, transitioning to the "leakage current increase stage," and further transitioning to the "switching delay deterioration stage," and determines the degradation stage to which the multidimensional state space representation vector belongs in the next moment based on this.
[0106] The server further defines a reinforcement learning action space for component replacement decisions, which includes a "continue monitoring and waiting" action and an "immediately execute replacement" action. When the server executes the "continue monitoring and waiting" action, the system continues to collect the multi-dimensional state space representation vector for the next time step and postpones the replacement decision. When the server executes the "immediate replacement" action, the system terminates the current electronic component lifespan prediction process and generates a replacement execution signal to the maintenance scheduling system. The server constructs a replacement decision reward function based on the multi-dimensional state space representation vector corresponding to the "immediately execute replacement" action and the remaining effective lifespan prediction distribution. When the immediate replacement time falls within the reliable remaining working time interval, the server outputs a positive replacement reward; when the immediate replacement time has exceeded the risk warning time point, the server outputs a negative replacement penalty to prevent the vehicle from continuing to operate in a high-risk state.
[0107] During the training of the lifetime prediction strategy network, the server introduces a replacement decision collaborative training mechanism, using the remaining effective lifetime prediction distribution output by the lifetime prediction strategy network as the basis for action selection. When the cumulative probability value of exceeding the risk warning time point in the distribution exceeds the emergency replacement initiation threshold, the server causes the action space to select to immediately execute the replacement action with a preset inertial probability bias. For the continued monitoring and waiting action, the server also dynamically adjusts the monitoring interval according to the dispersion of the remaining effective lifetime prediction distribution; when the prediction uncertainty increases, the server shortens the monitoring interval from 6 hours to 1 hour, and when the prediction uncertainty decreases, the server extends the monitoring interval to 12 hours. The server appends the newly collected data within the adjusted monitoring interval to the time-aligned running monitoring data stream, triggering an incremental update of the multidimensional state space representation vector sequence, and re-inputs the updated sequence into the training converged lifetime prediction strategy network to obtain the updated remaining effective lifetime prediction distribution. Then, based on this distribution, the replacement time window selection strategy is re-executed, forming a continuously adaptive lifetime prediction and replacement decision closed loop.
[0108] In this embodiment of the invention, after performing degradation trajectory enhancement modeling on the operation monitoring data stream and generating the state degradation trajectory vector set and lifetime stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, the method further includes:
[0109] The state degradation trajectory vector set and the lifetime stage transition probability array are subjected to cross-component degradation mode migration processing. The probability subarrays with shared degradation structure characteristics in the lifetime stage transition probability arrays corresponding to different electronic component identifiers are extracted, and the degradation stage clusters corresponding to the probability subarrays with shared degradation structure characteristics are marked as general degradation mode stage clusters.
[0110] The common degradation feature extraction process is performed on the general degradation mode stage cluster. All degradation stage center representation vectors belonging to the general degradation mode stage cluster from different electronic component identifiers are aggregated. The common factor decomposition process is performed on the aggregated degradation stage center representation vectors to obtain a general degradation mode representation vector basis set. The general degradation mode representation vector basis set contains a set of basis vectors for linearly representing the general degradation modes of different components.
[0111] The general degradation mode characterization vector basis set is injected into the channel cross-attention operation layer of the degradation trajectory encoder. A general degradation mode matching branch is added to the channel cross-attention operation layer of the degradation trajectory encoder. The general degradation mode matching branch calculates the mode matching similarity between the input load impact time series feature map, the environmental stress time series feature map and the electrical response time series feature map and each general degradation mode basis vector.
[0112] The pattern matching similarity output by the general degradation pattern matching branch is used to perform pattern enhancement processing on the multi-source stress coupling feature representation set. The pattern matching similarity is used as the pattern enhancement weight coefficient and multiplied with the corresponding time-domain feature vector in the multi-source stress coupling feature representation set to obtain the pattern-enhanced multi-source stress coupling feature representation set.
[0113] The degradation stage clustering analysis is re-executed on the enhanced multi-source stress coupling feature representation set to obtain the degradation stage division result based on the general degradation mode enhancement. The state degradation trajectory vector set and lifetime stage transition probability array corresponding to the electronic component identifier are updated based on the degradation stage division result based on the general degradation mode enhancement.
[0114] When training the lifetime prediction strategy network, meta-policy gradient sharing is performed on the lifetime prediction strategy networks corresponding to different electronic component identifiers. The gradients of the policy optimization loss functions corresponding to multiple electronic component identifiers are weighted and averaged to obtain the meta-policy gradient. The meta-policy gradient is used to update the weight parameters of the lifetime prediction strategy networks corresponding to each electronic component identifier simultaneously, thus obtaining a lifetime prediction strategy network trained across components.
[0115] A degradation trajectory prediction mechanism for cold start of novel electronic components is introduced. When a novel electronic component identifier not included in the training data appears, the initial stage operation monitoring data stream of the novel electronic component identifier is obtained, the initial state degradation trajectory vector set of the novel electronic component identifier is generated, and the initial state degradation trajectory vector set is processed by the general degradation mode characterization vector basis set to obtain the preliminary lifetime stage transition probability array estimate of the novel electronic component identifier.
[0116] The preliminary lifetime stage transition probability array estimate of the novel electronic component identifier is rapidly adapted to the lifetime prediction strategy network trained across components. The preliminary lifetime stage transition probability array estimate is used as the environmental prior knowledge of the lifetime prediction strategy network. The deviation between the remaining lifetime prediction mean parameter output of the lifetime prediction strategy network trained across components and the actual remaining lifetime observation value of the novel electronic component identifier is minimized and fine-tuned to obtain the adapted remaining lifetime prediction distribution of the novel component.
[0117] In an embodiment of the invention, for example, after generating state degradation trajectory vector sets and life stage transition probability arrays for multiple heavy-duty automotive electronic components, the server further performs cross-component degradation mode migration processing. The server simultaneously reads the life stage transition probability arrays corresponding to the engine control unit power switching devices, power management module capacitors, brake controller driver chips, and sensor interface board connectors, and compares the transition structures between degradation stages in each array. If the server identifies that multiple components exhibit similar probability subarrays that "transfer from the high-temperature stress accumulation stage to the electrical response shift stage, and then to the failure risk increase stage," it marks the degradation stage clusters corresponding to these probability subarrays with shared degradation structure characteristics as general degradation mode stage clusters.
[0118] The server then performs cross-component common degradation feature extraction on the general degradation mode stage clusters. The server aggregates degradation stage central representation vectors from different electronic component identifiers, such as aggregating the characteristics of increased on-state voltage drop of power switching devices, increased equivalent series resistance of capacitors, fluctuating contact resistance of connectors, and increased switching delay of driver chips. The server performs common factor decomposition on the aggregated central representation vectors to extract a set of general degradation mode representation vector bases that can jointly represent degradation laws such as "accumulated thermal stress," "increased electrical drift," "deteriorated insulation performance," and "deteriorated response delay." Each basis vector in this set is used to linearly represent the common degradation modes that can be transferred between different components.
[0119] The server injects the set of general degradation mode representation vectors into the channel cross-attention operation layer of the degradation trajectory encoder, and adds a general degradation mode matching branch within it. This branch calculates the mode matching similarity between the load impact time series feature map, environmental stress time series feature map, and electrical response time series feature map and each general degradation mode basis vector. When processing a certain power switching device, the server identifies that its feature of "continuous high temperature and increased on-state voltage drop" is highly similar to the general thermal degradation mode. Therefore, it uses this similarity as a mode enhancement weight coefficient and multiplies it with the corresponding time-domain feature vector in the multi-source stress coupling feature representation set to obtain the enhanced multi-source stress coupling feature representation set. Based on the enhanced features, the server re-performs degradation stage clustering analysis, making the previously unclear early degradation stage and mid-stage degradation stage more accurately distinguishable, and updates the state degradation trajectory vector set and lifetime stage transition probability array corresponding to the electronic component identifier accordingly.
[0120] During the lifetime prediction strategy network training phase, the server performs meta-policy gradient sharing processing on the strategy networks corresponding to different electronic component identifiers. The server calculates the gradients of the policy optimization loss function for power switching devices, capacitor modules, driver chips, and connectors respectively, and performs a weighted average according to the number of samples, the completeness of degradation stage coverage, and the prediction error weight to obtain the meta-policy gradient. The server uses this meta-policy gradient to synchronously update the weight parameters of each lifetime prediction strategy network, enabling different components to share degradation prediction experience and forming a cross-component collaboratively trained lifetime prediction strategy network.
[0121] When a new high-voltage drive module not included in the training data is added to the fleet, the server acquires the initial phase operation monitoring data stream corresponding to the identifier of this new electronic component and generates an initial state degradation trajectory vector set. The server uses a general degradation pattern representation vector basis set to perform pattern matching on this initial trajectory to obtain a preliminary lifetime stage transition probability array estimate. Subsequently, the server uses this estimated array as prior environmental knowledge and performs rapid adaptation processing with a lifetime prediction strategy network trained across components. Based on the actual remaining lifetime observations of the new component, the server fine-tunes the mean parameter of the remaining lifetime prediction output by the network to minimize the deviation, ultimately obtaining the adapted remaining lifetime prediction distribution of the new component.
[0122] Please refer to the following: Figure 3 , Figure 3 The schematic diagram of the degradation trajectory enhancement modeling process provided for the implementation of the present invention shows that, in the embodiment of the present invention, the degradation trajectory enhancement modeling process of the operation monitoring data stream is performed to generate a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set and the time-series electrical response parameter set. This can be implemented through the following example.
[0123] A degradation property constraint layer is constructed to reflect the internal micro-degradation mechanism of electronic components. The degradation property constraint layer includes material characteristic parameters, process characteristic parameters and operating limit characteristic parameters of electronic components. The material characteristic parameters are pre-stored in a one-to-one correspondence with the electronic component identifier. The process characteristic parameters are derived from the manufacturing process attribute records of electronic components. The operating limit characteristic parameters are derived from the technical specification documents of electronic components.
[0124] The material property parameters, process property parameters, and working limit property parameters of the degraded property constraint layer are injected into the degraded trajectory encoder. A degraded property constraint gating unit is added to the multi-layer temporal convolution operation layer of the degraded trajectory encoder. The degraded property constraint gating unit performs physical feasibility discrimination processing on the temporal features output after temporal convolution operation to suppress the temporal feature response amplitude that exceeds the allowable range of the working limit property parameters.
[0125] The stress damage accumulation path tracing process is performed on the multi-source stress coupling feature representation set generated after injecting the degraded physical property constraint layer. The evolution trajectory of the feature vector in the multi-source stress coupling feature representation set is calculated point by point along the time dimension. The damage accumulation direction vector after being restricted by the degraded physical property constraint layer is then calculated to generate a stress damage accumulation path curve. The stress damage accumulation path curve represents the evolution process of joint damage accumulation of electronic components under physical property constraints.
[0126] Based on the stress damage accumulation path curve, the cluster centers of the degradation stage clustering analysis are re-divided under physical property constraints. The time corresponding to the curvature change point on the stress damage accumulation path curve is taken as the degradation stage boundary point. The coverage time range of the degradation stage cluster is determined according to the path length of the stress damage accumulation path curve segment between the degradation stage boundary points, and the physical property constraint degradation stage division result is obtained.
[0127] The lifetime stage transition probability array corresponding to the electronic component identifier is updated according to the degradation stage division result of the physical property constraint. The number of degradation stage clusters and the time domain coverage in the lifetime stage transition probability array are consistent with the degradation stage division result of the physical property constraint. The transition probability value of the lifetime stage transition probability array is normalized and allocated according to the path characteristic distance of the stress damage accumulation path curve segment between each degradation stage boundary point.
[0128] The degradation trajectory vector mapping process is performed on the degradation stage division results of the physical property constraint. The cluster center feature vector of each degradation stage cluster in the degradation stage division results, the mean of the projection coordinates of all feature vectors in the degradation stage cluster along the stress damage accumulation path curve, and the time information of the degradation stage boundary point are fused and spliced to generate a state degradation trajectory vector set injected with physical property constraint information.
[0129] During the training of the lifetime prediction strategy network using the lifetime stage transition probability array, a degradation property constraint regularization term is introduced. This term performs a working limit characteristic parameter boundary check on the remaining effective lifetime prediction mean parameter output by the lifetime prediction strategy network. When the remaining effective lifetime prediction mean parameter exceeds the theoretical lifetime upper limit determined based on the working limit characteristic parameter, the loss value of the strategy optimization loss function is increased, thus constraining the prediction output of the lifetime prediction strategy network to conform to the physical degradation law.
[0130] In an embodiment of the invention, for example, when the server performs degradation trajectory enhancement modeling on the operation monitoring data stream corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05", it first constructs a degradation property constraint layer. The server reads the material characteristic parameters of the power switch device from the component property parameter library, including the thermal conductivity of the chip material, the coefficient of thermal expansion of the packaging material, the fatigue threshold of the solder joint, and the withstand voltage characteristics of the insulating medium; it reads process characteristic parameters such as bonding wire process, solder temperature profile, packaging batch, and heat dissipation substrate process from the manufacturing process attribute record; and it reads operating limit characteristic parameters such as maximum junction temperature, maximum on-state current, maximum leakage current, allowable voltage ripple, and switching delay limit from the technical specification document. The server organizes these parameters into a degradation property constraint layer that reflects the internal microscopic degradation mechanism, so that subsequent degradation judgment not only depends on data changes but is also constrained by the physical boundaries of the component.
[0131] The server injects the degradation property constraint layer into the degradation trajectory encoder and adds a degradation property constraint gating unit in the multi-layer temporal convolution operation layer. After the temporal convolution operation layer extracts timing features from load impact, ambient temperature rise, vibration stress, and electrical response offset, the degradation property constraint gating unit immediately performs a physical feasibility assessment on these features. For example, when a timing feature shows that the estimated junction temperature corresponding to the change in conduction current exceeds the upper limit allowed by the device's technical specifications, the server suppresses the response amplitude of that feature; when the leakage current growth trend is consistent with the material insulation degradation law, the server retains the feature and improves its reliability.
[0132] The server then performs stress damage accumulation path tracing on the multi-source stress coupling feature representation set generated after injecting property constraints. The server calculates the damage accumulation direction vector of the feature vector point-by-point along the time dimension under the constraints of the degraded property constraint layer. For example, it maps continuous high temperature, repeated current surges, and increased on-state voltage drop together as the thermal fatigue damage accumulation direction, and maps increased humidity, decreased insulation resistance, and increased leakage current together as the insulation degradation damage direction, thereby generating a stress damage accumulation path curve. This curve characterizes the combined damage evolution process caused by multi-source stresses to the power switching device under property constraints.
[0133] The server reclassifies the degradation stages based on the stress damage accumulation path curve. It identifies moments of significant curvature abrupt changes on the curve as degradation stage boundaries. For example, it marks the point where junction temperature cycling damage increases rapidly as the boundary between early and mid-stage degradation, and the point where leakage current accelerates as the boundary between mid-stage and high-risk degradation. Furthermore, it determines the coverage time range of each degradation stage cluster based on the path length of the path curve segments between adjacent boundaries, thus obtaining the property-constrained degradation stage classification results.
[0134] The server updates the lifetime stage transition probability array based on the partitioning results, ensuring that the number of degradation stage clusters and their temporal coverage are consistent with the property constraint degradation stages. The server assigns transition probabilities based on the path characteristic distance of the stress-damage accumulation path curve segments between each boundary point; the longer the path distance and the more drastic the damage change, the higher the probability of transitioning to a high-risk degradation stage. The server also fuses and splices together the cluster center feature vector of each degradation stage cluster, the mean of the projected coordinates of all feature vectors within that cluster on the stress-damage accumulation path curve, and the temporal information of the degradation stage boundary points to generate a state degradation trajectory vector set infused with property constraint information.
[0135] During the training of the lifetime prediction strategy network, the server further introduces a degradation property constraint regularization term. This regularization term performs a working limit boundary check on the mean parameter of the remaining effective lifetime prediction output by the network. When the mean remaining lifetime predicted by the network exceeds the theoretical lifetime upper limit derived from the maximum junction temperature, leakage current limit, and switching delay limit, the server increases the loss value of the strategy optimization loss function, forcing the lifetime prediction strategy network to output a remaining lifetime result that conforms to the actual physical degradation law of electronic components.
[0136] In this embodiment of the invention, after acquiring the operation monitoring data stream of heavy-duty vehicle electronic components, before performing degradation trajectory enhancement modeling processing on the operation monitoring data stream and generating the state degradation trajectory vector set and life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, the method further includes:
[0137] The operation monitoring data stream is processed to construct an operation condition diagram. With time as the horizontal axis and the combined parameter space of the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set as the vertical axis, the operation monitoring data stream is mapped into a multi-dimensional operation condition diagram representation. In the multi-dimensional operation condition diagram representation, each graph node corresponds to a multi-dimensional parameter combination value at a time point, and the edges between graph nodes reflect the sequential relationship in time.
[0138] The multidimensional operating condition graph representation is subjected to graph structure feature extraction processing to extract the node degree distribution characteristics, graph connectivity component characteristics, and closed loop detection characteristics of the multidimensional operating condition graph representation, and to construct an operating condition graph structure feature set, which includes a node degree distribution histogram description vector, a connected component identifier list, and a closed loop path list.
[0139] The operating condition diagram structure feature set is used to perform operating condition mode discrimination processing on the operating monitoring data stream. The node degree distribution histogram description vector is input into the pre-constructed operating condition mode classifier to obtain the operating condition mode identification result. The operating condition mode identification result includes a stable operating condition mode identifier, a periodic alternating operating condition mode identifier, or a drastic transient operating condition mode identifier.
[0140] Based on the operating condition mode recognition result, the network structure configuration parameters of the degenerate trajectory encoder are adaptively selected. When the operating condition mode recognition result is a stable operating condition mode identifier, the configuration parameters of the degenerate trajectory encoder with a deep temporal convolutional structure are selected. When the operating condition mode recognition result is a violent transient operating condition mode identifier, the configuration parameters of the degenerate trajectory encoder with a shallow temporal convolutional structure and an increased number of attention channels are selected.
[0141] The degenerate trajectory encoder is instantiated using the adaptively selected configuration parameters to generate a degenerate trajectory encoder instance that is self-adaptive to the current operating mode. The network depth, number of feature map channels, and number of attention heads of the degenerate trajectory encoder instance are based on the network structure configuration parameters.
[0142] The closed loop path list in the structural feature set of the operating condition diagram is subjected to cyclic operating condition annotation processing. The time period covered by each closed loop in the closed loop path list is marked as a cyclic load cycle segment. The cycle duration characteristics and parameter fluctuation amplitude characteristics within the cycle segment are extracted to generate cyclic operating condition description information.
[0143] The cyclic operating condition description information is injected into the degraded trajectory encoder instance. A cyclic sensing position encoding layer is added to the degraded trajectory encoder instance. The cyclic sensing position encoding layer constructs a periodic regular position encoding vector according to the period duration characteristics. The periodic regular position encoding vector is added and fused element by element with the time position encoding output of the multi-layer time convolution operation layer in the degraded trajectory encoder instance to generate a fused position encoding with cyclic operating condition sensing capability.
[0144] After obtaining the set of multi-source stress coupling feature representations output by the degradation trajectory encoder instance after injecting the cyclic operating condition perception capability, the set of multi-source stress coupling feature representations is subjected to subsequent degradation stage clustering analysis to generate a set of state degradation trajectory vectors and a life stage transition probability array injected with the adaptive characteristics of the operating condition mode.
[0145] In this embodiment of the invention, for example, after the server obtains the operational monitoring data stream corresponding to the electronic component identifier "HV-TRUCK-023-ECU-PWR-MOS-05", it first performs operational condition graph construction processing. The server maps the continuous monitoring data into a multi-dimensional operational condition graph representation, with time as the horizontal axis and the combined parameter space composed of load condition parameters, environmental stress parameters, and electrical response parameters as the vertical axis. Each graph node corresponds to a multi-dimensional parameter combination value at a sampling time point. For example, a node may simultaneously record traction current, engine torque, controller cavity temperature, vibration acceleration, output voltage ripple, and conduction voltage drop. The edges between adjacent nodes represent the temporal sequence of monitoring data, enabling the server to observe the process of the component from stable operation to stress accumulation and then to response shift in a graph structure.
[0146] The server then performs graph structure feature extraction processing on the multi-dimensional operating condition graph representation. The server statistically analyzes the degree distribution of each node in the graph, forming a node degree distribution histogram description vector; it identifies connected components formed during different operating stages and generates a list of connected component identifiers; it detects closed loops formed by the vehicle during repeated starts, accelerations, braking, and restarts, obtaining a list of closed loop paths. For heavy-duty transport vehicles in the mining area, the server identifies multiple closed loops in the graph structure consisting of "full-load start—climbing high load—slow braking—low-speed waiting," indicating that the electronic component is under long-term periodic alternating load conditions.
[0147] The server uses a set of structural features from the operating condition graph to perform operating condition mode discrimination. The server inputs the node degree distribution histogram description vector into a pre-built operating condition mode classifier to obtain the operating condition mode identification result. When the nodes in the graph structure change gradually and the connectivity is simple, the server outputs a stable operating condition mode identifier; when there are many closed loops with obvious periods, the server outputs a periodic alternating operating condition mode identifier; when nodes jump frequently and local edges change drastically, the server outputs a drastic transient operating condition mode identifier. For the aforementioned power switching device, the server identified a periodic alternating operating condition mode, indicating that the device is affected by both repetitive load cycles and temperature cycles.
[0148] The server adaptively selects the network structure configuration parameters of the degenerate trajectory encoder based on the operating condition pattern recognition results. For stable operating conditions, the server selects configuration parameters with deep temporal convolutional structures to capture long-term, slow degradation trends; for drastic transient operating conditions, the server selects shallow temporal convolutional structures and increases the number of attention channels to quickly capture short-term shocks and abrupt changes; for periodic alternating operating conditions, the server selects configuration parameters that combine moderate temporal convolutional depth with strong periodic awareness. The server instantiates the degenerate trajectory encoder using the selected configuration parameters, ensuring that the encoder's network depth, feature map channel count, and attention head count are all matched to the current operating condition pattern.
[0149] The server further processes the closed-loop path list with cyclic operating condition annotation, marking the time periods covered by each closed loop as cyclic load cycle segments and extracting cycle duration characteristics and parameter fluctuation amplitude characteristics within the cycle. For example, the server identifies that the vehicle completes a cycle of "loading start-up - high load hill climb-braking wait" every 18 minutes, with traction current fluctuation reaching 120A and controller cavity temperature fluctuation reaching 18℃ within the cycle, thereby generating cyclic operating condition description information. The server injects this cyclic operating condition description information into the degenerate trajectory encoder instance and adds a cyclic sensing position encoding layer to it. The cyclic sensing position encoding layer constructs a cyclic regular position encoding vector based on the cycle duration and fuses it element-wise with the temporal position encoding output of the multi-layer temporal convolution operation layer to generate a fused position encoding with cyclic operating condition sensing capability.
[0150] Finally, the server obtains the multi-source stress coupling feature representation set output by the degradation trajectory encoder instance after injecting cyclic operating condition perception capability. Since this feature set simultaneously reflects the load impact, environmental stress, electrical response, and periodic alternating operating condition patterns, the server can more accurately distinguish between normal periodic fluctuations and the true degradation trend when performing subsequent degradation stage clustering analysis on it, ultimately generating a state degradation trajectory vector set and a life stage transition probability array with adaptive operating condition mode characteristics.
[0151] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned reinforcement learning-based intelligent prediction method for the lifespan of heavy-duty automotive electronic components. Figure 4 As shown, Figure 4 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0152] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A heavy-duty vehicle electronic component life intelligent prediction method based on reinforcement learning, characterized in that, The method includes: The operation monitoring data stream of heavy-duty vehicle electronic components is acquired. The operation monitoring data stream includes electronic component identifiers, a set of time-series load condition parameters associated with the electronic component identifiers, a set of time-series environmental stress parameters associated with the electronic component identifiers, and a set of time-series electrical response parameters associated with the electronic component identifiers. The operation monitoring data stream is subjected to degradation trajectory enhancement modeling processing. Based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier are generated. A lifetime prediction policy network is constructed based on the state degradation trajectory vector set and the lifetime stage transition probability array, and the lifetime prediction policy network is trained using the lifetime stage transition probability array to obtain a set of weight parameters for the lifetime prediction policy network that has been trained and converged. The lifetime prediction strategy network, which has been trained and converged, is invoked to process the set of state degradation trajectory vectors to generate the remaining effective lifetime prediction distribution corresponding to the electronic component identifier. A set of component replacement scheduling instructions is generated based on the predicted distribution of remaining effective lifetime. The set of component replacement scheduling instructions includes a sequence of identifiers of electronic components to be replaced and corresponding time window information.
2. The method of claim 1, wherein, The degradation trajectory enhancement modeling process for the operation monitoring data stream, based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, generates a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier, including: The time-series load condition parameter set is subjected to load shock event detection processing. The load change time point and load change amplitude value are extracted from the time-series load condition parameter set to generate a load shock event sequence. Each load shock event in the load shock event sequence includes a load change time point label and a load change amplitude value label. The time-series environmental stress parameter set is segmented into environmental stress fluctuation segments to obtain a set of environmental stress fluctuation segments. Each environmental stress fluctuation segment in the set of environmental stress fluctuation segments includes the start time of environmental stress change, the end time of environmental stress change, and the gradient sequence of environmental stress change. The time-series electrical response parameter set is subjected to response anomaly offset identification processing. The starting point of the steady-state offset of the electrical response and the cumulative amount of the steady-state offset of the electrical response are extracted from the time-series electrical response parameter set to generate an electrical response offset feature sequence. The electrical response offset feature sequence includes the electrical response offset starting timestamp and the electrical response offset cumulative amount value. The load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence are input into a pre-constructed degenerate trajectory encoder. The degenerate trajectory encoder includes multiple layers of temporal convolution operation layers and channel cross-attention operation layers. The multiple layers of temporal convolution operation layers are used to extract multi-dimensional temporal features from the load impact event sequence, the set of environmental stress fluctuation segments, and the electrical response offset feature sequence to obtain load impact temporal feature maps, environmental stress temporal feature maps, and electrical response temporal feature maps. The channel cross-attention operation layers are used to perform cross-channel interactive weight allocation processing on the load impact temporal feature maps, environmental stress temporal feature maps, and electrical response temporal feature maps to generate a multi-source stress coupling feature representation set. The degradation stage clustering analysis is performed on the multi-source stress coupling feature representation set. The feature distribution distance between any two time-domain feature vectors in the degradation stage space is calculated, and the time-domain feature vectors that satisfy the feature distribution distance constraint are assigned to the same degradation stage cluster. The degradation stage division result corresponding to the electronic component identifier is obtained. The degradation stage division result includes several degradation stage clusters and the time-domain coverage of each degradation stage cluster. The degradation stage division results are processed by degradation stage transition relationship mining, the stage transition frequency between different degradation stage clusters within adjacent time windows is counted, and the stage transition frequency is normalized to obtain the lifetime stage transition probability array corresponding to the electronic component identifier. The row index and column index of the lifetime stage transition probability array correspond to different degradation stage clusters. Each element in the lifetime stage transition probability array represents the transition probability value from the degradation stage cluster corresponding to the row index to the degradation stage cluster corresponding to the column index. The degradation stage division results are processed by degradation trajectory vector mapping. The central feature vector of each degradation stage cluster is concatenated with the start and end time information of its time domain coverage to generate the state degradation trajectory vector set corresponding to the electronic component identifier. Each degradation trajectory element in the state degradation trajectory vector set includes the degradation stage central representation vector, the degradation stage start time label, and the degradation stage end time label.
3. The method according to claim 1, characterized in that, The process involves constructing a lifetime prediction policy network based on the state degradation trajectory vector set and the lifetime stage transition probability array, and training the lifetime prediction policy network using the lifetime stage transition probability array to obtain a set of weight parameters for a converged lifetime prediction policy network, including: The degraded trajectory vector set is subjected to temporal reorganization processing to obtain a degraded trajectory temporal sequence; An input state representation sequence is generated based on the degenerate trajectory time sequence; Construct a lifetime prediction policy network that includes state encoding, policy reasoning, and lifetime prediction output functions; The stage transition information corresponding to the input state representation sequence is determined based on the lifetime stage transition probability array. The reinforcement learning reward is calculated based on the stage transition information and the remaining effective lifetime prediction probability distribution output by the lifetime prediction strategy network, and the weight parameters of the lifetime prediction strategy network are updated based on the reinforcement learning reward. When the preset training convergence conditions are met, the set of weight parameters of the lifetime prediction strategy network that has achieved training convergence is obtained.
4. The method according to claim 1, characterized in that, The lifetime prediction strategy network, which has been trained and converged, is used to process the set of state degradation trajectory vectors to generate the remaining effective lifetime prediction distribution corresponding to the electronic component identifier, including: Extract the current state observation segment from the set of state degradation trajectory vectors; The current state observation segment is input into the lifetime prediction policy network that has been trained and converged to generate an initial remaining effective lifetime prediction distribution. Based on the lifetime stage transition probability array, extract the stage transition probability vector corresponding to the degradation stage at the end of the current state observation segment; The initial remaining effective lifetime prediction distribution is corrected by using the stage transition probability vector to obtain the corrected remaining effective lifetime prediction distribution. The modified remaining effective lifetime prediction distribution is subjected to time-series smoothing and normalization reshaping to obtain the remaining effective lifetime prediction distribution corresponding to the electronic component identifier.
5. The method according to claim 1, characterized in that, The step of generating a set of component replacement scheduling instructions based on the predicted distribution of remaining effective lifetime includes: The remaining effective lifetime prediction distribution is processed to locate the failure risk time point, so as to obtain the risk warning time point and the expected failure time point; Based on the predicted remaining effective lifetime distribution, the reliable remaining working time interval is extracted; A replacement time window selection strategy is constructed based on the risk warning time point, the expected failure time point, and the reliable remaining working time interval, and the replacement time window range is determined. A set of candidate replacement time points is generated based on the replacement time window range; The replacement cost is evaluated for the set of candidate replacement time points, and the optimal replacement time is selected; A set of component replacement scheduling instructions is generated based on the electronic component identifier and the optimal replacement time.
6. The method according to claim 1, characterized in that, After acquiring the operational monitoring data stream of heavy-duty vehicle electronic components, the method further includes: The operation monitoring data stream is subjected to multi-source heterogeneous sampling frequency alignment processing to obtain a time-aligned operation monitoring data stream; Construct a multidimensional state space representation vector sequence of the reinforcement learning environment based on the time-aligned running monitoring data stream; The state transition dynamics mechanism of the reinforcement learning environment is constructed based on the lifetime stage transition probability array. Define a reinforcement learning action space that includes continuing to monitor waiting actions and immediately executing replacement actions; A replacement decision reward function is constructed based on the remaining effective lifetime prediction distribution, the reliable remaining working time interval, and the risk warning time point; Action selection and monitoring time interval adjustment are performed based on the remaining effective lifetime prediction distribution, and new operation monitoring data streams are appended to the time-aligned operation monitoring data stream to update the multidimensional state space representation vector sequence, the remaining effective lifetime prediction distribution, and the replacement time window selection strategy.
7. The method according to claim 1, characterized in that, After performing degradation trajectory enhancement modeling on the operational monitoring data stream, and generating a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, the process further includes: The state degradation trajectory vector set and the lifetime stage transition probability array are subjected to cross-component degradation mode migration processing. The probability subarrays with shared degradation structure characteristics in the lifetime stage transition probability arrays corresponding to different electronic component identifiers are extracted, and the degradation stage clusters corresponding to the probability subarrays with shared degradation structure characteristics are marked as general degradation mode stage clusters. The common degradation feature extraction process is performed on the general degradation mode stage cluster. All degradation stage center representation vectors belonging to the general degradation mode stage cluster from different electronic component identifiers are aggregated. The common factor decomposition process is performed on the aggregated degradation stage center representation vectors to obtain a general degradation mode representation vector basis set. The general degradation mode representation vector basis set contains a set of basis vectors for linearly representing the general degradation modes of different components. The general degradation mode characterization vector basis set is injected into the channel cross-attention operation layer of the degradation trajectory encoder. A general degradation mode matching branch is added to the channel cross-attention operation layer of the degradation trajectory encoder. The general degradation mode matching branch calculates the mode matching similarity between the input load impact time series feature map, the environmental stress time series feature map and the electrical response time series feature map and each general degradation mode basis vector. The pattern matching similarity output by the general degradation pattern matching branch is used to perform pattern enhancement processing on the multi-source stress coupling feature representation set. The pattern matching similarity is used as the pattern enhancement weight coefficient and multiplied with the corresponding time-domain feature vector in the multi-source stress coupling feature representation set to obtain the pattern-enhanced multi-source stress coupling feature representation set. The degradation stage clustering analysis is re-executed on the enhanced multi-source stress coupling feature representation set to obtain the degradation stage division result based on the general degradation mode enhancement. The state degradation trajectory vector set and lifetime stage transition probability array corresponding to the electronic component identifier are updated based on the degradation stage division result based on the general degradation mode enhancement. When training the lifetime prediction strategy network, meta-policy gradient sharing is performed on the lifetime prediction strategy networks corresponding to different electronic component identifiers. The gradients of the policy optimization loss functions corresponding to multiple electronic component identifiers are weighted and averaged to obtain the meta-policy gradient. The meta-policy gradient is used to update the weight parameters of the lifetime prediction strategy networks corresponding to each electronic component identifier simultaneously, thus obtaining a lifetime prediction strategy network trained across components. A degradation trajectory prediction mechanism for cold start of novel electronic components is introduced. When a novel electronic component identifier not included in the training data appears, the initial stage operation monitoring data stream of the novel electronic component identifier is obtained, the initial state degradation trajectory vector set of the novel electronic component identifier is generated, and the initial state degradation trajectory vector set is processed by the general degradation mode characterization vector basis set to obtain the preliminary lifetime stage transition probability array estimate of the novel electronic component identifier. The preliminary lifetime stage transition probability array estimate of the novel electronic component identifier is rapidly adapted to the lifetime prediction strategy network trained across components. The preliminary lifetime stage transition probability array estimate is used as the environmental prior knowledge of the lifetime prediction strategy network. The deviation between the remaining lifetime prediction mean parameter output of the lifetime prediction strategy network trained across components and the actual remaining lifetime observation value of the novel electronic component identifier is minimized and fine-tuned to obtain the adapted remaining lifetime prediction distribution of the novel component.
8. The method according to claim 1, characterized in that, The degradation trajectory enhancement modeling process for the operation monitoring data stream, based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, generates a state degradation trajectory vector set and a life stage transition probability array corresponding to the electronic component identifier, including: Construct a degradation property constraint layer, which includes material property parameters, process property parameters, and operating limit property parameters of electronic components; The degenerate property constraint layer is injected into the degenerate trajectory encoder, and the temporal features are subjected to physical feasibility discrimination processing through the degenerate property constraint gating unit; Based on the degraded physical property constraint layer, stress damage accumulation path tracing is performed on the multi-source stress coupling feature representation set to generate stress damage accumulation path curves. Based on the stress damage accumulation path curve, the property constraint degradation stage division result is determined, and the state degradation trajectory vector set and the lifetime stage transition probability array are updated. During the training process of the lifetime prediction strategy network, a degradation property constraint regularization term is introduced to constrain the prediction output of the lifetime prediction strategy network to conform to the physical degradation law corresponding to the working limit characteristic parameter.
9. The method according to claim 1, characterized in that, After acquiring the operation monitoring data stream of heavy-duty vehicle electronic components, and before performing degradation trajectory enhancement modeling processing on the operation monitoring data stream and generating the state degradation trajectory vector set and life stage transition probability array corresponding to the electronic component identifier based on the time-series load condition parameter set, the time-series environmental stress parameter set, and the time-series electrical response parameter set, the method further includes: The operation monitoring data stream is processed to construct an operation condition diagram, resulting in a multi-dimensional operation condition diagram representation; The multidimensional operating condition diagram representation is subjected to graph structure feature extraction processing to obtain a set of operating condition diagram structure features; The operating condition pattern recognition result is obtained based on the structural feature set of the operating condition diagram. Based on the operating condition mode recognition results, the network structure configuration parameters of the degenerate trajectory encoder are selected, and a degenerate trajectory encoder instance that is self-adaptive to the current operating condition mode is generated. Based on the list of closed loop paths in the structural feature set of the operating condition diagram, cyclic operating condition description information is generated. The cyclic operating condition description information is injected into the degradation trajectory encoder instance, and a set of state degradation trajectory vectors and a lifetime stage transition probability array with adaptive characteristics of operating condition mode are generated based on the degradation trajectory encoder instance.
10. A lifespan intelligent prediction system for heavy-duty automotive electronic components based on reinforcement learning, characterized in that, Includes at least one service node; The service node includes a storage unit and a computing unit; the storage unit is used to store program code; the computing unit is used to run the program code to execute the reinforcement learning-based intelligent prediction method for the lifespan of heavy-duty vehicle electronic components as described in any one of claims 1 to 9.