Diesel engine-based carbon emission optimization method

By combining multi-source sensing and adaptive filtering technology with spatiotemporal graph convolutional networks and digital twins, the real-time monitoring problem of diesel engine crankshaft dynamic balance detection was solved, achieving optimization of diesel engine carbon emissions and improvement of operational stability.

CN121676162APending Publication Date: 2026-03-17NANJING VOCATIONAL UNIV OF IND TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing diesel engine crankshaft dynamic balancing testing technology cannot monitor dynamic changes in real time, making it difficult to deal with sudden imbalance problems, resulting in severe vibration of the whole engine, incomplete combustion, and increased carbon emissions.

Method used

Multi-source sensing and adaptive filtering techniques are used to extract crankshaft dynamic balance features. The spatiotemporal graph convolutional network and digital twin are combined for state prediction. Multi-objective reinforcement learning agents are used to generate collaborative optimization instructions to adjust fuel injection parameters and active electromagnetic balancer.

Benefits of technology

It achieves forward-looking optimization of diesel engine carbon emissions, improves combustion efficiency, reduces carbon emissions, and ensures stable operation and mechanical reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121676162A_ABST
    Figure CN121676162A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon emission optimization method based on a diesel engine, and relates to the field of diesel engines. According to the method, crankshaft dynamic balance characteristics are accurately extracted through a multi-source sensing and adaptive filtering technology, and state association and evolution laws of a mechanical structure are deeply excavated by using a space-time diagram convolutional network; a digital twinborn body fused with physical constraints is constructed to realize prospective prediction of carbon emission intensity, and finally, a collaborative optimization instruction is generated by means of a multi-target reinforcement learning agent to realize accurate regulation and control of an oil injection system and a balancer, so that the combustion efficiency of a diesel engine is remarkably improved, carbon emission is effectively reduced, and the service life of the diesel engine is prolonged. And meanwhile, the operation stability and the mechanical reliability are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of diesel engines, and more specifically, to a method for optimizing carbon emissions from diesel engines. Background Technology

[0002] Diesel engines, as an important power source, are widely used in ship propulsion systems, heavy-duty trucks, and stationary generator sets. These devices are often in harsh environments with high loads, variable speeds, and frequent changes in operating conditions. Internally, they face mechanical structural changes caused by crankshaft manufacturing tolerances, long-term wear, and thermal deformation. Externally, they are affected by multiple factors such as load fluctuations, ambient temperature changes, and vibration interference. As the core moving component of a diesel engine, the dynamic balance performance of the crankshaft directly affects the smoothness of the entire engine's operation and combustion efficiency. In actual operation, crankshaft dynamic balance deviations can easily cause severe vibrations in the entire engine, leading to decreased combustion chamber sealing, fuel injection timing deviation, and consequently, incomplete combustion, resulting in reduced fuel economy and increased harmful emissions.

[0003] Currently, the detection technology for dynamic balance of diesel engine crankshafts mainly relies on periodic offline testing, such as static verification using a balancing machine or laboratory measurement after disassembly. Although such methods can identify significant imbalances, they cannot capture dynamic changes under operating conditions in real time, nor can they cope with sudden imbalance problems. In addition, some existing online monitoring systems are based on simple vibration sensors and fixed threshold alarm mechanisms, which have defects such as low data sampling frequency, poor algorithm adaptability, and high false alarm rate. They lack the ability to accurately judge and compensate for dynamic imbalance states. Therefore, there is an urgent need in this field for a technical solution that can monitor the dynamic balance state of the crankshaft in real time, has high-frequency data acquisition and intelligent analysis capabilities, and can realize dynamic control and predictive maintenance to effectively optimize the carbon emission performance of diesel engines. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a carbon emission optimization method based on diesel engines. It accurately extracts crankshaft dynamic balance features through multi-source sensing and adaptive filtering technology, deeply mines the state correlation and evolution law of mechanical structures using spatiotemporal graph convolutional networks, constructs a digital twin that integrates physical constraints to achieve forward prediction of carbon emission intensity, and finally uses a multi-objective reinforcement learning agent to generate collaborative optimization instructions to solve the problems mentioned in the background technology.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] A method for optimizing carbon emissions based on diesel engines, characterized by comprising:

[0007] Step S1: Based on the multi-source sensors installed on the diesel engine crankshaft, the raw data of vibration acceleration, speed pulse sequence and temperature signal are collected synchronously; the raw data are fused using an adaptive weighted Kalman filter algorithm, and the time-frequency domain features are extracted from the filtered signal using a wavelet packet transform algorithm, thereby generating a multi-dimensional feature vector to describe the dynamic balance state of the crankshaft.

[0008] Step S2: Input the multidimensional feature vector into the pre-constructed spatiotemporal graph convolutional network model, and use the spatiotemporal graph convolutional network model to analyze the spatial topological relationship and temporal change relationship of the multidimensional feature vector on the graph structure abstracted from the crankshaft connecting rod piston mechanism, and output the mode probability distribution vector and the miscalculation prediction vector used to identify the current dynamic equilibrium state.

[0009] Step S3: Input the mode probability distribution vector and the imbalance prediction vector as the initial state into the digital twin of the diesel engine crankshaft. The digital twin is a long short-term memory network model that incorporates physical constraints. It is used to perform prediction calculations and output a sequence of vibration characteristics, imbalance state and predicted carbon emission intensity for multiple future time steps.

[0010] Step S4: Input the predicted carbon emission intensity sequence into a multi-objective reinforcement learning agent; the multi-objective reinforcement learning agent uses a reward function to perform decision-making operations and generate a collaborative optimization control instruction set; the collaborative optimization control instruction set is sent to the diesel engine controller, and the controller adjusts the fuel injection parameters or the active electromagnetic balancer according to the instruction set to achieve carbon emission optimization.

[0011] In the aforementioned method for optimizing carbon emissions based on a diesel engine, in step S1, the multi-source sensors include a high-frequency vibration acceleration sensor mounted on the crankcase, a photoelectric encoder mounted at the end of the crankshaft, and a temperature sensor embedded in the main bearing bore.

[0012] The raw data includes raw vibration acceleration signals from a high-frequency vibration acceleration sensor, raw rotational speed pulse sequences from an optical encoder, and raw temperature signals from a temperature sensor.

[0013] The step of fusing the original data using the adaptive weighted Kalman filter algorithm includes:

[0014] The original rotational speed pulse sequence is preprocessed, and the instantaneous angular velocity is calculated by measuring the pulse interval, and then the instantaneous rotational speed and instantaneous angular acceleration are further calculated.

[0015] Based on the absolute value of the instantaneous angular acceleration, the weights of the process noise covariance matrix in the filtering algorithm are dynamically adjusted. The larger the absolute value of the angular acceleration, the larger the value of the process noise covariance matrix.

[0016] The weights of the measurement noise covariance matrix are dynamically adjusted based on the absolute difference between the real-time temperature signal and the reference temperature value. The larger the absolute temperature difference, the larger the value of the measurement noise covariance matrix.

[0017] Using the dynamically adjusted process noise covariance matrix and measurement noise covariance matrix, the original vibration acceleration signal is filtered and calculated to output a high-fidelity vibration signal after noise reduction.

[0018] The aforementioned method for optimizing carbon emissions based on diesel engines includes the following step: extracting time-frequency domain features from the filtered signal using a wavelet packet transform algorithm.

[0019] Based on the instantaneous rotational speed calculated from the rotational speed pulse sequence, the core frequency band corresponding to the crankshaft rotational fundamental frequency and its second and third harmonics is determined in real time;

[0020] The filtered high-fidelity vibration signal is decomposed using wavelet packets, and the energy value of each wavelet packet node within the core frequency band is calculated.

[0021] Calculate the total energy of each core frequency band, and calculate the proportion of energy of each node in the total energy of that frequency band;

[0022] The energy entropy value of each core frequency band is calculated based on the energy ratio.

[0023] In the aforementioned method for optimizing carbon emissions based on diesel engines, step S2, the construction and feature mapping process of the graph structure abstracted from the crankshaft connecting rod piston mechanism, is as follows:

[0024] The multidimensional feature vector from step S1 is received, and the main journal of the crankshaft, the crank pin, and the large and small ends of the connecting rod are abstracted as nodes in the graph structure. The edges between the nodes are defined according to the actual physical connection relationship between the components. The edges represent the connection relationship, and the weight of the edges is calculated by the dynamic coupling coefficient between the components. Finally, a graph structure containing a set of nodes, a set of edges, and a weighted adjacency matrix is ​​formed, and the features in the multidimensional feature vector are assigned to the corresponding nodes in the graph structure according to their physical meaning, as the initial feature attributes of each node.

[0025] The aforementioned method for optimizing carbon emissions based on diesel engines involves analyzing spatial topological relationships as follows:

[0026] For each node in the graph structure, the feature information of all neighboring nodes is aggregated. During the aggregation operation, the influence of the features of neighboring nodes is adjusted according to the physical connection strength weights defined in the adjacency matrix. At the same time, the normalization coefficient is calculated based on the degree of the node itself and its neighbors, and the aggregated features are normalized. Finally, the normalized aggregated features are linearly transformed with the learnable convolution kernel parameters, and the updated feature vector of the node is obtained through a non-linear activation function.

[0027] The aforementioned method for optimizing carbon emissions based on diesel engines involves analyzing the time-series variation relationships as follows:

[0028] Based on spatial topological relationship analysis, a gate mechanism is used to perform convolution operations along the time dimension. The gate mechanism adaptively controls the retention ratio of historical state information and current candidate state information through an update gate calculated by a temporal convolution kernel. Both historical state information and current candidate state information are extracted from the node feature sequence through different temporal convolution kernels. Finally, the update gate is used to fuse the historical state and the current candidate state to form the feature representation of the node in the next time step.

[0029] The process of outputting the mode probability distribution vector and the miscalculation prediction vector to identify the current dynamic equilibrium state is as follows:

[0030] The global features of the graph extracted after processing by the spatiotemporal graph convolutional network are input into two parallel fully connected layers and an output layer, respectively. One output layer uses the Softmax function to output a pattern probability distribution vector, and the other output layer uses the linear regression function to output a miscalculation prediction vector.

[0031] In the aforementioned carbon emission optimization method based on diesel engines, step S3, which involves inputting the mode probability distribution vector and the miscalculation prediction vector as initial states into the digital twin, is as follows:

[0032] Extract the state pattern category with the highest probability from the pattern probability distribution vector, and combine the pattern category with the imbalance amplitude and imbalance phase in the imbalance prediction vector to form a physically interpretable initial state vector;

[0033] The initial state vector is input into a long short-term memory network model that incorporates physical constraints. The network model employs a composite loss function during training, which is composed of a weighted average of a data prediction error term and a physical law residual term. The physical law residual term is calculated by substituting the crankshaft angular acceleration predicted by the network into the diesel engine rigid body rotation dynamics equation.

[0034] The aforementioned carbon emission optimization method based on diesel engines involves the following process for performing forward-looking prediction calculations: A digital twin, starting from an initial state, recursively calculates using a long short-term memory network to predict vibration characteristic parameters, imbalance amplitude, and imbalance phase at multiple future moments step by step. Simultaneously, based on the currently predicted vibration amplitude and imbalance amplitude, a preset carbon emission intensity calculation model outputs the corresponding carbon emission intensity value in real time. Finally, the vibration characteristics, imbalance state data, and carbon emission intensity values ​​from multiple consecutive time steps are encapsulated into independent prediction sequences in chronological order. The vibration characteristic sequence includes the vibration spectrum characteristics at each moment, the imbalance state sequence includes the amplitude and phase combinations at each moment, and the carbon emission intensity sequence includes instantaneous carbon emission values ​​sorted by time.

[0035] In the aforementioned carbon emission optimization method based on diesel engines, step S4 involves constructing the reward function of the multi-objective reinforcement learning agent using a dynamic weight allocation mechanism. The process is as follows:

[0036] Based on the future multi-step carbon emission intensity values ​​extracted from the prediction sequence and the cumulative effect of the carbon emission intensity values, and combined with the amplitude of the real-time vibration prediction vector, a dynamic priority factor is calculated. The value of the dynamic priority factor is between 0 and 1, and it is used to adjust the weight ratio of the carbon emission optimization target in the reward function in real time. At the same time, the reward function uses the square of the vibration amplitude as a penalty term and the square of the change in the control command as a smoothness penalty term, which together constitute a reward function with multi-objective trade-offs.

[0037] When learning a strategy, the agent learns a mapping strategy from state to action by maximizing the cumulative reward, provided that the vibration constraint is met, i.e., the amplitude of the vibration prediction vector at the current moment is less than the preset vibration threshold.

[0038] The agent outputs a raw action vector, the dimension of which corresponds to the number of all adjustable parameters of the diesel engine.

[0039] The aforementioned method for optimizing carbon emissions based on diesel engines involves the following process: generating and issuing a collaborative optimization control command set; a multi-objective reinforcement learning agent calculates the original action vector based on a reward function; the action vector is decoded into a specific control command set, which includes a subset of injection parameter commands and a subset of active balancer commands; the injection parameter command subset includes main injection quantity commands, pre-injection quantity commands, and injection advance angle commands; the active balancer command subset includes compensation force amplitude commands and compensation force phase commands; before the commands are issued to the controller, boundary verification is performed by a preset safety verification unit; the safety verification unit incorporates the steady-state and transient safe operating boundary conditions of the diesel engine, compares the control commands with the preset safety boundaries, and if the control commands exceed the safety boundaries, the corresponding control commands are trimmed to the nearest boundary value; finally, the verified control command set is issued to the diesel engine controller for execution via the bus.

[0040] The beneficial effects achieved by this invention are as follows: The carbon emission optimization method based on diesel engines accurately extracts crankshaft dynamic balance features through multi-source sensing and adaptive filtering technology, deeply mines the state correlation and evolution law of mechanical structures using spatiotemporal graph convolutional networks, constructs a digital twin that integrates physical constraints to achieve forward prediction of carbon emission intensity, and finally uses a multi-objective reinforcement learning agent to generate collaborative optimization instructions to achieve precise control of the fuel injection system and balancer, thereby significantly improving the combustion efficiency of diesel engines, effectively reducing carbon emissions, and ensuring operational stability and mechanical reliability. Attached Figure Description

[0041] Figure 1 This is a flowchart of the carbon emission optimization method based on diesel engines according to the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0044] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0045] Example 1

[0046] This embodiment provides, for example Figure 1 The method for optimizing carbon emissions based on diesel engines, as shown, specifically includes the following steps:

[0047] Step S1: Based on the multi-source sensors installed on the diesel engine crankshaft, synchronously collect vibration acceleration. Rotational speed pulse sequence and temperature signal The original data is processed by an adaptive weighted Kalman filter algorithm to suppress noise, and a wavelet packet transform algorithm is used to extract time-frequency domain features from the filtered signal, thereby generating a multi-dimensional feature vector to describe the dynamic balance state of the crankshaft.

[0048] Step S2: Input the multidimensional feature vector into the pre-constructed spatiotemporal graph convolutional network model, and use the spatiotemporal graph convolutional network model to analyze the spatial topological relationship and temporal change relationship of the multidimensional feature vector on the graph structure abstracted by the crankshaft connecting rod piston mechanism, and output the mode probability distribution vector and the miscalculation prediction vector used to identify the current dynamic equilibrium state.

[0049] Step S3: Input the mode probability distribution vector and the imbalance prediction vector as the initial state into the digital twin of the diesel engine crankshaft. The digital twin is a long short-term memory network model that incorporates physical constraints. It is used to perform forward prediction calculations and output the vibration characteristics, imbalance state and predicted carbon emission intensity sequence for multiple future time steps.

[0050] Step S4: Input the predicted carbon emission intensity sequence into the multi-objective reinforcement learning agent; the multi-objective reinforcement learning agent uses the reward function to perform decision-making calculations and generate a collaborative optimization control instruction set; the collaborative optimization control instruction set is sent to the diesel engine controller to adjust the fuel injection parameters or the active electromagnetic balancer to achieve carbon emission optimization.

[0051] In this embodiment, it should be specifically noted that in step S1, the multi-source sensor includes a high-frequency vibration acceleration sensor mounted on the crankcase, with a sampling frequency of... An optical encoder installed at the end of the crankshaft and a temperature sensor embedded in the main bearing bore;

[0052] The raw data specifically includes raw vibration acceleration signals from a high-frequency vibration accelerometer, raw rotational speed pulse sequences from a photoelectric encoder, and raw temperature signals from a temperature sensor.

[0053] The specific steps for fusing the original data using the adaptive weighted Kalman filter algorithm are as follows:

[0054] First, the original rotational speed pulse sequence Preprocessing is performed, and the instantaneous angular velocity is calculated by measuring the pulse interval. And further calculate the instantaneous rotational speed. (Unit: revolutions per minute) and instantaneous angular acceleration (Used to characterize the intensity of motion); subsequently, based on instantaneous angular acceleration The absolute value of the angular acceleration is used to dynamically adjust the weights of the process noise covariance matrix in the filtering algorithm. The larger the absolute value of the angular acceleration, the larger the value of the process noise covariance matrix. The formula for the process noise covariance matrix is:

[0055] ;

[0056] in, The process noise covariance matrix represents the algorithm's estimate of the uncertainty of its own prediction model. This indicates that it is a diagonal matrix. The fundamental variance, representing process noise, is a fixed value pre-set based on sensor and system characteristics, and represents the basic model uncertainty under steady-state conditions. This represents the adjustment coefficient, a pre-set coefficient used to adjust the sensitivity of the process to changes in angular acceleration. This represents the absolute value of instantaneous angular acceleration, calculated in real-time from the engine speed signal. A larger value indicates more aggressive diesel engine operation, such as rapid acceleration or high load, and the more prone the predictive model is to failure. The formula's function is: the more aggressive the engine operation, i.e. Increase, the algorithm will automatically increase The value indicates that the algorithm will reduce its confidence in its own prediction model and will be more inclined to adopt the raw data measured by the sensors when filtering;

[0057] Simultaneously, the weights of the measurement noise covariance matrix are dynamically adjusted based on the absolute difference between the real-time temperature signal and the reference temperature value. The larger the absolute temperature difference, the larger the value of the measurement noise covariance matrix. The formula for the measurement noise covariance matrix is:

[0058] ;

[0059] in, Let represent the measurement noise covariance matrix, and let represent the algorithm's estimate of the uncertainty of the sensor measurements. The fundamental variance representing the measurement noise can be a fixed value preset based on the sensor's accuracy, and represents the sensor's basic measurement error under constant temperature conditions. This represents the adjustment coefficient, used to adjust the sensitivity of the measurement to changes in temperature noise. This represents the absolute difference between the current temperature and the reference temperature, used to quantify the intensity of thermal disturbances. Drastic temperature changes can affect the sensor's sensitivity and zero-point drift. The formula for the measurement noise covariance matrix utilizes additional errors; the greater the temperature change, the more significant the impact. Increase, it will automatically increase. The value of will reduce the confidence in the sensor measurement value, and will be more inclined to adopt the prediction value of its own model when filtering;

[0060] Finally, using the dynamically adjusted process noise covariance matrix and measurement noise covariance matrix described above, the original vibration acceleration signal is analyzed. Filtering calculations are performed to output a high-fidelity vibration signal after noise reduction. ;

[0061] The specific steps for extracting time-frequency domain features from the filtered signal using the wavelet packet transform algorithm are as follows:

[0062] First, the instantaneous rotational speed is calculated from the rotational speed pulse sequence. Real-time determination of the fundamental frequency of crankshaft rotation ( ) and its main second harmonic ( ) and triple frequency ( The corresponding core frequency band; then, the filtered high-fidelity vibration signal... Perform wavelet packet decomposition and calculate the energy value of each wavelet packet node within the core frequency band; then, calculate the total energy of each core frequency band and the proportion of each node's energy in the total energy of that frequency band; finally, calculate the energy entropy value of each core frequency band based on the energy proportion, using the following formula:

[0063] ;

[0064] in, This represents energy entropy, which is the degree of disorder in the distribution of vibrational energy within a specific frequency band. The larger the value, the more dispersed the energy distribution, which usually indicates a more abnormal state of imbalance. The total number of nodes refers to the total number of all wavelet packet nodes contained within a specific frequency band of a wavelet packet decomposition. Indicates the energy percentage, referring to the percentage within a specific frequency band. The energy of each wavelet packet node represents the proportion of the total energy in that frequency band, with values ​​ranging from 0 to 1. All nodes... The sum is 1, calculated as follows:

[0065] ,

[0066] The term represents the energy of a node, referring to the energy of the first node within a specific frequency band. The energy of a wavelet packet node is the sum of the squares of all wavelet packet coefficients of that node. The total energy of a frequency band refers to the sum of the energies of all wavelet packet nodes within a specific frequency band. The energy entropy values, instantaneous speed values, and real-time temperature values ​​of all core frequency bands are combined to form a multidimensional feature vector describing the dynamic balance state of the crankshaft.

[0067] This indicates how many different frequency band energy entropies are incorporated into the final feature vector, collectively used to describe the dynamic balance state of the crankshaft. This indicates that the feature row vector is transposed, representing a column vector.

[0068] In this embodiment, the process of constructing and feature mapping the graph structure abstracted from the crankshaft connecting rod piston mechanism in step S2 is as follows:

[0069] Receive the multidimensional feature vector from step S1, and abstract the crankshaft main journal, crank pin, and the large and small ends of the connecting rod as nodes in a graph structure; define the edges between nodes according to the actual physical connection relationships between the components, with each edge representing a connection relationship and its weight calculated from the dynamic coupling coefficient between the components; finally, a complete graph structure is formed, containing a set of nodes, a set of edges, and a weighted adjacency matrix. , Representing the graph structure, This represents the set of all nodes in a graph structure, where each node... This represents a physical component, such as a main journal or crank pin. This represents the set of all edges in a graph structure. An edge connects two nodes and represents the physical connection between components, such as the main journal being connected to the crank pin, or the crank pin being connected to the connecting rod. An adjacency matrix represents a graph structure and is used to describe the connection relationships between nodes in the graph structure. The element values ​​in the adjacency matrix... Represents a node and nodes The connection weight is a weighted value calculated based on the dynamic coupling coefficient, used to quantify the physical strength of the connection between nodes and to integrate the multidimensional feature vector. Each feature in the graph is assigned to a corresponding node in the graph structure based on its physical meaning, serving as the node's name. initial feature attributes ;

[0070] The specific process of analyzing spatial topological relationships is as follows:

[0071] For each node in the graph structure, the feature information of all its neighboring nodes is aggregated. During the aggregation operation, the influence of the neighboring node features is adjusted according to the physical connection strength weights defined in the adjacency matrix. Simultaneously, normalization coefficients are calculated based on the degree of the node itself and its neighbors to normalize the aggregated features. Finally, the normalized aggregated features are linearly transformed with learnable convolutional kernel parameters and then processed through a nonlinear activation function to obtain the updated feature vector for that node. This captures the structural dependencies within the crankshaft system. The formula for the updated feature vector is:

[0072] ;

[0073] in, Represents a node In the The updated feature vectors of the layer network These represent non-linear activation functions, including ReLU and the sigmoid function. Their purpose is to introduce non-linear transformations into the model, enhancing its expressive power. Indicates a node itself and the set of all neighboring nodes of this node. Each node in Perform calculations and summation to implement aggregation operations. This represents the symmetric normalization coefficient, used to normalize the aggregated features to eliminate the effects of uneven distribution of node degree (connection number). Represents a node The degree is usually calculated by adding the value after the self-loop, i.e. , It is a scalar value representing the total connection strength between a node and other nodes. Represents a node The degree, The learnable convolutional kernel weights are a matrix, representing parameters that the model continuously optimizes through gradient descent during training. They are used to perform a linear transformation on the aggregated features, i.e., feature mapping. In this formula, if... This indicates that the two nodes are not directly connected and no information is transmitted between them. The value of the node directly determines the node's value. Features of nodes The degree of influence (i.e., "the influence of moderating the features of neighboring nodes") Represents a node In the The current feature vector of the layer network;

[0074] The specific process of analyzing the temporal relationship is as follows:

[0075] Based on spatial topological relationship analysis, a gating mechanism is used to perform convolution operations along the time dimension. This gating mechanism adaptively controls the retention ratio of historical state information to current candidate state information through an update gate calculated by a temporal convolution kernel. The calculation formula for the update gate is as follows:

[0076] ;

[0077] in, The update gate is a gating signal output by the Sigmoid function, with a value between 0 and 1, which adaptively controls the proportion of historical states retained. This represents the Sigmoid activation function, which compresses the output to the (0,1) interval, thereby generating an effective gated signal. This represents the temporal convolution kernel parameters used to compute the update gate, and is a learnable weight matrix. This represents a 1D convolution operation along the time dimension, which is used to extract feature patterns from time-series data. Indicates the first The node feature sequence of a layer is a three-dimensional tensor, calculated as the number of nodes × feature dimension × time step T, representing the historical feature data of all nodes in the graph network over T consecutive time steps. This represents the update gate bias term; historical state information and current candidate state information are extracted from the node feature sequence through different temporal convolution kernels; finally, the update gate is used to fuse the historical state and the current candidate state to form the feature representation of the node in the next time step, thereby dynamically capturing the evolution trend of vibration features with crankshaft rotation angle, expressed as:

[0078] ;

[0079] in, This represents the output after the gated adaptive temporal convolution of this layer. Layer node characteristics, This represents the hyperbolic tangent activation function, which compresses the output to the (-1, 1) interval and is used to process candidate states. This represents element-wise multiplication, also known as the Hadamard product, used to fuse information according to the proportion of the gated signal. This represents the convolution kernel used to compute candidate states. and The model automatically optimizes its parameters through training, determining which temporal features to extract from historical data to generate gating and candidate states. Indicates the candidate state bias term;

[0080] The specific process for outputting the mode probability distribution vector and the miscalculation prediction vector used to identify the current dynamic equilibrium state is as follows:

[0081] The global features of the graph extracted after processing by the spatiotemporal graph convolutional network are input into two parallel fully connected layers and the output layer, respectively; one of the output layers uses the Softmax function to output a pattern probability distribution vector, the expression of which is:

[0082] ;

[0083] in, The pattern probability distribution vector is a column vector containing the probability that the current state is classified into each predefined pattern. This represents the probability that the dynamic equilibrium state belongs to a certain predefined pattern category, such as: Represents normal, Indicates a slight imbalance This represents a severe imbalance, with values ​​ranging from [0,1], and the sum of all elements being 1. This represents the total number of predefined dynamic equilibrium state modes. This means transposing the row vector to make it a column vector; the other output layer uses a linear regression function to output a miscalculation prediction vector, expressed as:

[0084] ;

[0085] in, The imbalance prediction vector is a column vector containing the imbalance physical quantities calculated from the data via regression. This represents the estimated magnitude of the imbalance, such as the product of mass and radius. It is a dimensionless relative scalar or a real estimate with physical units. This indicates the predicted imbalance phase, that is, the angular position of the imbalance mass block on the crankshaft. This means transposing the row vector so that it represents a column vector.

[0086] In this embodiment, in step S3, the mode probability distribution vector is... The sum of the predicted vectors The specific process of inputting the initial state into the digital twin is as follows: First, from the pattern probability distribution vector... Extract the state pattern category with the highest probability and match this pattern category with the loss prediction vector. The imbalance amplitude and imbalance phase are combined to form a physically interpretable initial state vector, expressed as:

[0087] ;

[0088] in, The initial state vector of the digital twin is represented by the discrete pattern. Continuous amplitude and continuous phase A new vector, formed by fusion, serves as the starting point for driving the digital twin's predictions. Subsequently, this initial state vector is input into a long short-term memory network model incorporating physical constraints. During training, the network model employs a composite loss function, which is a weighted sum of a data prediction error term and a physical law residual term. The physical law residual term is calculated by substituting the crankshaft angular acceleration predicted by the network into the diesel engine rigid body rotation dynamics equation. The expression for the composite loss function is:

[0089] ;

[0090] in, This represents the composite loss function, used for the overall optimization objective of training the digital twin LSTM model. A smaller value indicates more accurate predictions and a better alignment with physical laws. This represents the data fitting term, calculating the mean squared error (MSE) between the model's predicted values ​​and the actual observed values. It is a commonly used loss function in machine learning to ensure prediction accuracy. This indicates the calculation of the average coefficient. This indicates that the loss over N time steps is accumulated. The model represents the first time. The predicted output sequence for each time step includes multiple features such as vibration, imbalance, and carbon emissions. Indicates the first The real observation sequence at each time step is used to supervise the training and obtain the true values. Represents the physical constraint weight coefficient, a real number greater than 0, used to adjust the physical constraint. The greater the importance percentage in the composite loss function, the more strictly the model prediction results conform to the laws of physics.

[0091] This represents the physical residual term, quantifying the degree to which the model's predictions violate known physical laws. The inertia matrix of the crankshaft system is a constant matrix determined by the mass and geometric design parameters of components such as the diesel engine crankshaft and connecting rod, representing the rotational inertia of the system. The angular acceleration vector, representing the crankshaft rotation angle, is calculated by back-calculation from the vibration characteristic data predicted by the model, and represents the motion state of the system. The external torque vector is a torque estimated in real time from factors such as fuel injection parameters and load. It is the external torque applied to the system. The purpose of this formula is to force the trained model to satisfy this physical law as much as possible in its prediction results, thereby ensuring the rationality and credibility of the prediction.

[0092] During training, The forced prediction results satisfy the rigid body rotational dynamics equations It calculates the equations of motion for rigid bodies in physics. Ideally, the residual (unbalanced force / torque) is zero, ensuring that the prediction conforms to physical laws;

[0093] The specific process of performing forward-looking prediction computation is as follows: The digital twin starts from the initial state and performs recursive computation through the Long Short-Term Memory network. The formula for the physically constrained LSTM model is:

[0094] ;

[0095] in, Physically Constrained Long Short-Term Memory (LSTM) networks are the core computational units of the entire digital twin. Unlike standard LSTM, the training process of Physically Constrained Long Short-Term Memory networks is constrained by physical laws. Indicates in At any given moment, the state vector of the diesel engine crankshaft digital twin encapsulates the complete state of the system at that instant, including features such as vibration and imbalance. Indicates in At any given moment, the hidden state of the LSTM network represents the network's memory of previous historical sequence information. This represents the set of all learnable weight parameters within the LSTM network, including the weights and biases of the input gate, forget gate, and output gate. These parameters are fixed values ​​after the model is trained. This indicates that the LSTM network computes the following: The twin's predicted state vector at time 1. This indicates that the LSTM network has been updated. The hidden state at each time step will be passed to the next time step, predicting the vibration characteristic parameters, imbalance amplitude, and imbalance phase at multiple future time steps step by step. Simultaneously, based on the currently predicted vibration amplitude and imbalance amplitude, the carbon emission intensity value at the corresponding time step is output in real time through a pre-set carbon emission intensity calculation model. The formula for the carbon emission intensity calculation model is:

[0096] ;

[0097] in, This represents a carbon emission intensity calculation model, with mechanical states as input, including vibration amplitude. Imbalance Amplitude The output is carbon emissions. This indicates that the LSTM network computes the following: The twin's predicted state vector at time 1. Indicates in The predicted carbon emission intensity value at each time step is expressed in grams per kilowatt-hour; ultimately, the vibration characteristics of multiple consecutive time steps will be analyzed. Imbalanced state data and carbon emission intensity value The vibration features are encapsulated into independent prediction sequences in chronological order, where the vibration feature sequence contains the vibration spectrum features at each time step. The imbalance state sequence contains the amplitude and phase combination at each time step ( The carbon emission intensity sequence includes instantaneous carbon emission values ​​sorted by time. ), This indicates the total length of the predicted time series, i.e., how many future time steps' states are predicted.

[0098] In this embodiment, in step S4, the reward function of the multi-objective reinforcement learning agent is constructed using a dynamic weight allocation mechanism. Specifically, the process is as follows: First, based on the future multi-step carbon emission intensity values ​​extracted from the prediction sequence and their cumulative effects, and combined with the amplitude of the real-time vibration prediction vector, a dynamic priority factor is calculated. The formula for calculating the dynamic priority factor is:

[0099] ;

[0100] in, This represents a dynamic priority factor, a time-varying scalar whose value changes dynamically between 0 and 1. It is used to adjust the relative importance of carbon emission targets and fuel economy targets in the reward function in real time. The higher the value, the higher the priority of carbon reduction in the current state. The AI ​​will be more inclined to adopt carbon reduction strategies, which may slightly increase fuel consumption. This represents the Sigmoid function, which maps any real number calculated within the parentheses to the interval (0,1), such that the output is... These are probabilistic weight values ​​between 0 and 1. It is the normalized coefficient of carbon emissions. It is the normalized coefficient of vibration, due to the cumulative value of carbon emissions. and vibration amplitude These values ​​typically have different dimensions and orders of magnitude, so direct addition is meaningless. These two coefficients are used to adjust the two values ​​to a comparable order of magnitude, and together determine their magnitude. The value, Indicates from Add all the items together. This represents the discount factor, ranging from 0 to 1, indicating the varying degrees of importance the agent places on predictions made at different points in the future. This indicates that the importance of predictions further away from the current time decreases, causing the agent to focus more on recent carbon emission trends. Indicates the future number The predicted carbon emission intensity at each time step. The norm of the first-step vibration prediction vector is used to quantify the intensity of vibration at the current moment; the more intense the vibration, the larger the value. The dynamic priority factor, ranging from 0 to 1, is used to adjust the weight of carbon emission optimization in the reward function in real time. A larger dynamic priority factor indicates a higher priority for carbon emission optimization in the current state, and the corresponding weight of the carbon emission reward item in the reward function increases accordingly, while the weight of the fuel economy reward item decreases accordingly. Simultaneously, the reward function uses the square of the vibration amplitude as a penalty term and the square of the change in control command as a smoothness penalty term, together forming a multi-objective trade-off reward function, expressed as:

[0101] ;

[0102] in, The reward function is a scalar representing the final calculated value of the reward function. It indicates the immediate benefit or penalty gained by the reinforcement learning agent after performing an action in a specific state. The ultimate goal of the agent is to learn a policy that maximizes the total reward obtained from the environment. The cumulative sum, Represents a dynamic priority factor, when When the value is close to 1, the carbon emission reward item in the formula A higher weight indicates that the agent is more likely to take actions to reduce carbon emissions in the current state. Conversely, a lower weight indicates a lower weight. When the value is close to 0, the fuel consumption bonus item With greater weight given to [the technology / mechanics], the intelligent system will focus more on reducing fuel consumption. The change in carbon emissions is a scalar quantity, representing the change in the diesel engine's carbon emission intensity (g / kWh) after the agent's action, relative to before the action. The negative sign is placed in front because of the desire to reduce carbon emissions, therefore when When it is negative, This will turn into a positive value, thus giving the agent a positive reward; if emissions increase, a negative reward will be received. This represents the change in fuel consumption rate, a scalar quantity. It indicates the change in the diesel engine's instantaneous fuel consumption rate after the agent's action, relative to before the action. The negative sign indicates that, similar to the carbon emission rule, lower fuel consumption results in a positive reward, while higher fuel consumption results in a penalty. The vibration penalty coefficient is a preset, positive constant scalar that determines the proportion of the vibration penalty in the total reward. The weight or severity in the data. The larger the value, the more severe the penalty for exceeding the vibration limit. The norm of the predicted vibration vector for the first future step is a scalar value used to quantify the intensity of the vibration at the current moment; the more intense the vibration, the larger this value. This makes the effect of vibration intensity on reward non-linear; small vibrations result in lighter penalties, while large vibrations lead to very severe penalties, which meets the engineering requirements for vibration amplitude limiting. The smoothness penalty coefficient is a preset, positive constant scalar that determines the weight of the smoothness penalty term in the control instructions. The norm representing the change in control command. This indicates the control command generated this time. Compared with the previous instruction The difference between them The degree of drastic change in quantified control commands is penalized to prevent the agent from outputting jittery or abrupt control commands. Protective actuators such as fuel injectors and balancer motors ensure the smooth operation of the diesel engine. The more gradual the change in commands, the smaller the penalty.

[0103] When the agent learns a strategy, it does so under the premise of satisfying vibration constraints. , The vibration threshold is a preset, constant scalar value representing the maximum allowable vibration intensity of the system. This means that if the amplitude of the vibration prediction vector at the current moment is less than the preset vibration threshold, the system will maximize the cumulative reward. The method of learning the mapping strategy from state to action. , The state vector represents the state vector; ultimately, the agent outputs a raw action vector. The original action vector The dimension corresponds to the number of all adjustable parameters of the diesel engine, including fuel injection quantity, fuel injection timing, and balancer force;

[0104] The specific process of generating and issuing collaborative optimization control instruction sets is as follows: the multi-objective reinforcement learning agent calculates the original action vector based on its reward function; and decodes the action vector into a specific control instruction set. The instruction set includes a subset of injection parameter instructions and a subset of active balancer instructions. The injection parameter instruction subset includes main injection quantity instructions, pre-injection quantity instructions, and injection advance angle instructions. The active balancer instruction subset includes compensation force amplitude instructions and compensation force phase instructions. Before the instructions are sent to the controller, boundary verification is performed by a preset safety verification unit. The safety verification unit has built-in steady-state and transient safe operating boundary conditions for the diesel engine. It compares the control instructions with the preset safety boundaries. If the instructions exceed the safety boundaries, they are trimmed to the nearest boundary value. Finally, the control instruction set that has passed the verification is... The command is sent to the diesel engine controller via the bus for execution.

[0105] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0106] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0107] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0110] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0111] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method of optimizing carbon emissions based on a diesel engine, characterized by, The method comprises the following steps: Step S1, based on the multi-source sensor installed on the diesel engine crankshaft, synchronously collecting the original data of vibration acceleration, rotation speed pulse sequence and temperature signal; adopting adaptive weighted Kalman filtering algorithm to fuse the original data, and adopting wavelet packet transform algorithm to extract time-frequency domain features from the filtered signal, thereby generating a multi-dimensional feature vector for describing the dynamic balance state of the crankshaft; Step S2, inputting the multi-dimensional feature vector into a pre-constructed spatio-temporal graph convolution network model, analyzing the spatial topological relationship and time sequence change relationship of the multi-dimensional feature vector on the graph structure abstracted by the crankshaft connecting rod piston mechanism by using the spatio-temporal graph convolution network model, and outputting a mode probability distribution vector and an imbalance amount estimation vector for identifying the current dynamic balance state; Step S3, inputting the mode probability distribution vector and the imbalance amount estimation vector as initial states into the digital twin of the diesel engine crankshaft, the digital twin being a long short-term memory network model fused with physical constraints, for performing prediction operation and outputting a sequence of vibration features, imbalance state and predicted carbon emission intensity at future time steps; Step S4, inputting the sequence of predicted carbon emission intensity into a multi-objective reinforcement learning agent; the multi-objective reinforcement learning agent uses a reward function to perform decision operation and generates a collaborative optimization control instruction set; the collaborative optimization control instruction set is sent to the controller of the diesel engine, and the controller adjusts the injection parameters or the active electromagnetic balancer according to the instruction set, thereby realizing carbon emission optimization.

2. The method of claim 1, wherein: In the step S1, the multi-source sensor comprises a high-frequency vibration acceleration sensor arranged on the crankcase, an optical encoder installed at the end of the crankshaft, and a temperature sensor embedded at the main bearing hole; The original data comprises original vibration acceleration signals from the high-frequency vibration acceleration sensor, original rotation speed pulse sequences from the optical encoder, and original temperature signals from the temperature sensor; The step of fusing the original data by using the adaptive weighted Kalman filtering algorithm comprises: The original rotation speed pulse sequence is preprocessed, the instantaneous angular velocity is calculated by measuring the pulse interval, and the instantaneous rotation speed and instantaneous angular acceleration are further converted; According to the absolute value of the instantaneous angular acceleration, the weight of the process noise covariance matrix in the filtering algorithm is dynamically adjusted, and the larger the absolute value of the angular acceleration, the larger the value of the process noise covariance matrix; According to the absolute difference between the real-time temperature signal and the reference temperature value, the weight of the measurement noise covariance matrix is dynamically adjusted, and the larger the absolute difference in temperature, the larger the value of the measurement noise covariance matrix; The process noise covariance matrix and the measurement noise covariance matrix are adjusted dynamically, and the original vibration acceleration signal is filtered and calculated to output a high-fidelity vibration signal after noise reduction.

3. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 2, wherein: The step of extracting time-frequency domain features from the filtered signal by using the wavelet packet transform algorithm comprises: According to the instantaneous rotation speed calculated from the rotation speed pulse sequence, the core frequency band corresponding to the rotation base frequency and the second and third times of the rotation base frequency is determined in real time; The filtered high-fidelity vibration signal is subjected to wavelet packet decomposition, and energy values of each wavelet packet node in the core frequency band are calculated; The total energy of each core frequency band is calculated, and the proportion of the node energy in the total energy of the frequency band is calculated; The energy entropy value of each core frequency band is calculated based on the energy proportion.

4. The method of claim 3, wherein: In step S2, the construction of the graph structure abstracted from the crankshaft connecting rod piston mechanism and the feature mapping process are as follows: The multi-dimensional feature vector from step S1 is received, the main journal, crank pin, large end and small end of the connecting rod of the crankshaft are abstracted as nodes in the graph structure, the edges between the nodes are defined according to the actual physical connection relationship between the components, the edges represent the connection relationship, and the weight of the edge is calculated from the dynamic coupling coefficient between the components; finally, a graph structure including a node set, an edge set and a weighted adjacency matrix is formed, and each feature in the multi-dimensional feature vector is assigned to the corresponding node in the graph structure according to the physical meaning as the initial feature attribute of each node.

5. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 4, wherein: The process of analyzing the spatial topological relationship is as follows: For each node in the graph structure, aggregate the feature information of all adjacent nodes; during the aggregation operation, adjust the influence of the adjacent node features according to the physical connection strength weight defined in the adjacency matrix; at the same time, calculate the normalization coefficient according to the degree of the node itself and its neighbors, and normalize the aggregated features; finally, perform linear transformation on the normalized aggregated features and the learnable convolution kernel parameters, and pass them through a nonlinear activation function to obtain the updated feature vector of the node.

6. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 5, wherein: The process of analyzing the temporal change relationship is as follows: On the basis of the spatial topological relationship analysis, a gating mechanism is used to perform convolution operation along the time dimension; the gating mechanism adaptively controls the retention proportion of the historical state information and the current candidate state information through an update gate calculated by a time convolution kernel; the historical state information and the current candidate state information are extracted from the node feature sequence through different time convolution kernels; finally, the historical state and the current candidate state are fused by using the update gate to form the feature representation of the node at the next time step; The process of outputting the mode probability distribution vector and the imbalance amount estimation vector for identifying the current dynamic balance state is as follows: The global features of the graph extracted after the spatio-temporal graph convolution network processing are input into two parallel connected fully connected layers and output layers respectively; one of the output layers uses a Softmax function to output a mode probability distribution vector; the other output layer uses a linear regression function to output an imbalance amount estimation vector.

7. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 6, wherein: In step S3, the mode probability distribution vector and the imbalance amount estimation vector are input into the digital twin as the initial state. The state mode class with the highest probability is extracted from the mode probability distribution vector, and the mode class, the imbalance amplitude and the imbalance phase in the imbalance amount estimation vector are combined into a physically interpretable initial state vector. The initial state vector is input into a long short-term memory network model fused with physical constraints; the network model adopts a composite loss function in the training process, and the composite loss function is composed of a data prediction error term and a physical law residual term weighted; wherein the physical law residual term is calculated by substituting the network predicted crank angle acceleration into the rigid body rotation dynamics equation of the diesel engine.

8. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 7, wherein: The process of performing the forward-looking prediction operation is: the digital twin starts from the initial state, predicts the vibration characteristic parameters, imbalance amplitude and imbalance phase at multiple future time steps through recursive calculation of the long short-term memory network; at the same time, based on the currently predicted vibration amplitude and imbalance amplitude, the carbon emission intensity value at the corresponding time is output in real time through a preset carbon emission intensity calculation model; finally, the vibration characteristics, imbalance state data and carbon emission intensity value of continuous multiple time steps are encapsulated as independent prediction sequences in chronological order, wherein the vibration characteristic sequence contains the vibration frequency spectrum characteristics at each time, the imbalance state sequence contains the amplitude and phase combination at each time, and the carbon emission intensity sequence contains the time-ordered instantaneous carbon emission value.

9. A method of optimizing carbon emissions based on a diesel engine as claimed in claim 8, wherein: In step S4, the reward function of the multi-objective reinforcement learning agent is constructed using a dynamic weight allocation mechanism, and the process is: According to the future multi-step carbon emission intensity value extracted from the prediction sequence and the cumulative effect of the carbon emission intensity value, and combined with the amplitude of the real-time vibration prediction vector, a dynamic priority factor is calculated; the value of the dynamic priority factor is between 0 and 1, which is used to adjust the weight ratio of the carbon emission optimization target in the reward function in real time; at the same time, the reward function uses the square of the vibration amplitude as a penalty term, and uses the square of the control instruction change as a smoothness penalty term, to form a reward function that balances multiple objectives; When learning the strategy, the agent learns the mapping strategy from state to action by maximizing the cumulative reward under the premise of meeting the vibration constraint, i.e. the amplitude of the current time vibration prediction vector is less than the preset vibration threshold; The agent outputs an original action vector, and the dimension of the original action vector corresponds to the number of all adjustable parameters of the diesel engine.

10. The method of claim 9, wherein: The process of generating a collaborative optimization control instruction set and issuing the adjustment is: the multi-objective reinforcement learning agent calculates an original action vector according to the reward function; the action vector is decoded into a specific control instruction set, which includes an injection parameter instruction subset and an active balancer instruction subset; the injection parameter instruction subset includes main injection quantity instruction, pre-injection quantity instruction and injection advance angle instruction; the active balancer instruction subset includes compensation force amplitude instruction and compensation force phase instruction; before the instructions are issued to the controller, boundary checking is performed through a preset safety checking unit; the safety checking unit has built-in steady-state and transient safety operation boundary conditions of the diesel engine, and compares the control instructions with the preset safety boundary; if the control instructions exceed the safety boundary, the corresponding control instructions are clipped to the nearest boundary value; finally, the control instruction set that passes the verification is issued to the diesel engine controller through the bus for execution.