New energy truck carbon emission model adaptive optimization method and system

By dynamically adjusting the carbon emission model parameters through a hierarchical reinforcement learning reward function and a nonlinear correction factor, the problem of insufficient prediction accuracy of new energy trucks under complex operating conditions is solved, achieving adaptive optimization and high-precision prediction.

CN121504494APending Publication Date: 2026-02-10JIANGSU LINGHAO NETWORK TECH CO LTD

Patent Information

Application Number
CN202610036292.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing carbon emission models for new energy freight vehicles lack adaptive capabilities and struggle to achieve high-precision predictions under complex operating conditions. Existing reinforcement learning applications have failed to effectively optimize carbon emission models.

Method used

By acquiring operational data of new energy freight vehicles, a hierarchical reinforcement learning reward function is established, and nonlinear mapping and correction factors are used to dynamically adjust the parameters of the carbon emission model, forming an adaptive optimization method.

Benefits of technology

It achieves self-learning and self-optimization of carbon emission models under multiple and extreme operating conditions, improving prediction accuracy and stability, and adapting to emission characteristics under different vehicle, road and energy consumption conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504494A_ABST
    Figure CN121504494A_ABST
Patent Text Reader

Abstract

The invention discloses a new energy truck carbon emission model adaptive optimization method and system, and belongs to the technical field of model optimization, and the method specifically comprises the steps: obtaining the related data of a new energy truck in the operation process, generating a working condition feature vector corresponding to the parameters of a carbon emission model through nonlinear mapping, and carrying out the adaptive optimization of the carbon emission model based on the working condition feature vector, a layered reinforcement learning reward function is established, the reward function takes the minimum deviation between the predicted emission amount and the actual emission amount as a first-layer target and takes the energy consumption efficiency and the operation stability as a second-layer target, and in the reinforcement learning iteration process, a nonlinear correction factor is generated according to the feedback of the reward function; dynamically adjusting the coefficient of the carbon emission model by using the nonlinear correction factor, and performing adaptive optimization under multiple working conditions and extreme working conditions; according to the carbon emission model, continuous self-learning and self-optimization under multiple working conditions and extreme working conditions can be realized, and the adaptability, generalization and operation stability of the model are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of model optimization technology, specifically an adaptive optimization method and system for carbon emission models of new energy freight vehicles. Background Technology

[0002] With the widespread application of new energy trucks in urban logistics and long-distance transportation, their carbon emissions have gradually attracted attention. Although new energy vehicles do not rely on traditional fossil fuels during operation, their energy conversion process still indirectly leads to carbon emissions, especially under different loads, complex road conditions, and variable climate conditions, where actual emissions often deviate from theoretical calculations. Most existing carbon emission models are built based on fixed parameters and lack the ability to adapt to dynamic operating characteristics, resulting in insufficient prediction accuracy under complex operating conditions.

[0003] Existing research mainly uses empirical parameters or linear correction methods to calibrate models. These methods rely on preset conditions and are difficult to adapt to changes in vehicle operating conditions over long periods. In recent years, reinforcement learning has been gradually introduced into the field of vehicle operation optimization. Reinforcement learning can obtain feedback through continuous interaction with the environment and form adaptive strategies during the iteration process. However, most existing reinforcement learning applications focus on vehicle path planning and energy consumption optimization, and there is no mature method specifically designed for the adaptive optimization of carbon emission models for new energy freight vehicles. Therefore, how to establish a carbon emission model based on reinforcement learning technology that can dynamically adjust parameters and adapt to the characteristics of multiple operating conditions has become an urgent technical problem to be solved. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes an adaptive optimization method and system for carbon emission models of new energy freight vehicles.

[0005] To achieve the above objectives, the present invention provides the following technical solution: An adaptive optimization method for carbon emission models of new energy freight vehicles includes: The relevant data of new energy trucks during operation are obtained and generated as operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics. Based on the operating condition feature vector, a hierarchical reinforcement learning reward function is established. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability. During the reinforcement learning iteration process, a nonlinear correction factor is generated based on the feedback of the reward function. The coefficients of the carbon emission model are dynamically adjusted using the nonlinear correction factor, and adaptive optimization is performed under multiple operating conditions and extreme operating conditions.

[0006] Specifically, the step of acquiring relevant data of new energy trucks during operation and generating operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping includes: Acquire relevant data during the operation of new energy freight vehicles, including vehicle load, energy consumption parameters, and road condition characteristics; The acquired data is processed by time-series slicing to form a data sequence corresponding to the running cycle; The data sequence is subjected to multi-scale feature decomposition to obtain a multi-dimensional initial feature vector characterizing the vehicle's dynamic operation. The initial feature vector is subjected to a nonlinear mapping transformation to establish the correspondence between the vehicle operating state and the carbon emission model parameters; The results obtained through the nonlinear mapping are combined to form a working condition feature vector corresponding to the carbon emission model parameters.

[0007] Specifically, based on the aforementioned working condition feature vector, a hierarchical reinforcement learning reward function is established, including: Based on the operating condition feature vector, a first-level evaluation index related to the predicted emissions is generated, and the initial reward parameter corresponding to the first-level evaluation index is determined as the first-level reward parameter. Based on the operating condition feature vector, a second-layer evaluation index related to energy consumption efficiency and operational stability is generated, and the initial reward parameter corresponding to the second-layer evaluation index is determined as the second-layer reward parameter. A hierarchical relationship is established between the first-level evaluation indicators and the second-level evaluation indicators to form a multi-level structure of the reward function; Based on the data feedback during operation, the first-layer reward parameters and the second-layer reward parameters are dynamically weighted, and the weights are adjusted in real time during the iteration process. The weighted first-layer reward parameters are combined with the second-layer reward parameters to form a hierarchical reinforcement learning reward function.

[0008] Specifically, a hierarchical correlation is established between the first-level evaluation indicators and the second-level evaluation indicators to form a multi-level structure of the reward function, including: The first-level evaluation index and the second-level evaluation index are respectively feature-encoded to form an index representation that can be compared. A hierarchical constraint relationship is introduced between the indicator representations, so that the update process of the first-level evaluation indicator serves as the input condition for the second-level evaluation indicator. Based on the hierarchical constraint relationship, an indicator interaction path is constructed to transmit feedback information between different levels; The indicator interaction path is embedded into the reward function to generate a hierarchical reward function with a multi-level structure.

[0009] Specifically, the step of dynamically assigning weights to the first-layer reward parameters and the second-layer reward parameters based on data feedback during operation, and adjusting the weights in real time during the iteration process, includes: Acquire real-time feedback data during the operation, and match the real-time feedback data with the first-level reward parameters and the second-level reward parameters accordingly; Based on the matching results, initial weights are set for the first-level reward parameters and the second-level reward parameters; During the iterative update process, the initial weights are adaptively corrected based on the dynamic changes in the relevant data; The corrected weight values ​​are redistributed to the first-level reward parameters and the second-level reward parameters to form a dynamic weight allocation result that can be adjusted in real time.

[0010] Specifically, during the reinforcement learning iteration process, a nonlinear correction factor is generated based on the reward function feedback. This nonlinear correction factor is then used to dynamically adjust the coefficients of the carbon emission model, enabling adaptive optimization under multiple operating conditions and extreme conditions. This includes: During the reinforcement learning iteration process, feedback signals generated by the hierarchical reinforcement learning reward function are received, and the feedback signals are parsed and processed. Based on the analytical results, a correction vector is constructed to represent the adjustment requirements for carbon emission model parameters under different operating conditions; The correction vector is subjected to a nonlinear transformation to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions. The nonlinear correction factor is applied to the load factor, road condition factor and energy consumption factor in the carbon emission model to form a dynamically adjusted parameter set. During multiple iterations and updates, the parameter set is used to adaptively optimize the carbon emission model under various operating conditions and extreme conditions.

[0011] Specifically, the construction of the correction vector based on the analytical results, which characterizes the adjustment requirements for carbon emission model parameters under different operating conditions, includes: The analyzed feedback signals are classified according to load, energy consumption, and road condition to form a multi-dimensional feedback set; Feature factors related to the carbon emission model parameters are extracted from the multidimensional feedback set, and the feature factors are normalized. An adjustment vector based on normalized feature factors is constructed for the working condition dimension to characterize the independent adjustment requirements of carbon emission model parameters under different working conditions. The adjustment vectors of the aforementioned working condition dimensions are combined to generate a unified correction vector.

[0012] Specifically, the correction vector is subjected to a nonlinear transformation to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions, including: The correction vector is deconstructed into load dimension components, energy consumption dimension components, and road condition dimension components. Nonlinear transformations are performed on all dimensional components to generate intermediate correction results that reflect different operational characteristics; The intermediate correction results are interactively combined in a multi-dimensional space to generate a composite correction vector with correlation. Based on the composite correction vector, the mapping result is extracted to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions.

[0013] Specifically, the nonlinear correction factor is applied to the load factor, road condition factor, and energy consumption factor in the carbon emission model to form a dynamically adjusted parameter set, including: The nonlinear correction factor is mapped and matched with the load coefficient, road condition coefficient and energy consumption coefficient in the carbon emission model to determine the correction direction of each parameter. Based on the mapping and matching, the parameters are adjusted at the component level to generate load adjustment values, road condition adjustment values ​​and energy consumption adjustment values; The adjusted values ​​and correction directions of each parameter are iteratively synthesized with the corresponding original parameters to form dynamically updated parameter groups. The dynamically updated parameters are grouped and combined into a parameter set.

[0014] An adaptive optimization system for a carbon emission model of a new energy freight truck, used to implement the aforementioned adaptive optimization method for a carbon emission model of a new energy freight truck, includes: a data acquisition module, a reward function establishment module, and an optimization module; The data acquisition module is used to acquire relevant data of new energy trucks during operation, and generate operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics. The reward function establishment module establishes a hierarchical reinforcement learning reward function based on the operating condition feature vector. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability. The optimization module is used to generate a nonlinear correction factor based on the reward function feedback during the reinforcement learning iteration process, and to dynamically adjust the coefficients of the carbon emission model using the nonlinear correction factor to perform adaptive optimization under multiple operating conditions and extreme operating conditions.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention proposes an adaptive optimization method and system for carbon emission models of new energy freight vehicles. With the core objective of minimizing the deviation between predicted and actual emissions, it utilizes a hierarchical reward mechanism and a nonlinear correction factor to achieve real-time updates and cross-environmental migration of model parameters. Through this method, the carbon emission model can continuously learn and optimize under multiple operating conditions and extreme conditions, avoiding the problems of fixed parameters and difficulty in adapting to changes in operating characteristics in traditional models. It achieves high-precision prediction of emission characteristics for different vehicles, different roads, and different energy consumption states, significantly improving the model's adaptability, generalization ability, and operational stability. Attached Figure Description

[0016] Figure 1 A flowchart of an adaptive optimization method for carbon emission models of new energy freight vehicles provided by the present invention; Figure 2 The dynamic weight allocation flowchart provided by this invention; Figure 3 The adaptive optimization flowchart provided by this invention; Figure 4 This invention provides an adaptive optimization system architecture diagram for a carbon emission model of a new energy freight vehicle. Detailed Implementation

[0017] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. In addition, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0020] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0021] Example 1 Please see Figures 1-3 The present invention provides an embodiment of an adaptive optimization method for a carbon emission model of new energy freight vehicles, comprising the following specific steps: Step S1: Obtain relevant data of the new energy truck during operation, and generate a working condition feature vector corresponding to the carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics.

[0022] The specific steps of step S1 are as follows: Step S101: Obtain relevant data of the new energy truck during operation, including vehicle load, energy consumption parameters and road condition characteristics.

[0023] In this embodiment, multi-source data on the operation of new energy freight vehicles are systematically acquired and analyzed. Specifically, during vehicle operation, raw data such as vehicle load, drive energy consumption, motor output power, gradient, and road adhesion coefficient are collected in real time through on-board sensing units, energy management modules, and road information collection devices. Subsequently, the data is synchronized and denoised according to the time series, outliers are removed using data filtering algorithms, and load change rate, energy consumption growth rate, and road condition complexity factor are calculated through a feature extraction module. It should be noted that load data reflects the mass distribution characteristics of the vehicle under different operating conditions, energy consumption parameters characterize energy output and utilization efficiency, and road condition characteristics reflect the dynamic environmental constraints of vehicle operation.

[0024] Step S102: Perform time-series slicing on the acquired relevant data to form a data sequence corresponding to the running cycle.

[0025] In this embodiment, continuously collected operational data is sliced ​​according to time dimension and vehicle state changes to construct a data sequence that reflects the periodic operational characteristics of the vehicle. Specifically, firstly, based on key state events of vehicle operation as segmentation nodes, including the stages of starting, accelerating, constant speed, deceleration, and stopping, the collected load, energy consumption, and road condition data are divided into intervals. Secondly, within each operational interval, the data is aggregated using a time window based on the sampling frequency, so that each slice covers the complete energy output and environmental response characteristics. Then, a time synchronization mechanism is used to align different types of sensor data on the same time axis, ensuring that load changes, energy consumption fluctuations, and road condition characteristics have a consistent time reference. It should be noted that through this time-series slicing process, the originally continuous and redundant multidimensional data is transformed into a data sequence with stable periodic boundaries.

[0026] Step S103: Perform multi-scale feature decomposition on the data sequence to obtain a multi-dimensional initial feature vector characterizing the vehicle's dynamic operation.

[0027] In this embodiment, multi-scale feature decomposition is performed on the time-sliced ​​data sequence to extract dynamic features at different time granularities. Specifically, firstly, the load, energy consumption, and road condition data sequences are divided into multi-scale hierarchical layers, with short-term fluctuations, periodic trends, and long-term stable patterns serving as feature layers at different scales. Secondly, in each feature layer, feature extraction methods based on rate of change, fluctuation amplitude, and trend gradient are used to analyze local peaks, energy consumption growth trends, and road condition change intensity, capturing the vehicle's operating behavior at different time scales. Then, the feature results extracted at each level are vectorized and encoded so that each dimension corresponds to a specific operating condition response factor, such as load change rate, energy consumption conversion efficiency, and road surface complexity. It should be noted that the principle of this multi-scale feature decomposition is to express high-frequency disturbances and low-frequency trends simultaneously through the joint separation of time scale and feature space, thereby forming a multi-dimensional initial feature vector that can comprehensively characterize the vehicle's operating state.

[0028] Step S104: Perform a nonlinear mapping transformation on the initial feature vector to establish the correspondence between the vehicle operating state and the carbon emission model parameters.

[0029] In this embodiment, a dynamic correspondence between vehicle operating status and carbon emission model parameters is constructed through nonlinear mapping transformation. Specifically, firstly, the multidimensional initial feature vectors obtained in step S103 are reorganized to ensure that each dimension remains independently expressed in the data space. Then, a multi-layer nonlinear mapping structure is introduced to convert the high-order interaction information of the feature vectors into response modes in the parameter space step by step. The coupling changes of load features and energy consumption features are used to reflect the correction trend of load coefficients in the model, and the nonlinear combination of road condition features and energy consumption gradients is used to characterize the dynamic adjustment relationship of road condition correction coefficients. Then, through feature reconstruction operations, the outputs of different mapping layers are fused on a unified parameter dimension to form a parameter mapping matrix with hierarchical association. It should be noted that the principle of this nonlinear mapping transformation is to transform the originally dispersed operating status quantities into a structured representation that can drive the update of carbon emission model parameters through multi-level feature interaction and nonlinear association learning, thereby establishing a correspondence between vehicle operating status and model parameters.

[0030] Step S105: Combine the results obtained through the nonlinear mapping to form a working condition feature vector corresponding to the carbon emission model parameters.

[0031] In this embodiment, the multi-layer output results obtained through nonlinear mapping transformation are structurally combined to generate operating condition feature vectors that correspond one-to-one with the carbon emission model parameters. Specifically, firstly, the parameter response results output by each mapping layer are dimensionally identified and classified, and the results belonging to load correction, road condition correction, and energy consumption correction are divided into independent feature subsets. Secondly, in each subset, the mapping results under different time scales and spatial distributions are weighted and integrated using a feature weighting fusion method, retaining key operating features and reducing redundant correlations. Then, the integrated results are vectorized and encoded so that each vector dimension corresponds to a dynamic adjustment factor of a specific model parameter. Finally, all vectors are reconstructed according to the model parameter sequence to form operating condition feature vectors with a unified format. It should be noted that this step, through the hierarchical integration and vectorized reconstruction of the nonlinear mapping results, realizes an explicit mapping relationship between operating state features and model parameters, enabling the carbon emission model to achieve parameter-level response after inputting the operating condition feature vector.

[0032] Step S2: Based on the operating condition feature vector, establish a hierarchical reinforcement learning reward function. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability.

[0033] The specific steps of step S2 are as follows: like Figure 2As shown, step S201: Based on the operating condition feature vector, generate a first-level evaluation index related to the predicted emissions, and determine the initial reward parameter corresponding to the first-level evaluation index as the first-level reward parameter.

[0034] In this embodiment, a first-layer evaluation index related to the predicted emissions is constructed based on the aforementioned operating condition feature vector, and its corresponding initial reward parameters are determined. Specifically, firstly, the load features, energy consumption features, and road condition features in the operating condition feature vector are correlated, and the core components most relevant to emission changes are identified through a feature importance ranking method. Then, key variable groups representing emission trends are extracted from each component, and a first-layer evaluation index is generated to characterize the predicted emission error. This index can reflect the difference between the model output and the measured emissions in different operating cycles. Next, a reference emission range is established based on the vehicle's historical operating samples, and this range is used as the standard range of the evaluation index to calculate the weight benchmark of the initial reward parameters. It should be noted that the initial reward parameters are set based on the degree of deviation of the evaluation index, and the self-correction guidance of the model in the initial stage of reinforcement learning is achieved through weight allocation.

[0035] Step S202: Based on the operating condition feature vector, generate a second-level evaluation index related to energy consumption efficiency and operational stability, and determine the initial reward parameter corresponding to the second-level evaluation index as the second-level reward parameter.

[0036] In this embodiment, a second-layer evaluation index related to energy consumption efficiency and operational stability is generated based on the operating condition feature vector, and corresponding initial reward parameters are determined. Specifically, firstly, feature components representing energy usage intensity, drive power fluctuation, and road condition change frequency are extracted from the operating condition feature vector, and the energy consumption change pattern of the vehicle in different operating stages is identified through time correlation analysis. Secondly, the energy consumption features are jointly modeled with dynamic stability factors such as vehicle speed fluctuation and torque response rate to form a composite index system that simultaneously describes energy utilization efficiency and operational stability. Then, the second-layer evaluation index is calculated based on the energy efficiency change rate and stability deviation in the composite system, so that it can comprehensively reflect the energy conversion balance of the vehicle within a unit operating mileage. It should be noted that the initial reward parameters are determined based on the numerical distribution of the second-layer evaluation index. A joint constraint on energy efficiency and stability is established through interval division and weight assignment, so that the reinforcement learning model can simultaneously consider energy consumption optimization and operational smoothness control during the initial training stage.

[0037] Step S203: Establish a hierarchical relationship between the first-level evaluation indicators and the second-level evaluation indicators to form a multi-level structure of the reward function.

[0038] The specific steps of step S203 are as follows: Step S2031: Perform feature encoding on the first-layer evaluation index and the second-layer evaluation index respectively to form an index representation that can be compared.

[0039] In this embodiment, the first-level and second-level evaluation indicators are feature-encoded to form indicator representations for hierarchical comparison and correlation analysis. Specifically, firstly, the data structures of the two types of evaluation indicators are unified, converting the first-level indicators related to emission deviation and the second-level indicators related to energy efficiency and stability into vectorized expressions. Subsequently, normalization and scale balancing operations are performed within each indicator to eliminate differences in data magnitude and make them comparable. Next, a feature grouping coding strategy is introduced to classify and encode indicators with different physical attributes according to their functional types. For example, emission error features are labeled as performance factor vectors, and energy consumption stability features are labeled as behavioral factor vectors. Finally, the feature encoding results are integrated into the same representation space through a multi-dimensional index mapping method to ensure that different indicators can be semantically matched and hierarchically mapped at the encoding level. It should be noted that the principle of this step is to convert multi-source heterogeneous evaluation indicators into a unified feature representation structure, enabling the reinforcement learning model to perform interactive calculations and dynamic correlations of different level targets in subsequent stages.

[0040] Step S2032: Introduce hierarchical constraints between the indicator representations, so that the update process of the first-level evaluation indicators serves as the input condition for the second-level evaluation indicators.

[0041] In this embodiment, a hierarchical constraint relationship is established between the first-level evaluation indicators and the second-level evaluation indicators, so that the dynamic update result of the former serves as the input condition for the latter, thereby forming a dependent hierarchical structure. Specifically, firstly, the output change trend of the first-level evaluation indicators is monitored and modeled, and key change points in its time series are extracted as state transition signals. Secondly, this signal is correlated and mapped with the input parameters of the second-level evaluation indicators, so that energy efficiency and operational stability can perceive the adjustment direction of emission indicators in real time during calculation. Then, the dependency path between the two levels of indicators is defined through a hierarchical constraint matrix, which restricts the update process of the second-level indicators to be executed based on the feedback state of the first-level indicators, ensuring that the information transmission between different target levels has directionality and causal consistency. Finally, this constraint relationship is embedded in the reward calculation framework of reinforcement learning, so that the model automatically follows the hierarchical coupling rules between indicators during the training phase. It should be noted that the principle of this step is to incorporate emission accuracy optimization and energy efficiency stability control into a unified decision chain by establishing a hierarchical constraint mechanism, so that different levels of evaluation objectives have transferable constraint logic.

[0042] Step S2033: Construct an indicator interaction path based on the hierarchical constraint relationship to transmit feedback information between different levels.

[0043] In this embodiment, an indicator interaction path is constructed based on the established hierarchical constraint relationship to achieve feedback transmission and information sharing between different evaluation levels. Specifically, firstly, the output characteristics of the first-level evaluation indicators are state-encoded, converting the emission deviation change trend into a propagable feedback signal. Secondly, a signal receiving channel is set at the input end of the second-level evaluation indicators, and the feedback signal from the first level is embedded into the energy efficiency and stability calculation node through state indexing, enabling it to perceive changes in the upper level when updating parameters. Then, the transmission path and priority of the feedback signal are determined according to the hierarchical constraint matrix, so that the feedback of high-sensitivity indicators has priority in affecting key energy consumption variables, while the feedback of low-sensitivity indicators affects secondary parameters through a delayed update mechanism. Finally, the transmission mechanism of the interaction path is incorporated into the reinforcement learning loop, so that different levels form a continuous information feedback channel during the iteration process. It should be noted that this step, by constructing a multi-layered feedback transmission structure, enables dynamic linkage between carbon emissions and energy consumption targets, thereby establishing an inherent coupling relationship between hierarchical indicators at the algorithm level.

[0044] Step S2034: Embed the indicator interaction path into the reward function to generate a hierarchical reward function with a multi-level structure.

[0045] In this embodiment, the aforementioned established indicator interaction paths are embedded into the reward function structure to generate a hierarchical reward function with a multi-level structure. Specifically, firstly, the interaction paths of different levels of indicators are functionally encapsulated, and emission accuracy feedback and energy consumption stability feedback are defined as independent reward sub-channels. Secondly, a corresponding weight adjustment factor is set in each sub-channel to enable the feedback signals of different levels to have adaptive response capabilities when input into the reward function. Then, the reward signals of each sub-channel are dynamically integrated through a hierarchical aggregation mechanism, so that the immediate feedback of the first-level indicators drives the model correction first, while the stability constraint of the second-level indicators plays a balancing role in subsequent iterations. Finally, by introducing hierarchical control variables into the reinforcement learning main loop, the reward function can redistribute weights according to the state relationship between each level of indicators during the iteration process, forming a structured hierarchical reward system. It should be noted that this step, by embedding the indicator interaction paths into the reward function, realizes the construction of a composite optimization objective driven by multi-level indicators, giving the reward function hierarchical interpretability and dynamic adjustment characteristics.

[0046] Step S204: Based on the data feedback during the operation, dynamically allocate weights to the first-layer reward parameters and the second-layer reward parameters, and adjust the weights in real time during the iteration process.

[0047] The specific steps of step S204 are as follows: Step S2041: Obtain real-time feedback data during the operation, and match the real-time feedback data with the first-level reward parameters and the second-level reward parameters.

[0048] In this embodiment, a corresponding matching relationship is established between operational feedback data and hierarchical reward parameters to achieve dynamic reward association driven by real-time data. Specifically, firstly, feedback data of the new energy truck during operation is acquired in real time through the vehicle information acquisition module. The feedback data includes key operational parameters such as instantaneous energy consumption, emission estimation deviation, driving speed fluctuation, and load change. Secondly, the feedback data is preprocessed, and time synchronization and noise suppression algorithms are used to ensure the integrity and continuity of the data. Then, the preprocessed feedback data is input into the reward parameter matching module. Based on the data type and feature dimension, values ​​related to emission deviation are mapped to the first-level reward parameters, and values ​​related to energy efficiency and operational stability are mapped to the second-level reward parameters. Finally, an index mapping table is used to bind the feedback data to the reward parameters at each level, so that the operational status at each moment can be dynamically associated with the hierarchical reward function. It should be noted that this step provides data support for the hierarchical parameter update during the reinforcement learning process by constructing a matching channel between real-time feedback data and reward parameters.

[0049] Step S2042: Based on the matching results, set the initial weights for the first-level reward parameters and the second-level reward parameters.

[0050] In this embodiment, based on the aforementioned feedback matching results, the initial weight allocation of the first-layer reward parameters and the second-layer reward parameters is determined, establishing the initial optimization direction for reinforcement learning. Specifically, firstly, the distribution characteristics of the matched feedback data are statistically modeled to extract key features of emission deviation change rate and energy consumption fluctuation amplitude. Secondly, based on the frequency and sensitivity of each feature in historical operating cycles, a weight allocation matrix is ​​constructed to measure the relative influence of the first-layer and second-layer rewards on the overall optimization objective. Then, the reward parameters are initialized hierarchically according to the weight matrix output results, where the first-layer parameters, which are directly related to emission accuracy, are assigned higher weights, while the second-layer parameters, which are related to energy consumption stability, are dynamically proportionally set according to the current vehicle operating conditions. Finally, through a one-time iterative initialization operation, the initial weights are written into the reinforcement learning training framework as a reference baseline for subsequent dynamic adjustment of rewards.

[0051] Step S2043: During the iterative update process, the initial weights are adaptively corrected based on the dynamic changes of the relevant data.

[0052] In this embodiment, during the iterative update process of reinforcement learning, the initial weights are adaptively corrected based on the dynamic changes in the running data to achieve continuous balance adjustment of the hierarchical reward parameters. Specifically, firstly, the real-time data stream during the operation is windowed and key change indicators are extracted, including the fluctuation trend of emission deviation, the rate of change of energy consumption efficiency, and the time gradient of road condition complexity. Secondly, the above-mentioned change indicators are input into the weight correction module, and the response lag and sensitivity differences of each layer of reward parameters under the current operating condition are identified through the comparison of change rate and deviation accumulation analysis. Then, the weight correction coefficient is dynamically adjusted according to the identification results, so that the emission-related weights are increased when the deviation increases, while the energy consumption stability weights are preferentially strengthened when energy fluctuations intensify, thereby achieving real-time adaptive balance of the multi-objective optimization process. Finally, the corrected weights are written back to the reward function layer through an iterative feedback mechanism, so that the next round of reinforcement learning calculations automatically performs decision optimization based on the updated weight structure. It should be noted that this step introduces a weight self-adjustment mechanism in the iterative loop, enabling the reward function to continuously follow the changes in the running data to achieve proportional self-adjustment.

[0053] Step S2044: The corrected weight values ​​are redistributed to the first-level reward parameters and the second-level reward parameters to form a dynamic weight allocation result that can be adjusted in real time.

[0054] In this embodiment, the adaptively corrected weight values ​​are redistributed to construct a weight allocation result that can be dynamically updated according to changes in operating status. Specifically, firstly, the corrected weight vector output in step S2043 is normalized to ensure that the total weight of different levels of reward parameters remains constant, so as to avoid proportional drift during multi-objective optimization. Secondly, the normalized weight vector is hierarchically mapped, and the part related to emission accuracy is redistributed to the first-level reward parameters, and the part related to energy consumption balance and stability is redistributed to the second-level reward parameters, with the update cycle set according to the rate of change of operating data. Then, the weights are refreshed in real time through the dynamic allocation module. When the vehicle operating status changes abruptly or reaches a long-term steady state, the distribution ratio of each layer of weights on the time axis is automatically adjusted, thereby forming a dynamic weight structure that continuously changes with environmental feedback. Finally, this structure is written into the calculation node of the reward function, so that reinforcement learning can perform decision calculation with the latest weight configuration in each iteration. It should be noted that this step, by introducing a dynamic weight redistribution mechanism, enables different levels of reward parameters to maintain adaptive coordination in the time and operating condition dimensions.

[0055] Step S205: Combine the weighted first-layer reward parameters with the second-layer reward parameters to form a hierarchical reinforcement learning reward function.

[0056] In this embodiment, the dynamically weighted reward parameters of the first and second layers are structurally combined to generate a reinforcement learning reward function with hierarchical feedback characteristics. Specifically, firstly, the weighted reward parameters of the first and second layers are data aligned to ensure a unified index mapping relationship between them in the time series and feature space. Secondly, emission accuracy parameters and energy consumption stability parameters are jointly integrated using a hierarchical aggregation algorithm. During the aggregation process, a weight fusion strategy is set according to the degree of inter-layer dependency, so that the immediacy constraint of the first-layer objective and the stability constraint of the second-layer objective form a complementary relationship in the function structure. Then, a cross-adjustment term is introduced into the aggregation result to reflect the dynamic coupling effect between the two layers of parameters, thereby achieving self-balancing adjustment between emissions and energy consumption. Finally, the fused result is functionally encapsulated to construct a reward function with decomposable gradients and hierarchical dependency characteristics, which is used to guide the parameter update process of the reinforcement learning model. It should be noted that this step, by introducing hierarchical aggregation and interactive adjustment mechanisms into the reward calculation structure, enables hierarchical reinforcement learning to simultaneously express the weight association of different optimization objectives in a single reward function.

[0057] Step S3: During the reinforcement learning iteration process, a nonlinear correction factor is generated based on the feedback of the reward function. The coefficients of the carbon emission model are dynamically adjusted using the nonlinear correction factor, and adaptive optimization is performed under multiple operating conditions and extreme operating conditions.

[0058] like Figure 3 As shown, the specific steps of step S3 are as follows: Step S301: During the reinforcement learning iteration process, receive the feedback signal generated by the hierarchical reinforcement learning reward function, and parse and process the feedback signal.

[0059] In this embodiment, feedback signals output by the hierarchical reinforcement learning reward function are received during the iteration process of reinforcement learning, and are analyzed and structured to achieve parameterized driving of model updates. Specifically, firstly, real-time feedback data output by the reward function is collected in each iteration cycle. The data includes multi-dimensional signals such as emission prediction deviation, energy consumption change trend, and operational stability status. Secondly, the feedback signals are grouped according to their source hierarchy, so that the emission accuracy signal from the first layer and the energy efficiency balance signal from the second layer remain structurally independent and identifiable. Then, feature extraction and attribution analysis are performed on the feedback from each layer through the signal analysis module to extract key feature variables that affect the convergence direction of the model, such as feedback intensity, response delay, and weight sensitivity. Finally, a feedback analysis matrix is ​​generated based on the analysis results to provide data basis for the construction of subsequent correction vectors. It should be noted that this step, by receiving and analyzing the hierarchical reward feedback signals, makes the implicit reward structure in the reinforcement learning training process explicit, enabling the model to achieve adaptive parameter updates based on the hierarchical features of the feedback.

[0060] Step S302: Construct a correction vector based on the analysis results to represent the adjustment requirements for carbon emission model parameters under different operating conditions.

[0061] The specific steps of step S302 are as follows: Step S3021: Classify the parsed feedback signals according to load, energy consumption and road condition categories to form a multi-dimensional feedback set.

[0062] In this embodiment, the parsed feedback signals are categorized and organized to form a feedback set that reflects the multidimensional operating status of the vehicle. Specifically, firstly, the feedback signals from the hierarchical reward function are structurally decomposed, and parameter response values ​​related to vehicle operating characteristics are extracted. Pairing corrections are then performed based on the synchronicity of the feedback signals on the time axis. Secondly, based on the physical meaning of the vehicle operating characteristics, the signals are classified into three feature channels: load, energy consumption, and road condition. Load signals reflect the impact of vehicle load changes on the emission model, energy consumption signals characterize the dynamic changes in energy output efficiency, and road condition signals represent the degree of disturbance to the emission model response caused by road complexity. Then, within each category, temporal aggregation and variance normalization operations are used to ensure consistency in intensity and time scale for signals of the same category. Finally, the three processed signals are combined into a multidimensional feedback set, providing a basic data structure for subsequent feature extraction and correction vector construction. It should be noted that this step, by classifying complex feedback signals according to physical operating condition attributes, allows for the hierarchical representation of the influence of different operating factors.

[0063] Step S3022: Extract the feature factors related to the carbon emission model parameters from the multidimensional feedback set, and normalize the feature factors.

[0064] In this embodiment, feature factors that are directly or indirectly related to the carbon emission model parameters are identified and extracted from the multidimensional feedback set, and then normalized to establish a unified feature expression space. Specifically, firstly, feature sensitivity analysis is performed on the multidimensional feedback set obtained in step S3021, and key feature variables that have a significant impact on the changes in model parameters in the dimensions of load, energy consumption, and road conditions are identified using correlation screening and principal component decomposition. Secondly, attribute remapping is performed on the screened feature variables to convert nonlinear distribution features into measurable feature factors, ensuring that the mapping direction of features in different dimensions is consistent in the parameter space. Then, statistical normalization and interval recalibration techniques are used to process each feature factor to ensure consistency in numerical range, dimensions, and magnitude of change, thereby eliminating magnitude deviations between different types of signals. Finally, a feature factor index matrix is ​​established, and the normalized feature factors are classified and stored according to their corresponding model parameters, providing a standardized input basis for the construction of correction vectors. It should be noted that this step, through the extraction and normalization of feature factors, realizes a computable mapping from feedback data to the model parameter space.

[0065] Step S3023: Construct an adjustment vector for the operating condition dimension based on normalized feature factors to characterize the independent adjustment requirements of carbon emission model parameters under different operating conditions.

[0066] In this embodiment, adjustment vectors for operating conditions are constructed based on normalized feature factors to achieve independent correction of carbon emission model parameters under different operating environments. Specifically, firstly, feature mapping indices are established for load, energy consumption, and road condition dimensions according to the correlation between feature factors and model parameters, and the response direction of the feature factors in each dimension is paired with the corresponding parameter. Secondly, the feature contribution is calculated in each operating condition dimension through weighted aggregation, so that different feature factors form a measurable adjustment trend vector in the same dimension. Then, the adjustment trends of each dimension are orthogonalized to eliminate mutual interference between different operating condition features, so that each dimension vector can independently reflect its correction requirements for model parameters. Finally, the orthogonalized results are combined in dimensional order to generate a complete operating condition dimension adjustment vector, which can be used for fine-tuning of parameters in the subsequent correction factor construction stage. It should be noted that this step, by structuring the normalized feature factors into operating condition fractal vectors, realizes a direct mapping between operating state and model parameter correction, enabling the model to have decomposable parameter adaptive expression capabilities under multiple operating conditions.

[0067] Step S3024: Combine the adjustment vectors of the working condition dimension to generate a unified correction vector.

[0068] In this embodiment, the adjustment vectors of each operating condition dimension are structurally combined to generate a unified correction vector that can comprehensively characterize the global correction requirements of carbon emission model parameters. Specifically, firstly, the adjustment vectors of the three dimensions of load, energy consumption, and road condition are subjected to data alignment and scale standardization to ensure that the vectors of different dimensions have a consistent weight distribution in the numerical space. Secondly, dynamic fusion coefficients are set according to the importance and frequency of change of the operating conditions to weight and aggregate the adjustment vectors of each dimension, highlighting the characteristic dimensions that have a greater impact on the model parameters under the current operating conditions. Then, the weighted results are spatially reconstructed through a vector fusion mapping algorithm, transforming the adjustment trends of multiple dimensions into a set of globally consistent parameter correction expressions. Finally, the reconstructed vectors are rectified and indexed to ensure that each component maintains a one-to-one correspondence with the corresponding parameter position in the carbon emission model, thereby forming a unified correction vector. It should be noted that this step, through the fusion and spatial reconstruction of multi-dimensional vectors, transforms the scattered operating condition adjustment results into an overall operable correction instruction, enabling the model to respond to parameter shifts under different operating conditions in a globally consistent manner.

[0069] Step S303: Perform a nonlinear transformation on the correction vector to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions.

[0070] The specific steps of step S303 are as follows: Step S3031: Deconstruct the correction vector into load dimension components, energy consumption dimension components, and road condition dimension components.

[0071] In this embodiment, the key to step S3031 lies in the dimensional deconstruction of the unified correction vector, re-separating the implicit multi-condition features into independent components of three dimensions: load, energy consumption, and road condition. Specifically, firstly, according to the construction rules of the unified correction vector, the source features of each component are labeled and tracked, and the condition category to which each numerical element belongs is determined through parameter mapping index. Secondly, corresponding component sets are extracted for the three dimensions of load, energy consumption, and road condition, and the independent contribution rate of each dimension is calculated based on the feature correlation matrix to ensure that the components after deconstruction have non-overlapping influence domains in numerical terms. Then, the following steps are taken: The hierarchical decomposition algorithm removes the coupled signal portion, so that the load component retains only the influence of vehicle mass change, the energy consumption component reflects the energy conversion characteristics of the drive system, and the road condition component describes the degree of interference of road environment complexity on the model response. Finally, the three types of components are re-encoded into independent vector groups, providing a clear dimensional input structure for subsequent nonlinear transformation and parameter correction operations. It should be noted that this step establishes a mapping channel from global correction trends to dimensional influencing factors by deconstructing and regrouping the unified correction vector, enabling the model to identify the differentiated adjustment needs of carbon emission parameters under different operating conditions with finer granularity.

[0072] Step S3032: Perform nonlinear transformations on all dimensional components to form intermediate correction results that reflect different operating characteristics.

[0073] In this embodiment, the key to step S3032 lies in applying nonlinear transformations to the load component, energy consumption component, and road condition component obtained through dimensional deconstruction, respectively, to form intermediate correction results that can reflect different operating characteristics. Specifically, firstly, based on the physical properties and dynamic change characteristics of each dimensional component, nonlinear mapping models with different response curvatures are selected. The transformation of the load component focuses on reconstructing the relationship between mass fluctuation and power output, the transformation of the energy consumption component focuses on the nonlinear response analysis of energy utilization rate and instantaneous power, and the transformation of the road condition component focuses on modeling the complex correlation between road disturbance intensity and energy consumption fluctuation frequency. Secondly, within each dimension... The component values ​​are adaptively divided into intervals, and a multi-order response function is introduced to amplify or compress the local variation interval, thereby enhancing the model's sensitivity to nonlinear gradient changes. Then, the transformed results of each dimension are subjected to stability testing and dynamic smoothing to ensure that the numerical distributions of different dimensions are compatible. Finally, the processed results are defined as an intermediate correction result set, providing basic input for subsequent cross-mapping and global correction factor generation. It should be noted that this step achieves differentiated expression of features of each dimension through targeted nonlinear transformation, so that the influence of different operating factors can be quantified in an independent and adjustable form.

[0074] Step S3033: The intermediate correction results are interactively combined in a multidimensional space to generate a composite correction vector with correlation.

[0075] In this embodiment, the aforementioned intermediate correction results are interactively combined in a multi-dimensional space to generate a correlated composite correction vector, which is used to characterize the synergistic influence relationship between different operational features. Specifically, firstly, a unified spatial coordinate system is established for the intermediate correction results of the three dimensions of load, energy consumption, and road conditions, and the data of each dimension are mapped to a comparable coordinate domain using a normalized projection method. Secondly, a multi-dimensional cross-mapping algorithm is used to calculate the coupling strength between different dimensional components, and the influence of load changes on energy consumption response and the modulation relationship of road condition fluctuations on emission correction are represented in the form of a feature interaction matrix. Then, in the matrix... Based on the array, a hierarchical aggregation mechanism is introduced. Through nonlinear superposition and weighted fusion, the multidimensional interactive results are synthesized into a vector set with global coupling characteristics. Finally, the fused results are directionally normalized and the amplitude is recalibrated to ensure that the composite correction vector can maintain the independence of each dimension in the numerical space while reflecting its inherent correlation. It should be noted that this step, by introducing interactive combination and hierarchical aggregation process in multidimensional space, enables the independent correction effects of different working conditions to be synergistically expressed in a unified structure, thereby forming a composite correction vector that can dynamically reflect the comprehensive impact of load, energy consumption and road conditions.

[0076] Step S3034: Extract the mapping result based on the composite correction vector to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions.

[0077] In this embodiment, a nonlinear correction factor reflecting the characteristics of complex operating conditions is generated by extracting the mapping result of a composite correction vector. Specifically, firstly, spatial feature analysis is performed on the composite correction vector, and a hierarchical projection method is used to extract high-order feature components reflecting load coupling, energy consumption dynamics, and road condition disturbances to separate the nonlinear response modes between operating conditions. Secondly, the extracted high-order features are reconstructed using a nonlinear mapping model to transform the changing trends of different feature components into continuously differentiable response curves, so as to establish a correspondence between the correction factor and the model deviation at the parameter level. Then, an adaptive adjustment model is introduced into the reconstruction result. The model dynamically adjusts the curvature of the mapping function based on the complexity of the operating conditions and the density of the data distribution, enabling the generated correction factors to remain sensitive to changing environments. Finally, the adjusted nonlinear output values ​​are encapsulated into a set of correction factors, giving it the composite characteristics of multidimensional input and multi-layered response, and providing it in parameterized form for subsequent model updates. It should be noted that this step, through the combination of high-dimensional feature mapping and nonlinear response reconstruction, realizes the transformation from composite correction vectors to operable correction factors, enabling the carbon emission model to achieve dynamic parameter adjustment under complex operating conditions, and providing a nonlinear adjustment core for adaptive optimization driven by reinforcement learning.

[0078] Step S304: Apply the nonlinear correction factor to the load coefficient, road condition coefficient, and energy consumption coefficient in the carbon emission model to form a dynamically adjusted parameter set.

[0079] The specific steps of step S304 are as follows: Step S3041: Map and match the nonlinear correction factor with the load coefficient, road condition coefficient and energy consumption coefficient in the carbon emission model to determine the correction direction of each parameter.

[0080] In this embodiment, the key to step S3041 is to map and match the nonlinear correction factor with the load coefficient, road condition coefficient, and energy consumption coefficient in the carbon emission model to determine the correction direction of each parameter in the dynamic optimization process. Specifically, firstly, the core parameter set of the carbon emission model is structured and analyzed to establish a parameter attribute table, distinguishing different parameter dimensions that reflect vehicle load characteristics, energy consumption conversion efficiency, and road environment impact. Secondly, factor components related to the characteristics of each parameter dimension are identified in the nonlinear correction factor set, and their optimal mapping relationship is determined by feature similarity calculation. Then, the gradient direction of each factor's effect on the corresponding parameter is calculated based on the mapping relationship, and its positive or negative sign is used as the direction instruction for parameter adjustment. The correction direction of the load coefficient is mainly affected by the load change rate, the correction direction of the energy consumption coefficient is determined based on the energy utilization deviation, and the correction direction of the road condition coefficient is determined by the fluctuation trend of road complexity. Finally, the correction direction of each parameter is encoded into a parameter update vector to provide guiding information for subsequent dynamic iterative adjustments. It should be noted that this step achieves the directional driving of multidimensional factors on different types of parameters through the mapping and matching between nonlinear correction factors and model parameters. This enables the carbon emission model to have the ability to identify the parameter response direction for specific working conditions under the reinforcement learning framework, thus providing a logical starting point for the adaptive adjustment mechanism.

[0081] Step S3042: Based on the mapping and matching, adjust each parameter at the component level to generate load adjustment value, road condition adjustment value and energy consumption adjustment value.

[0082] In this embodiment, the core of step S3042 lies in adjusting the parameters in the carbon emission model at the component level based on mapping and matching, generating corresponding load adjustment values, road condition adjustment values, and energy consumption adjustment values. Specifically, firstly, according to the parameter correction direction determined in step S3041, the nonlinear correction factor is vectorized and paired with the original values ​​of each parameter to obtain the changing trend of each parameter in the multidimensional space. Secondly, the component offset is calculated in each parameter dimension according to the correction direction and rate of change, and a sensitivity coefficient is introduced for different parameter components to limit the proportional relationship between their adjustment magnitude and model stability. The system then integrates the adjustment offsets hierarchically using a component-weighted aggregation algorithm, ensuring that the load, road condition, and energy consumption parameters remain independently progressive rather than canceling each other out during the update process. Finally, the integrated adjustment results are converted into numerical parameter outputs, defined as load adjustment values, road condition adjustment values, and energy consumption adjustment values, respectively. It should be noted that this step achieves refined correction of model parameters through a component-level adjustment strategy, enabling the nonlinear correction factor to have differentiated effects in different dimensions, thereby ensuring that the carbon emission model can maintain the continuity of dynamic response and the stability of parameter adjustment under multiple operating conditions.

[0083] Step S3043: Iteratively synthesize the adjustment values ​​of each parameter, the correction direction of each parameter, and the corresponding original parameters to form a dynamically updated parameter group.

[0084] In this embodiment, the key to step S3043 is to iteratively synthesize the aforementioned parameter adjustment values ​​and their corresponding correction directions with the original parameters of the carbon emission model to form dynamically updated parameter groups. Specifically, firstly, parameter update channels are established based on the correction directions of each parameter, and the three types of parameters—load, energy consumption, and road conditions—are directionally superimposed according to their corresponding adjustment values. Secondly, iterative fusion operations are performed on each parameter channel, and the update step size is determined based on the correction magnitude and historical convergence trend in each round of calculation to prevent excessive parameter drift or adjustment lag. Then, the update process of each parameter is smoothed in the time dimension through weighted accumulation, so that short-term fluctuation signals are gradually absorbed while long-term trends are preserved. Finally, the parameter set that converges after multiple iterations is defined as a dynamic parameter group, where each group of parameters corresponds to a specific operating condition and time period. It should be noted that this step, by introducing a component-level iteration and directional synthesis mechanism, enables the parameters of the carbon emission model to maintain a real-time coupling relationship with the operating state during continuous learning, thereby realizing the self-evolution and steady-state update basis of the model in a dynamic environment.

[0085] Step S3044: Group the dynamically updated parameters into a parameter set.

[0086] In this embodiment, the dynamically updated parameter groups obtained through multiple rounds of iterative synthesis are structurally combined to form a parameter set that can support the adaptive adjustment of the carbon emission model. Specifically, firstly, each parameter group is indexed and organized according to its dimension and time series, establishing a one-to-one mapping relationship between the load parameter group, energy consumption parameter group, and road condition parameter group in the spatial dimension. Secondly, the update frequency and stability coefficient are calculated within each parameter group to determine its priority and weight ratio in the global set, thereby achieving hierarchical integration of multi-dimensional parameters. Then, the coupling relationship between different parameter groups is corrected through a set reconstruction algorithm to keep the short-term adjustment term and the long-term correction term balanced in the set structure, so as to avoid the parameters canceling each other out or being repeatedly adjusted under different operating conditions. Finally, the integrated parameter groups at each layer are combined in a matrix form to form a complete parameter set, providing an input benchmark for the next round of reinforcement learning updates of the model. It should be noted that this step, through the structured aggregation and hierarchical reconstruction of parameter groups, realizes the transition from single-dimensional dynamic correction to multi-dimensional parameter coupling, enabling the carbon emission model to have a sustainable iterative parameter management framework.

[0087] Step S305: During multiple iterations and updates, the carbon emission model under multiple operating conditions and extreme operating conditions is adaptively optimized using the parameter set.

[0088] In this embodiment, the core of step S305 lies in adaptively optimizing the carbon emission model under multiple operating conditions and extreme operating conditions based on the aforementioned parameter set during the multi-round iterative update process of reinforcement learning. Specifically, firstly, real-time operating condition feature data of the vehicle is collected at different operating stages, and dynamically matched with the corresponding parameter group in the parameter set to achieve synchronous mapping between the model input and the current environment. Secondly, according to the matching result, the corresponding parameter hierarchical structure in the set is called, so that load, energy consumption, and road condition parameters participate in the model calculation in a multi-dimensional space in the form of adaptive weights, thereby establishing a specific emission estimation path for different operating conditions. Then, in the reinforcement learning iteration... In the loop update, deviation detection is performed between the model output and the actual feedback. By comparing the difference between the predicted emissions and the measured values, a re-optimization mechanism within the parameter set is triggered, enabling the parameters to redistribute weights and achieve local convergence in different operating conditions. Finally, through multiple iterations, a stable parameter evolution trajectory is formed, allowing the model to maintain dynamic response capabilities under complex or abrupt operating conditions. This step, by integrating reinforcement learning with the adaptive parameter set, constructs a sustainable convergent multi-operating condition optimization system, enabling the carbon emission model to maintain prediction accuracy and adaptability under different load intensities, energy consumption patterns, and road complexity conditions, providing dynamic closed-loop support at the algorithm level for full-cycle emission optimization.

[0089] It should be noted that the emission deviation threshold, energy efficiency change threshold, operational stability threshold, and reward weight adjustment trigger threshold involved in this invention are all used to judge and adjust the vehicle's operating status and model output results. Their function is to distinguish the trigger conditions and adjustment intensity of model parameter adjustment under different operating conditions. The above-mentioned thresholds are not preset as fixed constants, but are determined comprehensively based on factors such as the distribution characteristics of historical operating data of new energy freight vehicles, vehicle load status, energy consumption level, and road condition complexity. In the specific implementation process, the variation range of each evaluation index under normal operating conditions is obtained by statistical analysis of historical operating data or sliding time window calculation, and this range is used as the reference benchmark for threshold setting. When the model prediction result or operating feedback data exceeds the reference range, it is determined that the corresponding threshold condition is triggered, thereby guiding the reinforcement learning reward weight or model parameters into the adjustment state.

[0090] Example 2 Please see Figure 4 Another embodiment of the present invention provides: an adaptive optimization system for carbon emission models of new energy freight vehicles, comprising: a data acquisition module, a reward function establishment module, and an optimization module; The data acquisition module is used to acquire relevant data of new energy trucks during operation, and generate operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics. The reward function establishment module establishes a hierarchical reinforcement learning reward function based on the operating condition feature vector. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability. The optimization module is used to generate a nonlinear correction factor based on the reward function feedback during the reinforcement learning iteration process, and to dynamically adjust the coefficients of the carbon emission model using the nonlinear correction factor to perform adaptive optimization under multiple operating conditions and extreme operating conditions.

[0091] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.

[0092] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive optimization method for a carbon emission model of new energy freight vehicles, characterized in that, include: The relevant data of new energy trucks during operation are obtained and generated as operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics. Based on the operating condition feature vector, a hierarchical reinforcement learning reward function is established. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability. During the reinforcement learning iteration process, a nonlinear correction factor is generated based on the feedback of the reward function. The coefficients of the carbon emission model are dynamically adjusted using the nonlinear correction factor, and adaptive optimization is performed under multiple operating conditions and extreme operating conditions.

2. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 1, characterized in that, The process of acquiring relevant data from the operation of new energy trucks and generating operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping includes: Acquire relevant data during the operation of new energy freight vehicles, including vehicle load, energy consumption parameters, and road condition characteristics; The acquired data is processed by time-series slicing to form a data sequence corresponding to the running cycle; The data sequence is subjected to multi-scale feature decomposition to obtain a multi-dimensional initial feature vector characterizing the vehicle's dynamic operation. The initial feature vector is subjected to a nonlinear mapping transformation to establish the correspondence between the vehicle operating state and the carbon emission model parameters; The results obtained through the nonlinear mapping are combined to form a working condition feature vector corresponding to the carbon emission model parameters.

3. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 2, characterized in that, Based on the aforementioned working condition feature vector, a hierarchical reinforcement learning reward function is established, including: Based on the operating condition feature vector, a first-level evaluation index related to the predicted emissions is generated, and the initial reward parameter corresponding to the first-level evaluation index is determined as the first-level reward parameter. Based on the operating condition feature vector, a second-layer evaluation index related to energy consumption efficiency and operational stability is generated, and the initial reward parameter corresponding to the second-layer evaluation index is determined as the second-layer reward parameter. A hierarchical relationship is established between the first-level evaluation indicators and the second-level evaluation indicators to form a multi-level structure of the reward function; Based on the data feedback during operation, the first-layer reward parameters and the second-layer reward parameters are dynamically weighted, and the weights are adjusted in real time during the iteration process. The weighted first-layer reward parameters are combined with the second-layer reward parameters to form a hierarchical reinforcement learning reward function.

4. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 3, characterized in that, A hierarchical relationship is established between the first-level evaluation indicators and the second-level evaluation indicators to form a multi-level structure of the reward function, including: The first-level evaluation index and the second-level evaluation index are respectively feature-encoded to form an index representation that can be compared. A hierarchical constraint relationship is introduced between the indicator representations, so that the update process of the first-level evaluation indicator serves as the input condition for the second-level evaluation indicator. Based on the hierarchical constraint relationship, an indicator interaction path is constructed to transmit feedback information between different levels; The indicator interaction path is embedded into the reward function to generate a hierarchical reward function with a multi-level structure.

5. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 4, characterized in that, The step of dynamically assigning weights to the first-layer reward parameters and the second-layer reward parameters based on data feedback during operation, and adjusting the weights in real time during the iteration process, includes: Acquire real-time feedback data during the operation, and match the real-time feedback data with the first-level reward parameters and the second-level reward parameters accordingly; Based on the matching results, initial weights are set for the first-level reward parameters and the second-level reward parameters; During the iterative update process, the initial weights are adaptively corrected based on the dynamic changes in the relevant data; The corrected weight values ​​are redistributed to the first-level reward parameters and the second-level reward parameters to form a dynamic weight allocation result that can be adjusted in real time.

6. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 5, characterized in that, During the reinforcement learning iteration process, a nonlinear correction factor is generated based on the reward function feedback. This nonlinear correction factor is then used to dynamically adjust the coefficients of the carbon emission model, enabling adaptive optimization under multiple operating conditions and extreme conditions. This includes: During the reinforcement learning iteration process, feedback signals generated by the hierarchical reinforcement learning reward function are received, and the feedback signals are parsed and processed. Based on the analytical results, a correction vector is constructed to represent the adjustment requirements for carbon emission model parameters under different operating conditions; The correction vector is subjected to a nonlinear transformation to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions. The nonlinear correction factor is applied to the load factor, road condition factor and energy consumption factor in the carbon emission model to form a dynamically adjusted parameter set. During multiple iterations and updates, the parameter set is used to adaptively optimize the carbon emission model under various operating conditions and extreme conditions.

7. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 6, characterized in that, The correction vector constructed based on the analytical results, representing the adjustment requirements for carbon emission model parameters under different operating conditions, includes: The analyzed feedback signals are classified according to load, energy consumption, and road condition to form a multi-dimensional feedback set; Feature factors related to the carbon emission model parameters are extracted from the multidimensional feedback set, and the feature factors are normalized. An adjustment vector based on normalized feature factors is constructed for the working condition dimension to characterize the independent adjustment requirements of carbon emission model parameters under different working conditions. The adjustment vectors of the aforementioned working condition dimensions are combined to generate a unified correction vector.

8. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 7, characterized in that, The correction vector is subjected to a nonlinear transformation to generate a nonlinear correction factor that reflects the characteristics of complex operating conditions, including: The correction vector is deconstructed into load dimension components, energy consumption dimension components, and road condition dimension components. Nonlinear transformations are performed on all dimensional components to generate intermediate correction results that reflect different operational characteristics; The intermediate correction results are interactively combined in a multi-dimensional space to generate a composite correction vector with correlation. Based on the composite correction vector, the mapping result is extracted to generate a nonlinear correction factor that can reflect the characteristics of complex working conditions.

9. The adaptive optimization method for carbon emission model of new energy freight vehicles as described in claim 8, characterized in that, The nonlinear correction factor is applied to the load factor, road condition factor, and energy consumption factor in the carbon emission model to form a dynamically adjusted parameter set, including: The nonlinear correction factor is mapped and matched with the load coefficient, road condition coefficient and energy consumption coefficient in the carbon emission model to determine the correction direction of each parameter. Based on the mapping and matching, the parameters are adjusted at the component level to generate load adjustment values, road condition adjustment values ​​and energy consumption adjustment values; The adjusted values ​​and correction directions of each parameter are iteratively synthesized with the corresponding original parameters to form dynamically updated parameter groups. The dynamically updated parameters are grouped and combined into a parameter set.

10. An adaptive optimization system for a carbon emission model of a new energy freight vehicle, used to implement the adaptive optimization method for a carbon emission model of a new energy freight vehicle as described in any one of claims 1-9, characterized in that, include: Data acquisition module, reward function establishment module, and optimization module; The data acquisition module is used to acquire relevant data of new energy trucks during operation, and generate operating condition feature vectors corresponding to carbon emission model parameters through nonlinear mapping. The relevant data includes vehicle load, energy consumption parameters and road condition characteristics. The reward function establishment module establishes a hierarchical reinforcement learning reward function based on the operating condition feature vector. The reward function includes a first-level objective of minimizing the deviation between predicted emissions and actual emissions, and a second-level objective of energy efficiency and operational stability. The optimization module is used to generate a nonlinear correction factor based on the reward function feedback during the reinforcement learning iteration process, and to dynamically adjust the coefficients of the carbon emission model using the nonlinear correction factor to perform adaptive optimization under multiple operating conditions and extreme operating conditions.

Citation Information

Patent Citations

  • Micro-grid dynamic cooperative scheduling method based on multi-mode reinforcement learning

    CN120601500A

  • Layered optimization regulation and control method for electricity-hydrogen coupling in multi-energy complementary system

    CN120745439A

  • Engineering vehicle carbon emission monitoring system and method

    CN120847331A

  • Network attack active defense strategy optimization method based on deep reinforcement learning

    CN120934876A

Cited By

  • Carbon emission optimization method and system considering safety and stability of energy system

    CN121860659A

  • A carbon emission optimization method and system considering safety and stability of energy system

    CN121860659B