Power domain controller configuration method, apparatus, device, and medium
By constructing a physical information dynamic network and adaptive control strategy patches, the disconnect between development and use of traditional power domain controllers is solved, enabling performance optimization and energy efficiency improvement throughout the vehicle's life cycle, coping with extreme operating conditions, and reducing risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONKWO COM
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional power domain controllers suffer from a disconnect between virtual simulation and physical vehicles, as well as between early development and later use. This makes it difficult to efficiently close the loop and provide feedback on policy issues, results in complex software code maintenance, makes it difficult to adapt to environmental changes, and leads to incomplete coverage of real vehicle road tests and difficulty in discovering potential risks.
By constructing and continuously evolving a dynamic network of physical information, the complex coupling relationships between variables in the dynamic system are captured in real time. Based on the minimization of the global network energy function, actuator control instructions are generated, adaptive control policy patches are synthesized in reverse, and meta-reinforcement learning is used to optimize the network construction rules to generate adaptive control policies.
It achieves improved adaptability of the power domain controller, optimizes vehicle performance and energy efficiency throughout its entire life cycle, effectively copes with extreme operating conditions, and reduces potential risks.
Smart Images

Figure CN122194710B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of electric control systems for new energy vehicles, and in particular to power domain controller configuration methods, devices, equipment and media. Background Technology
[0002] In the traditional development process of powertrain domain controllers, there is a common technical limitation of disconnect between virtual simulation and physical vehicles, and between early development and later use. On the one hand, policy problems discovered during the simulation verification phase cannot be efficiently and accurately fed back to the controller code in a closed loop, still relying on manual conversion and implementation, which is not only inefficient but also prone to introducing new errors. At the same time, to adapt to different powertrain configurations (such as EV, PHEV, etc.), multiple independent software codebases are often required, resulting in dispersed development resources, low reusability, and a long expansion cycle for new configurations.
[0003] On the other hand, once a vehicle rolls off the production line, its control strategies and parameters are usually fixed, making it difficult to adaptively adjust based on actual battery degradation, diverse user driving habits, and changing environmental conditions. This limits the vehicle's performance and energy efficiency optimization throughout its entire lifecycle. Furthermore, traditional real-vehicle road tests, limited by cost and safety, cannot fully cover various extreme and long-tail conditions, making it difficult to completely identify and eliminate potential risks before mass production. Summary of the Invention
[0004] This application provides a power domain controller configuration method, apparatus, device, and medium, which can shorten the development cycle of power domain control and achieve dynamic, global, and multi-objective optimal resource allocation.
[0005] On one hand, embodiments of this application provide a power domain controller configuration method, the method including: Acquire real-time multi-source data streams of the powertrain system and vehicle status events, wherein the real-time multi-source data streams include data from different sources collected in real time from the vehicle powertrain system and related sensors; Based on the real-time multi-source data stream and the vehicle state events, a physical information dynamic network is constructed and continuously evolved, wherein the nodes of the physical information dynamic network represent physical state variables, and the edges of the physical information dynamic network represent the dynamic coupling relationship between variables. Based on the physical information dynamic network, a global network energy function is constructed and minimized, wherein the low-energy state of the global network energy function corresponds to the coordinated working mode of the system; Based on the requirement of minimizing the global network energy function, the minimum information perturbation is calculated and applied to generate control commands for the actuator; Based on the long-term evolution topology of the physical information dynamic network, virtual physical constraints are synthesized in reverse, and adaptive control strategy patches are generated.
[0006] Optionally, constructing and minimizing a global network energy function based on the physical information dynamic network includes: Based on the degree of contradiction between the internal node states of the physical information dynamic network, a network tension term is constructed; Based on the disorder degree of the network structure of the physical information dynamic network, a network entropy term is constructed. Based on the constraints of the vehicle-level performance targets on the network state of the physical information dynamic network, a target term is constructed; The network tension term, network entropy term, and target term are linearly combined using time-varying coupling coefficients to form the global network energy function.
[0007] Optionally, the step of calculating and applying a minimum information perturbation based on the requirement of minimizing the global network energy function to generate control commands for the actuator includes: Based on the gradient information of the global network energy function under the current network state, and the constraints formed by the system physical equations, a Lagrange optimization problem is constructed. Solving the optimization problem yields the minimum change vector required to influence the state of each node in order to guide the energy function to decrease. The state components in the minimum change vector that correspond to those that the actuator can directly influence are converted into actual control commands.
[0008] Optionally, after generating control instructions for the actuator, the method further includes: Based on the physical information dynamic network, predict the evolution of the network state at the next moment after the execution of the control command; Acquire the actual sensor data at the next moment and calculate the difference between the evolution result and the network state induced by the actual data; If the difference exceeds a preset threshold, it is determined that the control command is inconsistent with the network causal model, triggering an immediate online correction of the weights of the relevant connection edges in the physical information dynamic network, and recalculating the control command based on the corrected network.
[0009] Optionally, the step of reverse-synthesizing virtual physical constraints based on the long-term evolution topology of the physical information dynamic network and generating adaptive control policy patches includes: The monitoring of the physical information dynamic network is conducted to determine whether a stable anomalous subgraph pattern emerges during its long-term evolution. The anomalous subgraph pattern includes a specific set of nodes and corresponding strong connection edges, and the energy level is consistently higher than the network baseline. The abnormal subgraph pattern is isolated from the physical information dynamic network and used as input to an independent constraint synthesis problem; The inversion algorithm is used to solve the following: under the assumption that there are certain additional, unmodeled physical constraints in the system, the global network energy function will form a local minimum at the currently observed anomalous subgraph structure; The additional physical constraints obtained from the solution are formally described as a single or set of virtual physical laws, which serve as the basis for the synthetic control strategy patch.
[0010] Optionally, generating the adaptive control strategy patch includes: The formally described virtual physical laws are compiled into a lightweight policy patching module. The policy patching module includes a state correction function and a rule injection unit. The state correction function is used to preprocess or correct the sensor data input to the traditional control algorithm according to the virtual physical laws. The rule injection unit is used to temporarily add extra terms corresponding to the virtual physical laws to the evolution equation of the physical information dynamic network. The policy patch module is received and securely loaded via the vehicle OTA channel.
[0011] Optionally, after synthesizing virtual physical constraints in reverse based on the long-term evolution topology of the physical information dynamic network and generating adaptive control policy patches, the method further includes: Obtain a metacognitive dataset within a historical time period, which includes all network evolution trajectories, energy function history, control command sequences, and final performance indicators; When cloud or vehicle-side computing power allows, initiate a meta-reinforcement learning process, which is used to optimize the hyperparameters of the construction rules of the physical information dynamic network, the dynamic coupling coefficient generation strategy of the global network energy function, and the heuristic rules of the reverse synthesis virtual physical constraints. The optimization strategies obtained from meta-learning are encapsulated into metacognitive update packages and sent to vehicles via in-vehicle OTA.
[0012] On the other hand, embodiments of this application provide a power domain controller configuration device, the device comprising: The acquisition module is used to acquire real-time multi-source data streams of the power system and vehicle status events. The real-time multi-source data streams include data from different sources collected in real time from the vehicle power system and related sensors. The first construction module is used to construct and continuously evolve a physical information dynamic network based on the real-time multi-source data stream and the vehicle state event, wherein the nodes of the physical information dynamic network represent physical state variables, and the edges of the physical information dynamic network represent the dynamic coupling relationship between variables. The second construction module is used to construct and minimize a global network energy function based on the physical information dynamic network, wherein the low energy state of the global network energy function corresponds to the coordinated working mode of the system. The first generation module is used to calculate and apply minimum information perturbation based on the requirement of minimizing the global network energy function, and generate control instructions for the actuator; The second generation module is used to reverse synthesize virtual physical constraints based on the long-term evolution topology of the physical information dynamic network, and generate adaptive control strategy patches.
[0013] In another aspect, embodiments of this application provide an electronic device, the device comprising: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the power domain controller configuration method as described in the first aspect.
[0014] In another aspect, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the power domain controller configuration method as described in the first aspect.
[0015] The power domain controller configuration method, apparatus, device, and medium of this application can capture the complex dynamic coupling relationships between various variables of the power system in real time by constructing and continuously evolving a physical information dynamic network. Based on minimizing the global network energy function, it generates refined actuator control commands, effectively solving the problems of traditional controller strategies being rigid and unable to adapt to environmental changes. Furthermore, by analyzing the long-term evolution topology of the network, reverse-synthesizing virtual physical constraints, and generating adaptive strategy patches, this method can compensate for the shortcomings of existing models, improve vehicle performance and energy efficiency optimization throughout its entire lifecycle, and effectively cope with extreme operating conditions that are difficult to cover in traditional road tests, reducing potential risks. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a power domain controller configuration method provided in an embodiment of this application; Figure 2 This is a structural block diagram of a power domain controller configuration device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0018] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0019] To address the problems of existing technologies, this application provides a power domain controller configuration method, apparatus, device, and medium. In this application, a dynamically evolving physical information network can be constructed to capture the complex dynamic coupling relationships between various variables in the power system in real time. Based on minimizing the global network energy function, refined actuator control commands are generated, effectively solving the problems of traditional controller strategies being rigid and unable to adapt to environmental changes. Furthermore, by analyzing the long-term evolution topology of the network, reverse-synthesizing virtual physical constraints, and generating adaptive strategy patches, this method can compensate for the shortcomings of existing models, improve vehicle performance and energy efficiency optimization throughout its entire lifecycle, and effectively cope with extreme operating conditions that are difficult to cover in traditional road tests, reducing potential risks.
[0020] The configuration method of the power domain controller provided in the embodiments of this application will be introduced first below.
[0021] Figure 1 A flowchart illustrating a power domain controller configuration method according to an embodiment of this application is shown. Figure 1 As shown, the configuration method for the power domain controller may include S101-S105: S101 acquires real-time multi-source data streams of the powertrain system and vehicle status events.
[0022] In this application embodiment, the power system generally refers to the collection of all components in a vehicle responsible for generating and transmitting power, such as the engine, motor, battery, transmission, drive shaft, etc. Its operating status directly affects the vehicle's performance, energy consumption, and emissions.
[0023] Real-time multi-source data stream: This data stream refers to data collected in real time from various sources by the vehicle's powertrain system and related sensors, such as engine speed, torque, temperature, battery voltage, current, state of charge (SOC), vehicle speed, and accelerator pedal opening. This data is acquired in the form of a continuous stream to reflect the system's immediate state.
[0024] Vehicle status events: These events refer to state changes or external conditions that occur during vehicle operation and have specific meanings or trigger specific logic, such as driving mode switching, fault diagnostic code triggering, charging status changes, road type recognition, and weather condition changes. These events typically require the controller to respond accordingly.
[0025] S102 constructs and continuously evolves a physical information dynamic network based on real-time multi-source data streams and vehicle status events.
[0026] In this embodiment, a physical information dynamic network is used: this network is an abstract model to characterize the physical state variables and their interactions in a dynamic system. The network represents physical state variables through nodes and dynamic coupling relationships between variables through edges. This network can continuously evolve based on real-time data, reflecting the dynamic characteristics of the system's behavior.
[0027] Specifically, a node in a physical information dynamic network represents a specific physical state variable in a dynamic system, such as the temperature, pressure, or rotational speed of a component, or the state of charge or health of a battery.
[0028] Edge: In a physical information dynamic network, an edge represents the interaction or influence relationship between different physical state variables. This relationship can be causal, associative, or constraining, and its strength and properties can change dynamically over time.
[0029] S103, based on the physical information dynamic network, construct and minimize a global network energy function.
[0030] Specifically, the global network energy function is a scalar function used to quantify the overall "energy" level of a physical information dynamic network at a given moment. The low-energy state of this function typically corresponds to a coordinated, stable, and efficient operating mode of the dynamic system. Specifically, the low-energy state of the global network energy function corresponds to the coordinated operating mode of the system. By minimizing this energy function, the system can be guided to evolve towards the desired operating state.
[0031] S104, based on the requirement of minimizing the global network energy function, calculates and applies minimum information perturbation to generate control commands for the actuator.
[0032] Minimal information disturbance: This disturbance refers to the minimal intervention required on the system in terms of information or energy consumption to guide the global network energy function towards a lower energy state. This disturbance is typically achieved by adjusting the control input of the actuator.
[0033] S105: Based on the long-term evolution topology of the physical information dynamic network, reverse synthesize virtual physical constraints and generate adaptive control strategy patches.
[0034] First, real-time acquisition of multi-source data streams can be achieved in various ways. For example, an independent sensor network can be used to aggregate sensor data from various subsystems such as the engine, transmission, battery, and motor to a central data processing unit. These data streams can be analog or digital signals, received via a data acquisition card or CAN bus interface. Another approach is to periodically read parameters stored within the vehicle's ECU via the On-Board Diagnostics (OBD) interface and include them as part of the data stream. Acquisition of vehicle status events can rely on preset triggering conditions; for example, when the vehicle speed exceeds a certain threshold, the driver switches driving modes, or the system detects a specific fault code, corresponding event signals are generated and transmitted.
[0035] Furthermore, the construction of physical information dynamic networks can adopt a predefined structure based on expert knowledge, that is, the initial network topology can be manually established according to the physical connections and known coupling relationships of the dynamic system.
[0036] For example, a connection can be pre-defined between the engine speed node and the transmission input speed node, and between the battery state of charge node and the motor power demand node. The continuous evolution of this network can be achieved by periodically updating the node state values and edge weights. For instance, at fixed time intervals, the latest sensor data can be mapped to the corresponding node states, and the edge weights can be adjusted according to simple rules (such as moving averages) to reflect changes in the correlation between variables. In this network, nodes are used to represent physical state variables, such as engine temperature, battery voltage, and vehicle speed, while edges are used to represent the dynamic coupling relationships between variables, such as the influence of engine speed on vehicle speed and the limitation of battery state of charge on motor output power.
[0037] Then, the sum of squared deviations between the current state values of all nodes in the network and the preset ideal state values can be used as one component of the energy function, and the sum of the weights of all edges can be used as another component. These two components are then simply added together. The low-energy state of this energy function is designed to correspond to the coordinated operating mode of the system. For example, when all node states are close to their ideal values and the coupling relationships in the network are stable, the value of the energy function is considered low, and the system is determined to be in a coordinated operating mode.
[0038] This can then be achieved by iterating through all possible actuator control combinations and evaluating the impact of each combination on the energy function. For example, several discrete values of motor torque can be enumerated, and the changes these values cause to the network state can be simulated to select the control quantity that will reduce the energy function the most. This minimum information disturbance is applied to the actuator, for example, by sending a specific torque command to the motor controller or a shift command to the transmission. This generates control commands for the actuators, designed to guide the power system toward a lower energy state of the energy function.
[0039] Finally, this can be achieved through manual analysis of historical network data. For example, engineers can periodically review long-term change logs of the network topology, identify recurring specific connection patterns or anomalous behaviors, and infer potential unmodeled physical constraints based on experience. Based on these inferences, simple rules or lookup tables can be manually written as virtual physical constraints. For instance, if it is found that battery charging efficiency is consistently lower than expected at a certain temperature, a virtual constraint can be synthesized specifying that the charging current should be limited at that temperature. Subsequently, an adaptive control strategy patch is generated. This patch can be a pre-written control logic fragment for a specific operating condition, which is activated when the virtual physical constraint is met to correct or enhance the existing control strategy.
[0040] In this embodiment, a dynamically evolving physical information network can be constructed to capture the complex dynamic coupling relationships between various variables in the powertrain system in real time. Based on minimizing the global network energy function, refined actuator control commands are generated, effectively solving the problems of traditional controller strategies being rigid and unable to adapt to environmental changes. Furthermore, by analyzing the long-term evolution topology of the network, reverse-synthesizing virtual physical constraints, and generating adaptive strategy patches, this method can compensate for the shortcomings of existing models, improve vehicle performance and energy efficiency optimization throughout its entire lifecycle, and effectively address extreme conditions that are difficult to cover with traditional road tests, reducing potential risks.
[0041] In some other embodiments, S102 may include: Based on the degree of contradiction between the internal node states of the physical information dynamic network, a network tension term is constructed; Based on the disorder degree of the network structure of the physical information dynamic network, a network entropy term is constructed. Based on the constraints of vehicle-level performance targets on the network state of the physical information dynamic network, a target term is constructed. The network tension term, network entropy term, and objective term are linearly combined with time-varying coupling coefficients to form the global network energy function.
[0042] Specifically, when constructing the network tension term, when there are deviations between certain node states (such as engine torque demand and the current gear position of the transmission) that do not conform to preset physical laws or optimization objectives, these deviations can be quantified as tension. This term can be constructed by defining ideal relationships or constraints between the states of each node and calculating the degree of deviation between the current actual state and the ideal state. For example, it can be calculated using squared error, absolute error, or prediction error based on a physical model to reflect the degree of "disharmony" within the system. The introduction of this term allows the energy function to actively eliminate or alleviate internal contradictions during the minimization process, promoting coordinated work among components.
[0043] Simultaneously, when constructing the network entropy term, this term is used to measure the structural complexity, randomness, or unpredictability of a physical information dynamic network. A high-entropy network typically implies structural instability, low information transmission efficiency, or redundancy / chaos. For example, entropy can be calculated by analyzing the weight distribution of edges, the uniformity of node connection patterns, or the volatility of node states. The introduction of this term aims to guide physical information dynamic networks towards a more stable, simpler, and more predictable structure, thereby improving the system's robustness and control efficiency. For example, metrics such as Shannon entropy based on information theory or structural entropy from graph theory can be used to quantify the network's disorder.
[0044] Furthermore, when constructing the objective term, this term associates the abstract physical information dynamic network state with specific vehicle performance indicators (such as fuel economy, emissions, driving comfort, and power response). For example, based on a preset performance indicator function, the impact of the current network state on vehicle performance can be evaluated and incorporated into the energy function as a penalty or reward term. When the network state deviates from the vehicle-level performance target, the value of this objective term increases, thereby guiding the network state to adjust towards meeting or optimizing vehicle performance during the energy function minimization process.
[0045] Building upon this, the network tension term, network entropy term, and objective term are linearly combined with time-varying coupling coefficients to form the global network energy function. This linear combination provides a flexible and interpretable mechanism for balancing internal system coordination, structural stability, and external performance objectives. The time-varying coupling coefficients are crucial, allowing for dynamic adjustment of the relative importance of each component based on the vehicle's real-time operating conditions, driver intent, environmental conditions, or system priorities. For example, in scenarios prioritizing optimal power response, the coupling coefficient of objective terms related to power performance can be increased; conversely, in scenarios prioritizing fuel economy, the coupling coefficient of objective terms related to fuel consumption can be increased. These time-varying coupling coefficients can be learned and adjusted online through a pre-defined operating condition mapping table, a rule-based expert system, or machine learning algorithms (such as reinforcement learning) to ensure that the global network energy function dynamically adapts to complex vehicle operating environments and control requirements.
[0046] In this embodiment, the complex state of the physical information dynamic network can be comprehensively and multidimensionally quantified. The network tension term effectively captures the potential contradictions and conflicts between various physical state variables within the system, prompting the system to tend towards an internally coordinated and consistent state. The network entropy term guides the network structure to evolve towards a more stable and orderly direction, preventing the system from falling into a chaotic or unpredictable state. The objective term directly integrates the vehicle-level performance objective into the energy function, ensuring that minimizing the energy function not only signifies internal coordination and structural stability but also directly serves the overall performance optimization of the vehicle. The introduction of the time-varying coupling coefficient allows the energy function to dynamically adjust the weights of each component according to real-time operating conditions and control requirements, thereby flexibly balancing internal coordination, system stability, and external performance objectives under different driving scenarios. This refined energy function construction method enables the low-energy state of the global network energy function to more accurately and robustly correspond to the optimal coordinated operating mode of the powertrain system, providing a solid foundation for the subsequent generation of efficient and adaptive control commands, and significantly improving the intelligence and adaptability of the powertrain domain controller.
[0047] In some other embodiments, S103 may include: Based on the gradient information of the global network energy function under the current network state, and the constraints formed by the system physical equations, a Lagrangian optimization problem is constructed. Solving the optimization problem yields the minimum change vector required to influence the state of each node in order to guide the energy function to decrease. The state components in the minimum change vector that correspond to those that the actuator can directly influence are converted into actual control commands.
[0048] In this embodiment, firstly, a Lagrange optimization problem is constructed based on the gradient information of the global network energy function under the current network state and the constraints formed by the system's physical equations. The Lagrange optimization problem is a method for finding the extremum of a function under given constraints. Its core lies in combining the minimization objective of the global network energy function with the constraints formed by the actual physical equations of the system to form a mathematically solvable optimization framework. This ensures that while seeking to reduce the energy function, the generated control commands will not violate physical laws or the inherent operational limitations of the system. Specifically, the global network energy function can be used as the objective function, while the system's physical equations (e.g., vehicle dynamics models, engine combustion models, battery charging and discharging models, etc.) serve as equality or inequality constraints. These physical equations describe the intrinsic relationships between system state variables (such as speed, torque, voltage, current, etc.) and the influence of actuators (such as throttle opening, braking pressure, motor current, etc.) on these state variables. By introducing Lagrange multipliers, the constraints are integrated into the objective function, thereby transforming the constrained optimization problem into an unconstrained optimization problem for solution.
[0049] Secondly, the optimization problem is solved to obtain the minimum change vector applied to the states of each node to guide the energy function downward. Solving the Lagrange optimization problem aims to find an optimal solution that represents the direction and magnitude of state changes that, while satisfying all physical constraints, allows for the fastest decrease in the global network energy function. This "minimum change vector" embodies the principle of "minimum information disturbance," that is, driving the system towards a lower energy state (coordinated working mode) in the most economical and efficient way, avoiding unnecessary or excessive control actions. Iterative optimization algorithms can be used to solve this problem, such as gradient descent, Newton's method, quasi-Newton methods, and sequential quadratic programming (SQP). These algorithms gradually approach the optimal solution by calculating the gradient information of the objective function and constraints. In each iteration, the algorithm calculates the direction and step size of the next state change based on the current state and gradient information, until the convergence condition is met or the preset number of iterations is reached. The obtained minimum change vector contains the ideal change amount of all node states in the network (including directly controllable and indirectly controllable states).
[0050] Finally, the state components in the minimum change vector that correspond to those directly affected by the actuator are converted into actual control commands. The minimum change vector contains the ideal changes of all node states, but the actuator can only directly control a portion of these states. This step transforms the abstract ideal state changes into concrete physical quantities that the actuator can understand and execute, thereby achieving precise control of the vehicle's powertrain. First, it is necessary to identify which components in the minimum change vector correspond to physical states that the actuator can directly affect. For example, if a node state represents engine speed, and engine speed can be controlled by throttle opening or fuel injection quantity, then this component needs to be mapped to the corresponding actuator command. This conversion typically involves a control mapping function or inverse dynamics model that can calculate the corresponding actuator input based on the required state change. For example, the target speed change can be converted into a throttle opening command or a motor torque command through lookup tables, model-based inverse calculations, or PID controllers. The final generated actual control commands are digital or analog signals that can be directly sent to vehicle actuators (such as engine controllers, transmission controllers, battery management systems, etc.).
[0051] By combining the minimization requirement of the global network energy function with the constraints imposed by the system's physical equations, a Lagrangian optimization problem is constructed and solved. This allows for the precise calculation of the minimum state change vector required to guide the energy function downwards. This ensures that the generated control commands not only effectively propel the system towards a coordinated operating mode but also strictly adhere to the physical limitations and dynamic characteristics of the vehicle's powertrain, avoiding unrealistic or excessively perturbed control. Finally, by converting the state components directly affected by the actuators from these minimum change vectors into actual control commands, precise, efficient, and physically feasible control of the actuators is achieved, thereby improving the response speed, control accuracy, and system stability of the power domain controller under complex operating conditions.
[0052] In other embodiments, control commands for the actuators are generated by constructing a physical information dynamic network and minimizing the global network energy function. However, during actual vehicle operation, due to environmental changes, component aging, or complex factors such as lack of modeling, the pre-constructed physical information dynamic network may not accurately reflect the real-time dynamics of the system. This can lead to deviations between the generated control commands and the actual system response, thereby affecting the accuracy and reliability of the control. Therefore, after S104, the method may include: Based on physical information dynamic networks, predict the evolution of the network state at the next moment after the execution of control commands; Acquire the actual sensor data at the next moment and calculate the difference between the evolution result and the network state induced by the actual data; If the difference exceeds the preset threshold, it is determined that the control command is inconsistent with the network causal model, triggering an immediate online correction of the weights of the relevant connection edges in the physical information dynamic network, and recalculating the control command based on the corrected network.
[0053] In this embodiment, after generating control commands for the actuators, the system first predicts the evolution of each node (representing physical state variables) and edge (representing dynamic coupling relationships between variables) in the network at the next moment, based on the currently constructed physical information dynamic network and the impact of the control commands on the system state. This prediction process utilizes the system dynamics model inherent in the physical information dynamic network, simulating or calculating the system response under the control commands to obtain an expected future network state. This prediction result provides a benchmark for subsequent evaluation of actual system behavior.
[0054] Subsequently, the system acquires actual sensor data at the next moment, reflecting the true operating state of the vehicle's powertrain after executing control commands. By processing and mapping this actual sensor data, the system obtains the actual physical state variable values corresponding to nodes in the physical information dynamic network, thereby inducing the actual network state at the next moment. Next, the system calculates the difference between the previously predicted network state and the currently induced network state. This difference quantifies the degree of inconsistency between model predictions and actual observations and can be calculated using various mathematical methods, such as Euclidean distance, weighted squared difference, or error functions based on specific physical meanings.
[0055] If the calculated difference exceeds a preset threshold, it indicates a significant inconsistency between the currently generated control commands and the causal model of the system represented by the Physical Information Dynamics Network (PIN). This may mean that some parameters in the PSN (especially the edge weights) no longer accurately reflect the dynamic coupling relationships of the actual system. In this case, the system will immediately trigger an online correction of the relevant edge weights in the PSN. The correction process aims to adjust the network parameters to better fit the actual observation data, thereby improving the accuracy of the model. After the correction is completed, the system will recalculate the control commands based on the updated PSN to ensure that the new commands can more accurately guide the system to the desired coordinated operating mode.
[0056] This application enables the system to promptly detect inconsistencies between control commands and the network causal model through real-time prediction, actual data feedback, and discrepancy calculation. Upon detecting significant discrepancies, it immediately triggers online correction of the weights of key connections in the physical information dynamic network, ensuring that the network model remains highly consistent with the actual physical system. Based on this, the control commands are recalculated, significantly improving the adaptability, accuracy, and robustness of the control, avoiding performance degradation due to model drift, and ensuring the coordinated operation and optimal energy efficiency of the power system under various working conditions.
[0057] In some other embodiments, S105 may include: Monitor whether stable anomalous subgraph patterns appear in the long-term evolution of the physical information dynamic network. Anomalous subgraph patterns include a specific set of nodes and corresponding strong connection edges, and the energy level is consistently higher than the network baseline. The abnormal subgraph pattern is isolated from the physical information dynamic network and used as an input to an independent constraint synthesis problem; The inversion algorithm is used to solve the following: under the assumption that there are certain additional, unmodeled physical constraints in the system, the global network energy function will form a local minimum at the currently observed anomalous subgraph structure; The additional physical constraints obtained from the solution are formally described as a single or set of virtual physical laws, which serve as the basis for the synthetic control strategy patch.
[0058] In the embodiments of this application, long-term evolution refers to the changes in parameters such as the structure, node state, and edge weight of a physical information dynamic network over a long period of time. This may be caused by system wear and tear, environmental changes, or unmodeled external interference.
[0059] Stable anomalous subgraph patterns refer to the persistent characteristics of certain local network structures (i.e., subgraphs) that are inconsistent with the overall network baseline behavior during the long-term evolution of the network. For example, the energy level is abnormally high, and this anomalous state is not a transient phenomenon, but a persistent one.
[0060] Specific nodes refer to physical state variables in anomaly subgraph patterns that exhibit abnormal behavior or have abnormally strong coupling relationships with other nodes.
[0061] Strongly connected edges refer to significant, unconventional dynamic coupling relationships between specific nodes, indicating that there are close, potentially poorly understood interactions between them.
[0062] A sustained energy level above the network baseline indicates that the local system represented by the subgraph is in a state of high energy consumption, high inconsistency, or high uncertainty, deviating from the system's expected coordinated operating mode. Potential subgraph structures can be identified by periodically performing snapshot analysis on the dynamic physical information network and utilizing graph theory algorithms (such as community detection and subgraph isomorphism detection). Simultaneously, statistical methods (such as moving averages and exponential smoothing) are used to evaluate the subgraph's energy level and compare it with a pre-defined or dynamically calculated "network baseline" (e.g., the mean energy distribution or threshold obtained through training with historical normal operation data). When the energy level of a subgraph consistently exceeds the baseline threshold over multiple consecutive time windows, it can be determined as a stable anomalous subgraph pattern.
[0063] Isolation refers to logically or computationally separating the identified anomalous subgraph patterns from the overall physical information dynamic network, making them an independent research object for in-depth analysis without being disturbed by the complexity of other parts of the network.
[0064] The independent constraint synthesis problem refers to an optimization or inversion problem specifically designed for this anomalous subgraph pattern, with the goal of discovering the underlying physical constraints that lead to the anomalous behavior.
[0065] The structure, node states, edge weights, and persistently high energy levels of the anomalous subgraph are used as known conditions for solving this independent problem. Isolation can be achieved by creating a copy of the subgraph or by focusing only on the nodes and edges of the subgraph during computation. The independent constraint synthesis problem can be modeled as a reverse engineering problem, with inputs including the topology of the anomalous subgraph, time-series data of node states, and their corresponding energy function values.
[0066] An inversion algorithm is a computational method that infers the cause or underlying mechanism from observation results. In this context, it refers to the reverse deduction of possible physical laws or constraints that are not captured by the current physical information dynamic network model by analyzing the observed behavior of anomaly subgraphs (i.e., their persistent high-energy states).
[0067] Additional, unmodeled physical constraints refer to physical laws that were not considered or could not be accurately modeled during system design or initial modeling, such as increased friction due to component wear, changes in elastic modulus caused by material fatigue, or fluid dynamic effects under specific environments. The existence of these constraints causes the system to exhibit abnormalities under specific conditions.
[0068] The fact that the global network energy function forms a local minimum at the currently observed anomalous subgraph structure means that if these additional constraints do exist, then after considering these constraints, the system state corresponding to the anomalous subgraph will no longer be a high-energy state, but will reach a relatively stable, low-energy state under the new constraints, i.e., a local minimum. Inversion algorithms can employ various mathematical optimization techniques, such as gradient descent-based optimization, Monte Carlo methods, genetic algorithms, or Bayesian inference. Specifically, a parameterized physical constraint model can be constructed, and then the model parameters can be iteratively adjusted so that, after introducing these parameterized constraints, the energy function value of the anomalous subgraph converges to a local minimum, and this local minimum matches the observed anomalous behavior.
[0069] Formal description refers to transforming abstract physical constraints derived through inversion algorithms into a structured, computable expression, such as mathematical equations, logical rules, or state machines.
[0070] Formal descriptions can employ difference equations, algebraic equations, inequalities, conditional statements (such as IF-THEN rules), or finite state machines. For example, if an inversion algorithm discovers that the friction of a component increases non-linearly with increasing temperature, it can be formalized as a virtual physical law: `F_friction = f(Temperature, wear_coefficient)`. These virtual physical laws can then be integrated into the controller's decision logic as a basis for modifying existing control strategies.
[0071] This application can continuously monitor the evolution of the physical information dynamic network, promptly detect and isolate stable anomalous subgraph patterns, and accurately locate the source of anomalous behavior in the system. Furthermore, it utilizes an inversion algorithm to deduce the underlying virtual physical laws leading to these anomalies, enabling the system to "learn" unknown physical mechanisms from the data. Formalizing these virtual physical laws and using them as the basis for patching control strategies allows the controller to adaptively adjust its behavior, thereby providing robust and effective responses to complex and dynamically changing real-world conditions. This significantly improves the long-term adaptability, reliability, and performance stability of the dynamic domain controller.
[0072] In other embodiments, generating an adaptive control policy patch includes: The formally described virtual physical laws are compiled into a lightweight policy patching module. The policy patching module includes a state correction function and a rule injection unit. The state correction function is used to preprocess or correct the sensor data input to the traditional control algorithm according to the virtual physical laws. The rule injection unit is used to temporarily add extra terms corresponding to the virtual physical laws to the evolution equation of the physical information dynamic network. The policy patch module is received and securely loaded via the vehicle OTA channel.
[0073] In this embodiment, to enable execution by the power domain controller, it needs to be converted into executable code or configuration. The compilation process transforms these high-level descriptions into low-level instruction sets or data structures of a specific format. For example, they can be compiled into dynamic link libraries (DLLs), shared objects (SOs), script files, or configurable parameter sets. The policy patch module is designed to be "lightweight," meaning it has a small code size, low runtime resource consumption (such as CPU and memory), and can be loaded and unloaded quickly to minimize the impact on the main control system performance and support rapid iterative updates.
[0074] The strategy patching module can include a state correction function and a rule injection unit. The state correction function is designed to preprocess or correct the sensor data input to the traditional control algorithm. When virtual physics laws reveal a deviation or unmodeled correlation between sensor data and the actual physical state, the state correction function adjusts the original sensor data in real time according to these laws. For example, if virtual physics laws indicate that a temperature sensor reading is generally high under specific operating conditions, the state correction function can apply a correction factor or offset. This can be implemented as a data transformation pipeline, filtering, scaling, offsetting, or performing more complex nonlinear transformations on the sensor data before it enters the traditional control loop. This allows the traditional control algorithm to make decisions based on more accurate data that better conforms to newly discovered physical constraints without modifying its own logic. The rule injection unit is used to temporarily add extra terms corresponding to the virtual physics laws to the evolution equations of the physical information dynamic network. Unlike the state correction function, which focuses on data input, the rule injection unit directly acts on the network's dynamic model. When virtual physics laws reveal new dynamic coupling relationships between nodes (physical state variables) in a network, or when existing coupling relationships need adjustment, the rule injection unit dynamically modifies the network's adjacency matrix, weight parameters, or evolution rules. For example, if virtual physics laws indicate that two previously considered independent physical states are strongly correlated under specific conditions, the rule injection unit can introduce a cross term reflecting this correlation into the network's evolution equation, enabling the physical information dynamic network to more accurately simulate the system's real behavior. This injection is "temporary," meaning it can be activated when needed and removed when no longer required, providing great flexibility.
[0075] Furthermore, policy patch modules are received and securely loaded via the vehicle's OTA (Over-The-Air) channel. The vehicle OTA channel is a wireless communication mechanism that allows vehicles to remotely receive software updates and configurations. The reception process involves establishing a secure communication link, such as using encryption protocols (like TLS / SSL) to ensure the confidentiality and integrity of data transmission. Secure loading refers to the rigorous verification performed on the received policy patch module by the system, including but not limited to digital signature verification (confirming the legitimacy of the patch's source), integrity verification (preventing data tampering during transmission), compatibility checks (ensuring the patch is compatible with the current system version), and sandbox environment pre-loading tests (verifying the patch's behavior in an isolated environment). Only patches that pass all security verifications are loaded into the powertrain domain controller and activated. This mechanism enables vehicles to continuously and remotely obtain the latest adaptive control policies without requiring users to visit service stations, significantly improving system maintenance efficiency and response speed.
[0076] By compiling virtual physics laws into lightweight policy patch modules, which include state correction functions and rule injection units, the system can adaptively adjust in two complementary ways: the state correction function indirectly introduces new physical constraints by preprocessing sensor data without modifying the core logic of traditional control algorithms, reducing the complexity of introducing new risks; while the rule injection unit can modify the evolution equations of the physical information dynamic network at a deeper level, directly reflecting new dynamic coupling relationships, thereby improving the accuracy and predictive ability of the network model. Furthermore, by receiving and securely loading the policy patch modules through the vehicle's OTA channel, these adaptive policies can be remotely, quickly, and securely deployed to the vehicle. This allows the power domain controller to continuously learn and adapt to the vehicle's operating environment, component aging, or unmodeled physical phenomena, significantly enhancing the robustness, adaptability, and long-term reliability of the control system, and avoiding performance degradation or safety hazards caused by system behavior inconsistent with the model.
[0077] In some other embodiments, after S105, the method may include: Obtain metacognitive datasets for historical time periods. Metacognitive datasets include all network evolution trajectories, energy function history, control command sequences, and final performance indicators. When cloud or vehicle-side computing power allows, initiate the meta-reinforcement learning process. The meta-reinforcement learning process is used to optimize the hyperparameters of the construction rules of the physical information dynamic network, the dynamic coupling coefficient generation strategy of the global network energy function, and the heuristic rules for reverse synthesis of virtual physical constraints. The optimization strategies obtained from meta-learning are encapsulated into metacognitive update packages and sent to vehicles via in-vehicle OTA.
[0078] In this embodiment, the metacognitive dataset acquired over a historical time period can include all network evolution trajectories, energy function history, control command sequences, and final performance indicators. It refers to the data set that records and summarizes the system's operating state, decision-making process, and results during operation. This dataset not only includes direct sensor data or control outputs but also focuses on the system's internal "thinking" process and "learning" effects. Specifically, the network evolution trajectory records the structural changes (nodes, edges, and their weights) of the physical information dynamic network at different points in time; the energy function history records the numerical changes of the global network energy function in different control cycles, reflecting the stability and optimization level of the system's coordinated working mode; the control command sequence records all control commands issued by the system to the actuators; and the final performance indicators quantify the effects achieved by these control commands in the actual physical system, such as fuel economy, emissions, power response, and driving smoothness. This data forms the basis for the system's high-level learning and self-optimization. This metacognitive dataset can be stored by setting up a dedicated log recording module within the power domain controller or by uploading the relevant data in real time to an onboard data recording unit or a cloud server. Data acquisition should have high temporal resolution and completeness to ensure that it can accurately reflect the system's behavior patterns under different operating conditions.
[0079] Building upon this foundation, and provided the computing power in the cloud or on-vehicle environment permits, a meta-reinforcement learning process is initiated. This process optimizes the hyperparameters of the construction rules for the physical information dynamic network, the dynamic coupling coefficient generation strategy for the global network energy function, and the heuristic rules for back-synthesizing virtual physical constraints. Meta-reinforcement learning is a high-level learning paradigm designed to enable learning systems to "learn how to learn." In this context, it does not directly learn control strategies, but rather learns how to generate or adjust the underlying parameters and rules of those strategies. Specifically, the goals of the meta-reinforcement learning process are: optimizing the hyperparameters of the construction rules for the physical information dynamic network, which determine how the network extracts physical information and dynamic relationships from raw data (e.g., node connection thresholds, edge weight initialization methods, network update frequency, etc.); optimizing the dynamic coupling coefficient generation strategy for the global network energy function, which affects the relative weights of the network tension term, network entropy term, and objective term in the energy function, thus determining the system's trade-offs between contradictions, disorder, and performance objectives under different operating conditions; and optimizing the heuristic rules for back-synthesizing virtual physical constraints, which guide the system inferring virtual physical constraints from anomalous subgraph patterns. Through meta-reinforcement learning, the system can autonomously discover and adjust these underlying mechanisms to adapt to a wider range of more complex operating environments. This meta-reinforcement learning process can be trained offline on high-performance cloud computing clusters, using large-scale historical data for iterative optimization; or it can be learned online or semi-online on vehicle-side domain controllers or central computing platforms with sufficient computing power to achieve faster adaptability. This process typically involves defining a meta-reward function that evaluates the merits of the current hyperparameters and policy based on a final performance metric, and updating the meta-policy through algorithms such as meta-policy gradient and meta-Q-learning, thereby generating better underlying rules.
[0080] Furthermore, the optimization strategies obtained from meta-learning are encapsulated into a metacognitive update package and delivered to the vehicle via onboard OTA (Over-The-Air). The metacognitive update package is a software module that encapsulates the optimization results obtained from the meta-reinforcement learning process. These optimization results include, but are not limited to, updated network construction hyperparameters, code or configuration of dynamic coupling coefficient generation strategies, and heuristic rule sets for inverse synthesis of virtual physical constraints. This update package aims to deliver high-level "learning how to learn" results to the vehicle in a structured and deployable manner. Onboard OTA delivery is a wireless update technology that allows vehicles to receive and install software updates via cellular networks or Wi-Fi without needing to visit a service center. This method ensures that vehicles can obtain the latest optimization strategies in a timely and secure manner, thereby continuously improving the performance and adaptability of their power domain controllers. The metacognitive update package can be a compressed file containing configuration parameter files, algorithm module code, or model weights. During the delivery process, the security (e.g., encryption, digital signatures) and integrity of data transmission must be ensured, and strict verification and validation must be performed on the vehicle side to prevent malicious tampering or update failure. After the vehicle receives the update package, the onboard software update management system is responsible for decompressing, installing and activating the new optimization strategy. This usually requires restarting the relevant control modules for the new strategy to take effect.
[0081] By leveraging meta-reinforcement learning and utilizing metacognitive data such as network evolution trajectories, energy function histories, control command sequences, and final performance indicators over historical time periods, the system autonomously learns and adjusts the hyperparameters of the physical information dynamic network construction rules, the dynamic coupling coefficient generation strategy of the global network energy function, and the heuristic rules for reverse-synthesizing virtual physical constraints. This allows the system to move beyond relying on preset fixed parameters or manual tuning, continuously optimizing its learning and adaptive capabilities based on actual operational data and performance feedback. When the optimized strategy obtained through meta-learning is distributed to the vehicle via onboard OTA, the vehicle's power domain controller can construct and evolve the physical information dynamic network in a more optimized manner, more accurately minimize the global network energy function to generate control commands, and more effectively reverse-synthesize virtual physical constraints and generate adaptive control strategy patches. This continuous self-optimization mechanism significantly improves the robustness, adaptability, and overall performance of the power domain controller in the face of complex and variable operating conditions, ensuring that the system can maintain an optimal or near-optimal operating state over a long period, thus effectively solving the technical problem of the difficulty in continuously adaptively optimizing parameters and strategies in traditional methods.
[0082] In other embodiments, this application also provides a power domain controller configuration system, wherein the power domain controller configuration system may include: Virtual Validation Layer: As the core of policy exploration and training, it automatically explores extreme conditions through reinforcement learning and generates Pareto optimal policies.
[0083] Edge Adaptation Layer: As a policy execution and correction engine, it is embedded in the vehicle domain controller and corrects the simulation-real vehicle differences online through lightweight transfer learning.
[0084] Physical execution layer: Includes motor controller, battery management system, etc., which executes power output and collects data to drive the continuous evolution of virtual model.
[0085] The system connects the virtual layer and the edge layer through an OTA bidirectional channel, forming a closed loop.
[0086] In this embodiment, the virtual verification layer may include: Extreme operating condition generator: It adopts a generative adversarial network (GAN) to generate extreme high temperature, high altitude, and cold extreme operating conditions on the NEDC / WLTC basic operating conditions. The feasibility is ensured by physical rule verification, and about 10,000 valid new operating conditions are generated every day. Reinforcement learning trainer: Based on the PPO algorithm, it is trained on a digital twin model and outputs Pareto optimal policy parameters; Configuration compiler generator: Automatically compiles JSON configuration descriptions into C code compliant with the AUTOSAR standard, generating ARXML interface descriptions; Virtual twin database: Stores millions of labeled working conditions to support continuous model evolution.
[0087] The edge adaptation layer may include: Transfer learning module: Uses Online SVR to correct simulation-real vehicle residuals, eliminating the need for full retraining; Configuration interpretation engine: dynamically loads JSON configuration packages, with hot-switching latency <50ms; Online prediction module: Deploys a lightweight LSTM network to predict braking intentions 2-5 seconds in advance; Safety monitoring module: three-indicator (parameter deviation, energy consumption deterioration, failure rate) graded protection, automatic rollback when exceeding limits; Runtime data caching: minute-level operation condition segment caching, automatic marking and uploading of abnormal events; Physical execution layer: includes motor controller, battery management system, vehicle control unit (VCU), thermal management controller, etc. It communicates with the edge layer in real time via CAN / CAN-FD bus, executes commands and feeds back sensor data.
[0088] High-value data bidirectional transmission channel: OTA bidirectional channel, downlink transmission strategy parameter packet (256 bytes), configuration description packet (128 bytes), prediction model packet (<100KB), uplink transmission condition segment packet (2KB), residual data packet, and fault event packet.
[0089] Based on the power domain controller configuration method provided in the above embodiments, this application also provides specific implementations of the power domain controller configuration device. Please refer to the following embodiments.
[0090] First see Figure 2 The power domain controller configuration device 200 provided in this application embodiment may include: The acquisition module 201 is used to acquire real-time multi-source data streams of the power system and vehicle status events. The real-time multi-source data streams include data from different sources collected in real time from the vehicle power system and related sensors. The first construction module 202 is used to construct and continuously evolve a physical information dynamic network based on real-time multi-source data streams and vehicle state events. The nodes of the physical information dynamic network represent physical state variables, and the edges of the physical information dynamic network represent the dynamic coupling relationship between variables. The second construction module 203 is used to construct and minimize a global network energy function based on the physical information dynamic network, wherein the low energy state of the global network energy function corresponds to the coordinated working mode of the system. The first generation module 204 is used to calculate and apply minimum information perturbation according to the requirement of minimizing the global network energy function, and generate control instructions for the actuator; The second generation module 205 is used to reverse synthesize virtual physical constraints based on the long-term evolution topology of the physical information dynamic network, and generate adaptive control strategy patches.
[0091] As an optional implementation, the first building module 202 can also be used for: Based on the degree of contradiction between the internal node states of the physical information dynamic network, a network tension term is constructed; Based on the disorder degree of the network structure of the physical information dynamic network, a network entropy term is constructed. Based on the constraints of vehicle-level performance targets on the network state of the physical information dynamic network, a target term is constructed. The network tension term, network entropy term, and objective term are linearly combined with time-varying coupling coefficients to form the global network energy function.
[0092] As an optional implementation, the second building module 203 can also be used for: Based on the gradient information of the global network energy function under the current network state, and the constraints formed by the system physical equations, a Lagrangian optimization problem is constructed. Solving the optimization problem yields the minimum change vector required to influence the state of each node in order to guide the energy function to decrease. The state components in the minimum change vector that correspond to those that the actuator can directly influence are converted into actual control commands.
[0093] As an optional implementation, the first generation module 204 can also be used for: Based on physical information dynamic networks, predict the evolution of the network state at the next moment after the execution of control commands; Acquire the actual sensor data at the next moment and calculate the difference between the evolution result and the network state induced by the actual data; If the difference exceeds the preset threshold, it is determined that the control command is inconsistent with the network causal model, triggering an immediate online correction of the weights of the relevant connection edges in the physical information dynamic network, and recalculating the control command based on the corrected network.
[0094] As an optional implementation, the second generation module 205 can also be used for: Monitor whether stable anomalous subgraph patterns appear in the long-term evolution of the physical information dynamic network. Anomalous subgraph patterns include a specific set of nodes and corresponding strong connection edges, and the energy level is consistently higher than the network baseline. The abnormal subgraph pattern is isolated from the physical information dynamic network and used as an input to an independent constraint synthesis problem; The inversion algorithm is used to solve the following: under the assumption that there are certain additional, unmodeled physical constraints in the system, the global network energy function will form a local minimum at the currently observed anomalous subgraph structure; The additional physical constraints obtained from the solution are formally described as a single or set of virtual physical laws, which serve as the basis for the synthetic control strategy patch.
[0095] As an optional implementation, the second generation module 205 can also be used for: The formally described virtual physical laws are compiled into a lightweight policy patching module. The policy patching module includes a state correction function and a rule injection unit. The state correction function is used to preprocess or correct the sensor data input to the traditional control algorithm according to the virtual physical laws. The rule injection unit is used to temporarily add extra terms corresponding to the virtual physical laws to the evolution equation of the physical information dynamic network. The policy patch module is received and securely loaded via the vehicle OTA channel.
[0096] As an optional implementation, the second generation module 205 can also be used for: Obtain metacognitive datasets for historical time periods. Metacognitive datasets include all network evolution trajectories, energy function history, control command sequences, and final performance indicators. When cloud or vehicle-side computing power allows, initiate the meta-reinforcement learning process. The meta-reinforcement learning process is used to optimize the hyperparameters of the construction rules of the physical information dynamic network, the dynamic coupling coefficient generation strategy of the global network energy function, and the heuristic rules for reverse synthesis of virtual physical constraints. The optimization strategies obtained from meta-learning are encapsulated into metacognitive update packages and sent to vehicles via in-vehicle OTA.
[0097] Figure 3 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0098] An electronic device may include a processor 301 and a memory 302 storing computer program instructions.
[0099] Specifically, the processor 301 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0100] Memory 302 may include mass storage for data or instructions. For example, and not limitingly, memory 302 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 302 may include removable or non-removable (or fixed) media, or memory 302 may be non-volatile solid-state memory. Memory 302 may be internal or external to the integrated gateway disaster recovery device.
[0101] In one instance, memory 302 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0102] Memory 302 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the power domain controller configuration method according to the first aspect of this disclosure.
[0103] The processor 301 reads and executes computer program instructions stored in the memory 302 to achieve... Figure 1 A power domain controller configuration method in the illustrated embodiment.
[0104] In one example, the electronic device may also include a communication interface 303 and a bus 304. For example, Figure 3 As shown, the processor 301, memory 302, and communication interface 303 are connected through bus 304 and complete communication with each other.
[0105] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0106] Bus 304 includes hardware, software, or both, that couples components of an electronic device together. For example, and not as a limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 304 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0107] The electronic device can execute the power domain controller configuration method in the embodiments of this application, thereby achieving the combination Figures 1-2 The described method and apparatus for configuring a power domain controller.
[0108] Furthermore, in conjunction with the power domain controller configuration methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the power domain controller configuration methods in the above embodiments.
[0109] In an optional embodiment, in conjunction with the power domain controller configuration method in the above embodiments, this application embodiment can provide a computer program product to implement it. The instructions in the computer program product are executed by the processor of the electronic device, enabling the electronic device to implement any of the power domain controller configuration methods in the above embodiments.
[0110] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0111] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0112] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0113] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0114] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for configuring a dynamic domain controller, characterized in that, include: Acquire real-time multi-source data streams of the powertrain system and vehicle status events, wherein the real-time multi-source data streams include data from different sources collected in real time from the vehicle powertrain system and related sensors; Based on the real-time multi-source data stream and the vehicle state events, a physical information dynamic network is constructed and continuously evolved, wherein the nodes of the physical information dynamic network represent physical state variables, and the edges of the physical information dynamic network represent the dynamic coupling relationship between variables. Based on the physical information dynamic network, a global network energy function is constructed and minimized, wherein the low-energy state of the global network energy function corresponds to the coordinated working mode of the system; Based on the requirement of minimizing the global network energy function, the minimum information perturbation is calculated and applied to generate control commands for the actuator; Based on the long-term evolution topology of the physical information dynamic network, virtual physical constraints are synthesized in reverse, and adaptive control strategy patches are generated.
2. The method according to claim 1, characterized in that, The process of constructing and minimizing a global network energy function based on the physical information dynamic network includes: Based on the degree of contradiction between the internal node states of the physical information dynamic network, a network tension term is constructed; Based on the disorder degree of the network structure of the physical information dynamic network, a network entropy term is constructed. Based on the constraints of the vehicle-level performance targets on the network state of the physical information dynamic network, a target term is constructed; The network tension term, network entropy term, and target term are linearly combined using time-varying coupling coefficients to form the global network energy function.
3. The method according to claim 1 or 2, characterized in that, The step of calculating and applying minimum information perturbation based on the requirement of minimizing the global network energy function, and generating control commands for the actuator, includes: Based on the gradient information of the global network energy function under the current network state, and the constraints formed by the system physical equations, a Lagrange optimization problem is constructed. Solving the optimization problem yields the minimum change vector required to influence the state of each node in order to guide the energy function to decrease. The state components in the minimum change vector that correspond to those that the actuator can directly influence are converted into actual control commands.
4. The method according to claim 1, characterized in that, After generating control instructions for the actuator, the method further includes: Based on the physical information dynamic network, predict the evolution of the network state at the next moment after the execution of the control command; Acquire the actual sensor data at the next moment and calculate the difference between the evolution result and the network state induced by the actual data; If the difference exceeds a preset threshold, it is determined that the control command is inconsistent with the network causal model, triggering an immediate online correction of the weights of the relevant connection edges in the physical information dynamic network, and recalculating the control command based on the corrected network.
5. The method according to claim 1, characterized in that, The step of reverse-synthesizing virtual physical constraints based on the long-term evolution topology of the physical information dynamic network and generating adaptive control policy patches includes: The monitoring of the physical information dynamic network is conducted to determine whether a stable anomalous subgraph pattern emerges during its long-term evolution. The anomalous subgraph pattern includes a specific set of nodes and corresponding strong connection edges, and the energy level is consistently higher than the network baseline. The abnormal subgraph pattern is isolated from the physical information dynamic network and used as input to an independent constraint synthesis problem; The inversion algorithm is used to solve the following: when there are any additional, unmodeled physical constraints in the system, the global network energy function will form a local minimum at the currently observed anomalous subgraph structure, thus obtaining the additional physical constraints; The additional physical constraints obtained from the solution are formally described as a single or set of virtual physical laws, which serve as the basis for the synthetic control strategy patch.
6. The method according to claim 5, characterized in that, The generation of adaptive control strategy patches includes: The formally described virtual physical laws are compiled into a lightweight policy patching module. The policy patching module includes a state correction function and a rule injection unit. The state correction function is used to preprocess or correct the sensor data input to the traditional control algorithm according to the virtual physical laws. The rule injection unit is used to temporarily add extra terms corresponding to the virtual physical laws to the evolution equation of the physical information dynamic network. The policy patch module is received and securely loaded via the vehicle OTA channel.
7. The method according to claim 1, characterized in that, After synthesizing virtual physical constraints in reverse based on the long-term evolution topology of the physical information dynamic network and generating adaptive control policy patches, the method further includes: Obtain a metacognitive dataset within a historical time period, which includes all network evolution trajectories, energy function history, control command sequences, and final performance indicators; When cloud or vehicle-side computing power allows, initiate a meta-reinforcement learning process, which is used to optimize the hyperparameters of the construction rules of the physical information dynamic network, the dynamic coupling coefficient generation strategy of the global network energy function, and the heuristic rules of the reverse synthesis virtual physical constraints. The optimization strategies obtained from meta-learning are encapsulated into metacognitive update packages and sent to vehicles via in-vehicle OTA.
8. A power domain controller configuration device, characterized in that, The device includes: The acquisition module is used to acquire real-time multi-source data streams of the power system and vehicle status events. The real-time multi-source data streams include data from different sources collected in real time from the vehicle power system and related sensors. The first construction module is used to construct and continuously evolve a physical information dynamic network based on the real-time multi-source data stream and the vehicle state event, wherein the nodes of the physical information dynamic network represent physical state variables, and the edges of the physical information dynamic network represent the dynamic coupling relationship between variables. The second construction module is used to construct and minimize a global network energy function based on the physical information dynamic network, wherein the low energy state of the global network energy function corresponds to the coordinated working mode of the system. The first generation module is used to calculate and apply minimum information perturbation based on the requirement of minimizing the global network energy function, and generate control instructions for the actuator; The second generation module is used to reverse synthesize virtual physical constraints based on the long-term evolution topology of the physical information dynamic network, and generate adaptive control strategy patches.
9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the power domain controller configuration method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the power domain controller configuration method as described in any one of claims 1-7.