Satellite energy system intelligent control method and storage medium
Through the combination of the large language sub-model and intelligent decision-making framework, real-time data analysis and optimization control of satellite energy systems are realized, and the problem of balanced multi-task needs in dynamic space environments is solved, the flexibility and stability of the system are improved, and the risk of failure is reduced.
Patent Information
- Application Number
- CN202510813906.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing satellite energy management systems cannot efficiently balance multi-task requirements in dynamic space environments, lack flexibility, are difficult to cope with complex task changes and fluctuations in energy demand, and lack in-depth analysis of data, resulting in an increase in the risk of system failure.
The large language sub-model is used to perform semantic understanding and feature correlation analysis, and combined with the intelligent decision-making framework to generate state data and action data. Through the reward function optimization strategy, real-time control of satellite energy systems is achieved, including dynamic adjustment of battery status, solar panel output, load equipment power consumption and environmental conditions.
It realizes efficient management and optimization of satellite energy systems in dynamic space environments, improves the adaptability and stability of the system, reduces the risk of failure, and ensures the smooth completion of the mission.
Smart Images

Figure CN120335379A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of satellite energy systems, and in particular, to an intelligent control method and storage medium for a satellite energy system. Background Art
[0002] With the rapid development of satellite technology and the increasing complexity of space missions, the management of satellite energy systems has become a key link to ensure the stable operation of satellites and the successful completion of missions. Satellite energy systems need to efficiently manage energy resources in a highly dynamic and changing space environment to meet complex mission requirements and environmental challenges.
[0003] In related technologies, most traditional satellite energy management systems adopt rule-based control strategies or simple optimization algorithms. These methods can meet basic requirements in specific scenarios, but in the face of complex space environments and dynamic missions, they often show problems such as insufficient flexibility and poor adaptability, and are unable to effectively respond to various emergencies and energy demand fluctuations; and due to the lack of the ability to deeply analyze data, existing systems often have difficulty accurately predicting potential failures, which increases the risk of system failures; in addition, traditional methods often ignore the dynamic changes of environmental conditions and mission requirements when formulating optimization strategies, resulting in low energy utilization efficiency and the inability to fully exert the system performance.
[0004] Currently, there is no effective solution to the problems in related technologies that are difficult to adapt to dynamic space environments and cannot efficiently balance multi-mission requirements. Summary of the Invention
[0005] Embodiments of this application provide an intelligent control method, device, system, electronic device, and storage medium for a satellite energy system to at least solve the problems in related technologies that are difficult to adapt to dynamic space environments and cannot efficiently balance multi-mission requirements.
[0006] In a first aspect, embodiments of this application provide an intelligent control method for a satellite energy system, including:
[0007] Obtain real-time operation data of the satellite energy system;
[0008] Input the real-time operation data into a trained satellite energy system environment model, and use the large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the real-time operation data, and extract key features;
[0009] Using the intelligent decision-making framework in the satellite energy system environment model, based on the key features, generate state data corresponding to the state space; based on the key features and the state data, generate action data corresponding to the action space, and based on the key features, the state data, and the reward function, generate a reward prediction value; wherein, the intelligent decision-making framework includes the state space, the action space, and the reward function;
[0010] Based on the action data and the reward prediction value, generate an optimization strategy; based on the optimization strategy, perform real-time control operations on the satellite energy system.
[0011] In some embodiments, the method further includes:
[0012] Based on the key features, combined with the preset satellite mission objectives and environmental constraints, generate optimization objectives and constraints;
[0013] Based on the optimization objectives, define the reward function;
[0014] Based on the constraints, define the state space and the action space.
[0015] In some embodiments, the state space includes the current battery charge state, the current output power of the solar panels, the current power consumption demand of the load devices, the current environmental conditions, the priority and real-time requirements of the satellite mission, and the health state of the energy system;
[0016] The action space includes the solar panel angle adjustment amount, the satellite attitude adjustment amount, and the energy scheduling strategy adjustment amount;
[0017] Each sub-objective in the reward function includes energy utilization efficiency, system stability, and mission completion.
[0018] In some embodiments, the generating an optimization strategy based on the action data and the reward prediction value includes:
[0019] With the goal of maximizing the reward prediction value, select the optimal action from the action data and generate an optimization strategy based on the optimal action.
[0020] In some embodiments, the method further includes:
[0021] Use the large language sub-model to deeply analyze the operating state of the satellite energy system and determine whether the current operating state of the satellite energy system is abnormal;
[0022] If the current operating state is an abnormal mode, generate a warning message and continue to use the intelligent decision-making framework to generate the state data, the action data, and the reward prediction value;
[0023] If the current operating state is not the abnormal mode, directly use the intelligent decision-making framework to generate the state data, the action data, and the reward prediction value.
[0024] In some embodiments, determining whether the current operating state of the satellite energy system is abnormal includes:
[0025] Generate a set of safe states based on preset environmental constraint conditions;
[0026] Determine whether the current operating state of the satellite energy system belongs to the set of safe states;
[0027] If the current operating state of the satellite energy system belongs to the set of safe states, the current operating state is the abnormal mode;
[0028] If the current operating state of the satellite energy system does not belong to the set of safe states, the current operating state is not the abnormal mode.
[0029] In some embodiments, the training process of the satellite energy system environment model includes:
[0030] Construct a simulation environment of the satellite energy system and obtain training data of the satellite energy system;
[0031] Input the training data into the initial environment model and run the initial environment model in the simulation environment; wherein, the initial environment model includes an initial large language sub-model and an optimization algorithm module;
[0032] Use the initial large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the training data, generate training key features, construct a large language loss function, and substitute the training key features into the large language loss function for calculation, and adjust the parameters of the initial large language sub-model with the goal of minimizing the calculation result of the large language loss function to obtain the trained large language sub-model;
[0033] Input the training key features into the optimization algorithm module, construct an optimization loss function, and adjust the parameters of the optimization algorithm module with the goal of minimizing the calculation result of the optimization loss function to obtain the trained intelligent decision-making framework;
[0034] Based on the large language sub-model and the intelligent decision-making framework, obtain the trained satellite energy system environment model.
[0035] In some embodiments, obtaining the training data of the satellite energy system includes:
[0036] Generate empirical data using the state transition function in the simulation environment; the state transition function is used to describe the state change of the satellite energy system under a given current state, execution action, and environmental random factors;
[0037] Obtain the training data based on the empirical data.
[0038] In some embodiments, the step of inputting the training key features into the optimization algorithm module, constructing an optimization loss function, and adjusting the parameters of the optimization algorithm module with the goal of minimizing the calculation result of the optimization loss function to obtain the trained intelligent decision-making framework includes:
[0039] The optimization algorithm module includes a policy network and a value network;
[0040] Based on the training key features, generate training state data corresponding to the state space; via the policy network, based on the training key features and the training state data, generate an action probability distribution, and select training action data corresponding to the action space based on the action probability distribution;
[0041] Via the value network, based on the training state data, the training action data, and the reward function, calculate the immediate reward value of the current training action data;
[0042] According to the training state data, the training action data, the immediate reward value, and the next training state data generated by the state transition function, calculate the result of the optimization loss function, and adjust the parameters of the optimization algorithm module with the goal of minimizing the result of the optimization loss function to obtain the trained intelligent decision-making framework.
[0043] In a second aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the satellite energy system intelligent control method as described in the first aspect above.
[0044] Compared with the related art, the satellite energy system intelligent control method provided by the embodiments of the present application realizes closed-loop control by obtaining satellite energy data in real time, parsing key features through a large language sub-model, generating state and action data, predicting rewards, and optimizing strategies, solves the problems of being difficult to adapt to the dynamic space environment and unable to efficiently balance multi-task requirements, and realizes the efficient management and optimization of the satellite energy system.
[0045] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0047] Figure 1 is a hardware structure block diagram of a terminal of the intelligent control method for a satellite energy system according to an embodiment of the present invention;
[0048] Figure 2 is a flowchart of the intelligent control method for a satellite energy system according to an embodiment of the present application;
[0049] Figure 3 is a flowchart of the intelligent control method for a satellite energy system according to a preferred embodiment of the present application;
[0050] Figure 4 is a framework diagram of a data preprocessing and feature extraction module in the intelligent control method for a satellite energy system according to a preferred embodiment of the present application;
[0051] Figure 5 is a framework diagram of state analysis and optimization requirement assessment in the intelligent control method for a satellite energy system according to a preferred embodiment of the present application;
[0052] Figure 6 is a framework diagram of an intelligent decision-making in the intelligent control method for a satellite energy system according to a preferred embodiment of the present application. Detailed Embodiments
[0053] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without creative efforts fall within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.
[0054] Reference to "embodiments" in the present application means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0055] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those of ordinary skill in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a limitation in quantity and may represent a singular or plural number. The terms "comprising", "including", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order of the objects.
[0056] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. Taking running on a terminal as an example, Figure 1 is the hardware structure block diagram of the terminal of the intelligent control method for the satellite energy system according to the embodiment of the present invention. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0057] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the intelligent control method of the satellite energy system in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0058] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0059] This embodiment provides an intelligent control method for a satellite energy system. Figure 2 is a flowchart of the intelligent control method for a satellite energy system according to an embodiment of the present application, as Figure 2 shown, and this process includes the following steps:
[0060] Step S201, obtain the real-time operation data of the satellite energy system;
[0061] Among them, the real-time operation data of the satellite energy system is collected in real time through sensors deployed on the satellite. The real-time operation data includes power system data (such as battery voltage, current, temperature, charge and discharge status, etc.), load system data (power consumption of each device, working status (such as communication module, sensor, thruster, etc.)), and satellite environment data (solar irradiance, space radiation dose, satellite surface temperature, etc.). The collected data is transmitted back to the ground station in real time through the on-board data transmission system. The real-time operation data collected by the sensors in this step covers three dimensions of power supply, load, and environment, providing complete input for subsequent model analysis, and through real-time data collection by the sensors, it can ensure timely monitoring and analysis of the state of the satellite energy system.
[0062] Step S202: Input the real-time operation data into the trained satellite energy system environment model. Use the large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the real-time operation data, and extract key features.
[0063] Among them, when inputting the real-time operation data into the already trained satellite energy system environment model, first perform data preprocessing on the real-time operation data. For example, cleaning (removing invalid values, duplicate values, and outliers to ensure data integrity and accuracy), denoising (eliminating noise in the data through filtering algorithms or statistical methods to improve the signal-to-noise ratio of the data), and normalization (unifying and standardizing data with different dimensions and magnitudes). A large language sub-model is embedded in the satellite energy system environment model. This sub-model uses the semantic understanding ability of the pre-trained large language model to perform semantic understanding and feature correlation analysis on the preprocessed real-time operation data, correlate historical data with real-time data, and extract key features that are crucial for the energy system state assessment and optimization decision-making. Key features are high-level abstract indicators extracted from real-time operation data that can reflect the core state of the system, integrating the dynamic changes of the power system, load devices, and environmental conditions, providing input for the intelligent decision-making framework and supporting the generation of optimization strategies. Key features include the health status of the power system (such as the degree of battery aging), the power consumption trend of the load system (such as periodic high-load periods), and the dynamic changes of environmental conditions, etc. The large language sub-model outputs key features for use by the subsequent intelligent decision-making framework. In this step, preprocessing operations such as cleaning, denoising, and normalization are performed before the data is input into the large language sub-model to improve data quality and provide a reliable basis for subsequent analysis; through the in-depth analysis of the real-time operation data by the large language sub-model, the state changes and potential faults of the energy system can be more accurately identified. Based on the extracted key features, the intelligent decision-making framework can generate optimization strategies faster, improving decision-making efficiency and accuracy.
[0064] Step S203: Use the intelligent decision-making framework in the satellite energy system environment model to generate state data corresponding to the state space based on the key features; generate action data corresponding to the action space based on the key features and state data, and generate a reward prediction value based on the key features, state data, and reward function; where the intelligent decision-making framework includes a state space, an action space, and a reward function.
[0065] Among them, the state space is the set of all parameters describing the real-time operating state of the satellite energy system, covering key dimensions such as energy supply and demand, environmental conditions, and mission requirements; the action space is the set of all optimization operations that the satellite energy system can execute, achieving efficient energy utilization by adjusting hardware or scheduling strategies; the reward function is used to quantify the quality of each action, comprehensively considering energy efficiency, system stability, and mission completion through mathematical formulas to guide the model to learn the optimal strategy. The intelligent decision-making framework perceives the environment through the state space, executes control through the action space, and evaluates the effect through the reward function, transforming the complex optimization problem of the satellite energy system into a learnable decision-making process. State data is generated based on the key features extracted in the above steps (such as power supply health, load trend, environmental dynamic changes), and the state data includes data such as the output power of the solar panel, the total power consumption demand of the load device, and the light intensity. Based on the fused key features and state data, candidate actions in the action space are output, that is, action data. The reward prediction value of the candidate action is calculated using the formula of the reward function, and the formula is as follows:
[0066]
[0067] In the above formula, the energy utilization efficiency (η) is determined by the effective energy utilization rate, the stability (σ) is calculated by the degree of parameter deviation from the safe range, and the mission completion (T) is weighted by the priority, 、 、 are weight coefficients used to balance the priorities of different optimization objectives.
[0068] Specifically, if the key features are: identifying the correlation between "battery aging leading to capacity decline" and "light interruption in the shadow area"; predicting that the light cannot be restored within the next 30 minutes and battery power supply is required. The generated action data are: Action 1 adjusts the satellite attitude and tries to get out of the shadow area in advance (but may increase energy consumption); Action 2 suspends low-priority tasks and allocates all available energy (200W) to communication tasks; Action 3 reduces the voltage of the load device (implied action) to reduce instantaneous power consumption. Next, the reward function is evaluated: If Action 1 successfully gets out of the shadow area, the energy utilization efficiency η increases, but the attitude adjustment consumes energy (the stability increases), and the reward may be R = 0.7; Action 2 directly ensures mission completion (T = 0.9), but the risk of battery depletion is high (σ = 0.5), and the reward R = 0.85; Action 3 has the best stability (σ = 0.2), but the mission completion degree decreases (T = 0.6), and the reward R = 0.65.
[0069] In this step, the comprehensive energy efficiency, system stability, and task completion are integrated through the reward function to scientifically balance different optimization objectives and avoid system imbalance caused by single - indicator optimization. An action strategy is generated in real - time to ensure rapid response in a dynamic space environment (such as sudden changes in light and load fluctuations). By abstracting the complex energy system optimization problem into an intelligent decision - making problem and defining a clear state space, action space, and reward function, the problem becomes more structured and solvable.
[0070] In step S204, an optimization strategy is generated based on the action data and the predicted reward value. Based on the optimization strategy, real - time control operations are performed on the satellite energy system.
[0071] Among them, through the action data (such as adjusting the angle of the solar panel, satellite attitude, etc.) and the corresponding predicted reward value, the action with the highest reward value is selected as the optimal strategy. For example, if the predicted reward value R = 0.85 of the action "pause low - priority tasks" is higher than other actions, it is determined as the optimization strategy. The abstract action is converted into specific control instructions (such as "deactivate the observation device"), and the strategy is executed through the satellite control module for real - time control operations. For example: adjust the angle of the solar panel to maximize light capture; dynamically allocate energy to prioritize high - priority tasks (such as communication equipment). In addition, the system state after execution (such as battery charge recovery, task completion) can be monitored to update the model parameters to optimize subsequent strategies. Finally, the real - time operating state of the satellite energy system, the execution effect of the optimization strategy, and the key performance indicators can be visually displayed through the human - machine interaction interface, and a detailed optimization report can be generated to provide decision - making support for the operation and maintenance personnel. Through real - time control operations in this step, dynamic adjustment can be made according to the real - time state and task requirements of the satellite energy system to ensure the efficient and stable operation of the system.
[0072] Through the above steps, this application first obtains real-time operation data, uses the large language sub-model to perform semantic parsing and feature association on the original data, breaks through the limitations of traditional threshold judgment or simple statistical models, and captures the deep associations of multi-source heterogeneous data through context understanding ability (such as performing semantic-level association analysis on the temperature fluctuations of the solar wings and the changes in load power consumption). Subsequently, a dynamic state space is constructed in the intelligent decision-making framework, mapping environmental parameters, task priorities, etc. into high-dimensional feature vectors. Compared with traditional look-up table methods or fixed rule bases, it can generate a refined action space including "adjusting the angle of the solar wings", "dynamic power allocation", etc. in real time, and quantify the policy value through a reward function (such as energy utilization rate × task completion weight coefficient). Finally, in the policy generation stage, algorithms such as Monte Carlo tree search are used to dynamically select the action sequence with the highest reward prediction value, such as giving priority to ensuring the power supply of the communication module while enabling the backup payload. Compared with traditional priority queue algorithms, it can complete the calculation of the multi-objective Pareto optimal solution in milliseconds. This method effectively solves the conflict between the lag in adapting to the dynamic environment and resource allocation, thus solving the technical problems that the satellite energy system is difficult to adapt to environmental changes and cannot efficiently balance multi-task requirements in the dynamic space environment. Its essence is to transform the discrete control problem into an optimal path search problem in a continuous decision space through a reinforcement learning framework driven by a large language model, realizing the full-link intelligence of environmental perception, feature extraction, and policy generation.
[0073] In some of these embodiments, the method further includes:
[0074] Based on the key features, combined with the preset satellite mission objectives and environmental constraint conditions, generate optimization objectives and constraint conditions;
[0075] Based on the optimization objectives, define a reward function;
[0076] Based on the constraint conditions, define the state space and the action space.
[0077] Among them, based on the key features (such as the battery health status E health , the load power consumption trend P load , the dynamic environmental changes T env ), the preset satellite mission objectives (mission priority R task ), and environmental constraints (such as light intensity, battery capacity limit), generate optimization objectives and constraint conditions. The optimization objectives include maximizing energy efficiency, ensuring system stability, and meeting mission requirements; the constraint conditions include power output capacity, load power consumption requirements, and dynamic environmental changes.
[0078] Specifically, the optimization objectives are as follows:
[0079] Maximize energy utilization efficiency: By optimizing power output and load distribution, reduce energy waste and improve energy utilization rate. The objective function can be expressed as:
[0080] ;
[0081] In the above formula, η is the energy utilization efficiency, P used is the effectively utilized energy, and P total is the total available energy.
[0082] Guarantee system stability: Ensure that the power supply system and load devices operate within a safe range to avoid system failures caused by insufficient energy or overload. The objective function can be expressed as:
[0083] ;
[0084] In the above formula, σ is the system stability index, S i is the actual value of the i-th system parameter, S target,i is the target value, and S max,i and S min,i and are the upper and lower safety limits of the parameter respectively.
[0085] Meet satellite mission requirements: Dynamically adjust energy allocation according to the priorities and real-time requirements of satellite missions to ensure the smooth completion of critical missions. The objective function can be expressed as:
[0086] ;
[0087] In the above formula, T is the mission completion degree, is the priority weight of the j-th mission, and C j is the mission completion status (1 indicates completed, 0 indicates not completed).
[0088] The constraint conditions are as follows:
[0089] Output capacity of the power supply system: Considering limitations such as battery capacity and output power of solar panels, ensure that the power supply system operates within a safe range. The constraint condition can be expressed as:
[0090]
[0091] Among them, is the output power of the power supply system, and are the lower and upper limits of the output power respectively.
[0092] Power consumption requirements of load devices: Reasonably allocate energy according to the power consumption characteristics and mission priorities of each device to avoid local overload or insufficient energy. The constraint condition can be expressed as:
[0093]
[0094] Among them, is the power consumption of the th load device, and
[0095] Dynamic changes in environmental conditions: Considering the impact of environmental factors such as light intensity, temperature, and radiation on the energy system, dynamically adjust the optimization strategy. The constraint conditions can be expressed as:
[0096]
[0097] where is the impact factor of environmental conditions, and are the lower and upper limits of the impact factor respectively.
[0098] Real-time operating status of the satellite energy system: Based on real-time monitoring data, evaluate the system status and adjust the optimization strategy to ensure the stability and efficiency of the system. The constraint conditions can be expressed as:
[0099]
[0100] where is the current operating status, and
[0101]
[0102]
[0103]
[0104] is the set of safe states.
[0105] By comprehensively considering the above optimization objectives and constraint conditions, a scientific and reasonable energy management strategy can be formulated to achieve the efficient and stable operation of the satellite energy system, convert the satellite mission objectives described in natural language (such as "giving priority to ensuring scientific experiment payloads") into mathematical constraints, and make the optimization direction strictly align with the actual needs of the system, avoiding the problem of fuzzy objectives in the traditional threshold control method.
[0105] In some of these embodiments, the state space includes the current battery charge state, the current output power of the solar panels, the current power consumption requirements of the load devices, the current environmental conditions, the priority and real-time requirements of the satellite mission, and the health state of the energy system;
[0106]
[0107] where is the current battery charge state, is the current output power of the solar panel, is the current power consumption demand of the load device, is the current environmental conditions such as ambient temperature and light intensity, is the priority and real-time requirements of the satellite mission, is the health status of the energy system.
[0108] The action space A is used to describe the optimization control operations on the satellite energy system and is defined as:
[0109]
[0110] where, is the adjustment amount of the solar panel angle, is the adjustment amount of the satellite attitude, is the adjustment of the energy scheduling strategy (such as task priority adjustment, etc.).
[0111] The reward function R is used to evaluate the effect of each action. The sub-goals include energy utilization efficiency, system stability, and mission completion. The reward function is defined as:
[0112]
[0113] where, is the energy utilization efficiency and its calculation formula is ; σ is the system stability index and its calculation formula is ; is the mission completion and its calculation formula is ; , , are the weight coefficients used to balance the priorities of different optimization goals.
[0114] The state space is generated by a large language sub-model through semantic understanding and feature extraction of real-time operation data, integrating multi-dimensional data such as power supply, load, and environment; the action space is generated by the policy network of deep reinforcement learning based on the state space and the reward function to generate candidate action sequences; the reward function adjusts the weights of each sub-goal in real time according to environmental changes (such as sudden failures and task urgency). In this step, the reward function comprehensively considers energy efficiency, system stability, and task completion to avoid imbalance caused by single-goal optimization. The reward function quantifies complex goals into learnable numerical values to generate optimal actions; the state space is updated in real time, and the action space responds quickly to ensure the real-time adaptability of the system. The action space strictly follows the power output capacity and load demand limits to avoid risks of overload or insufficient energy, fuses real-time data with historical features to support accurate decision-making, scientifically balances efficiency, stability, and task requirements, and reduces manual intervention, improving the satellite's survival ability and mission success rate in complex space environments.
[0115] In some of these embodiments, an optimized policy is generated based on action data and predicted reward values, including:
[0116] With the goal of maximizing the predicted reward value, select the optimal action from the action data and generate an optimized policy based on the optimal action.
[0117] Among them, based on the action data (candidate actions in the action space) and their corresponding predicted reward values (such as the reward for "adjusting the solar panel angle" R = 0.6 and the reward for "pausing low-priority tasks" R = 0.85), through the value network of deep reinforcement learning, select the action with the highest reward value as the optimal policy. For example, when the satellite enters the shadow area, "pausing low-priority tasks" is selected because of its higher reward value. Then map the optimal action to specific control instructions, such as pausing the observation device (saving 200W power consumption); adjusting the solar panel angle (from 30° to 45° to maximize light capture). Send instructions through the satellite control module (such as driving the motor to adjust the solar panel angle). In addition, parameters such as battery power and task completion after execution can be tracked in real time, and the execution results are fed back to the model to update the parameters of the policy network and improve the accuracy of subsequent decision-making. In this embodiment, by selecting the action with the highest reward value, energy waste is reduced, and scenarios such as sudden changes in light and load fluctuations are responded to in real time.
[0118] In some of these embodiments, the method further includes:
[0119] Utilize a large language sub-model to deeply analyze the operating state of the satellite energy system to determine whether the current operating state of the satellite energy system is abnormal;
[0120] If the current operating state is an abnormal mode, generate a warning message and continue to use the intelligent decision-making framework to generate state data, action data, and predicted reward values;
[0121] If the current operating state is not the abnormal mode, directly utilize the intelligent decision-making framework to generate state data, action data, and reward prediction values.
[0122] Among them, utilize the large language sub-model to perform semantic correlation analysis on the preprocessed real-time operating data (power supply system, load system, environmental data) and historical data, extract key features, and judge whether the current state is abnormal based on feature thresholds or dynamic rules (such as the battery temperature exceeding the safe range, the load power consumption not matching the task priority). If it is determined to be abnormal, generate a warning message (such as "Battery temperature exceeds the limit, risk level: high"), and trigger the intelligent decision-making framework to generate state data, action data, and reward prediction values; if the state is normal, directly enter the intelligent decision-making framework to generate optimization strategies. In this embodiment, through the semantic correlation ability of the large language model, quickly identify abnormal modes such as battery aging and sudden load increase, generate early warnings in advance (such as "Light interruption risk warning"), and avoid system crashes or task interruptions; in the abnormal state, the intelligent decision-making framework combines the warning information to generate targeted strategies (such as pausing non-critical tasks) to ensure the stable operation of core functions; in the normal state, optimize energy allocation based on real-time state data to balance multi-task requirements such as efficiency, stability, and task needs.
[0123] In some of these embodiments, determining whether the current operating state of the satellite energy system is abnormal includes:
[0124] Generate a set of safe states based on preset environmental constraint conditions;
[0125] Determine whether the current operating state of the satellite energy system belongs to the set of safe states;
[0126] If the current operating state of the satellite energy system belongs to the set of safe states, the current operating state is the abnormal mode;
[0127] If the current operating state of the satellite energy system does not belong to the set of safe states, the current operating state is not the abnormal mode.
[0128] Among them, based on preset environmental constraint conditions (such as battery voltage range, temperature threshold, load power consumption upper limit, etc.), define the set of safe states. For example, the safe range of battery voltage is 3.3V ≤ V battery ≤ 4.2V, and the safe range of temperature is T env ∈[-20°C, 50°C]. Combine the safe ranges of multi-dimensional parameters into the set of safe states State safe in the high-dimensional space. Based on the state data, obtain the current operating state State current , and the current operating state includes the battery power P battery , the solar output power P solar, ambient temperature T env etc. If State current State safe , it is determined as an abnormal mode and a warning is triggered; if State current ∈State safe , it is determined as a normal mode and directly enters the optimization decision-making process. In addition, the large language sub-model can perform auxiliary analysis, combine historical data with real-time features, identify latent anomalies (such as the capacity decay trend caused by battery aging), and enhance the comprehensiveness of the determination. In the abnormal mode, specific warning content (such as "Battery voltage exceeds the limit: the current value is 4.3V, and the safe range is 3.3V–4.2V") is generated and pushed to the operation and maintenance interface. In this embodiment, by comparing the preset constraint conditions with the real-time state, parameter out-of-bounds (such as temperature over-limit, voltage anomaly) is quickly identified, avoiding the lag of the traditional rule engine, comprehensively considering the safety ranges of multiple parameters such as power supply, load, and environment, and avoiding system-level cascading failures.
[0129] In some of these embodiments, the training process of the satellite energy system environment model includes:
[0130] Construct a simulation environment for the satellite energy system and obtain training data for the satellite energy system;
[0131] Input the training data into the initial environment model and run the initial environment model in the simulation environment; among them, the initial environment model includes an initial large language sub-model and an optimization algorithm module;
[0132] Use the initial large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the training data, generate training key features, construct a large language loss function, and substitute the training key features into the large language loss function for calculation. With the goal of minimizing the calculation result of the large language loss function, adjust the parameters of the initial large language sub-model to obtain a trained large language sub-model;
[0133] Input the training key features into the optimization algorithm module, construct an optimization loss function, and with the goal of minimizing the calculation result of the optimization loss function, adjust the parameters of the optimization algorithm module to obtain a trained intelligent decision-making framework;
[0134] Based on the large language sub-model and the intelligent decision-making framework, obtain a trained satellite energy system environment model.
[0135] Among them, based on the physical models of the satellite energy system (such as power supply system, load system, environmental conditions), a high-fidelity simulation environment is constructed to simulate dynamic scenarios such as light intensity fluctuations, load demand changes, and equipment aging. Preset tasks (such as orbit adjustment, emergency communication tasks) are run in the simulation environment, and real-time data such as the voltage / current of the power supply system, the power consumption of load devices, and the environmental temperature are collected to form training data. The initial environment model includes an initial large language sub-model (Large Language Model, LLM large language model) and an optimization algorithm module (deep reinforcement learning network). The training data is input into the initial environment model. The initial large language sub-model performs preliminary feature extraction on it, and the optimization algorithm generates an initial action strategy. The simulation environment updates the state according to the action and feedbacks the reward value. Specifically, the initial large language sub-model performs semantic association analysis on the training data and extracts key features (such as battery health, load power consumption trend). A large language loss function is constructed to calculate the difference between the output features of the initial large language sub-model and the real state of the simulation environment:
[0136]
[0137] In the above formula, y i is the real feature label provided by the simulation environment (such as the battery capacity attenuation rate).
[0138] Then, the parameters of the initial large language sub-model are adjusted by the gradient descent method to minimize the loss function and improve the feature extraction accuracy, and a trained large language sub-model is obtained.
[0139] The training key features extracted by the initial large language sub-model (such as E health 、P load ) are input into the optimization algorithm module, an optimization loss function is constructed, and the parameters of the optimization loss function are adjusted through experience replay and gradient descent to maximize the cumulative reward, and a trained intelligent decision-making framework is obtained. The trained large language sub-model is combined with the intelligent decision-making framework. The large language sub-model is responsible for real-time feature extraction, and the intelligent decision-making framework generates an action strategy based on the features, and a trained satellite energy system environment model is obtained.
[0140] In this embodiment, the simulation environment can reproduce the high dynamics of the space environment (such as the interruption of light in the shadow area), ensure that the training data covers various extreme situations, avoid the high costs of real satellite experiments, quickly generate large-scale labeled data, verify the feasibility of the initial model through the simulation environment, identify early defects, and prevent satellite system failures caused by incorrect actions; the large language sub-model mines deep associations in the data through semantic understanding (such as the relationship between battery temperature and load fluctuations), improves the quality of features, and the trained large language sub-model can identify hidden anomalies (such as voltage fluctuations caused by battery aging); the trained optimization algorithm can quickly generate actions to respond to dynamic environmental changes; the semantic understanding of the large language sub-model and the decision-making ability of the optimization algorithm complement each other, improving the overall model performance, enabling the model to adapt to unknown environments (such as sudden radiation interference), and ensuring the stable operation of the satellite in complex scenarios.
[0141] In some of these embodiments, obtaining training data for the satellite energy system includes:
[0142] Using the state transition function in the simulation environment to generate empirical data; the state transition function is used to describe the state change of the satellite energy system under a given current state, executed action, and environmental random factors;
[0143] Based on the empirical data, obtaining the training data.
[0144] Among them, a high-fidelity simulation environment including a power system, a load system, and environmental dynamics is constructed to simulate scenarios such as changes in light intensity, load fluctuations, and equipment aging. The state transition function is s t+1 =f(s t , a t , e t ) where s t is the current state data (such as battery power, load demand), a t is the executed action data (such as adjusting the angle of the solar panel), and e t is the environmental random factor (such as sudden light changes, equipment noise). Running the model in the simulation environment, recording the state, action, reward, and next state at each step to form an empirical data tuple (s t , a t , r t , s t+1 ). Storing the empirical data in the empirical pool D to support subsequent random sampling to break data correlations. Sampling data in batches from the empirical pool D as training data for updating the parameters of the large language sub-model and the optimization algorithm module. The training data includes state data s t , action data a t , immediate reward value r t , and next state data s t+1 . This embodiment introduces the environmental random factor e tSimulate uncertainties such as light fluctuations and equipment noise to enhance the model's adaptability to dynamic environments; sample data in batches from the experience pool to improve data utilization rate, accelerate model convergence, and use random sampling to reduce data correlation and enhance model generalization, enabling stable decision-making in unknown scenarios (such as sudden radiation interference).
[0145] In some of these embodiments, the training key features are input into the optimization algorithm module to construct an optimization loss function, and the parameters of the optimization algorithm module are adjusted with the goal of minimizing the calculation result of the optimization loss function to obtain a trained intelligent decision-making framework, including:
[0146] The optimization algorithm module includes a policy network and a value network;
[0147] Based on the training key features, generate training state data corresponding to the state space; via the policy network, based on the training key features and the training state data, generate an action probability distribution, and select training action data corresponding to the action space based on the action probability distribution;
[0148] Via the value network, based on the training state data, the training action data, and the reward function, calculate the immediate reward value of the current training action data;
[0149] According to the training state data, the training action data, the immediate reward value, and the next training state data generated by the state transition function, calculate the result of the optimization loss function, and adjust the parameters of the optimization algorithm module with the goal of minimizing the result of the optimization loss function to obtain a trained intelligent decision-making framework.
[0150] Among them, the optimization algorithm module includes a policy network (Policy Network, π) and a value network (Value Network, Q). Input the training key features and the state data into the policy network, and output the action probability distribution. The action probability distribution represents the probability of each action in the action space being selected given the current state data. For example: action probability = [0.6, 0.3, 0.1] corresponds to A = {Δθ solar , ΔS schedule , Δθ attitude}, Δθ solar is the adjustment amount of the solar panel angle, Δθ attitude is the adjustment amount of the satellite attitude, ΔS schedule is the adjustment of the energy scheduling policy (such as task priority adjustment, etc.), and select an action according to the probability distribution (such as ). Evaluate the long-term rewards of the value network actions to guide the optimization direction of the policy network. Input the state data and action data into the value network, calculate the immediate reward value based on the reward function, and quantify the action effect. For example, the reward for adjusting the angle of the solar panel is R = 0.85, which is better than other actions. The loss functions of the policy network and the value network in the optimization algorithm module are respectively defined as:
[0151] ;
[0152] In the above formula, L π is the loss function of the policy network, and the goal is to maximize the expected value of the long-term cumulative reward; E is the expectation; Q(s t , a t ; θ Q ) is the state-action value function output by the value network, representing the expected long-term reward value of executing action a t under state s t ; s t is the state data at time step t; π(s t ; θ π ) is the action a π generated by the policy network according to state s t under parameters θ t ; θ π are the parameters of the policy network, which are optimized by the gradient descent method; θ Q are the parameters of the value network, which are fixed for the training of the policy network (to avoid fluctuations in the target value).
[0153] ;
[0154] In the above formula, L Q is the loss function of the value network, and the goal is to make the predicted value approach the true target value (minimize the mean square error); N is the amount of training data sampled from the experience pool; Q(s i , a i ; θ Q ) is the predicted value of the value network for state s Q and action a i under parameters θ i ; r i is the immediate reward value obtained after executing action a i ; γ is the discount factor (0 ≤ γ ≤ 1), which balances the importance of current rewards and future rewards (e.g., γ = 0.9 means more emphasis on recent rewards); is the action a i+1 that maximizes the output of the target value network (with parameters θ Q′ ) in the next state s i+1 ; θ Q′are the parameters of the target value network, which are synchronized from θ Q periodically (during the stable training process).
[0155] The optimized loss function includes the loss function of the policy network and the loss function of the value network. Taking the minimization of the optimized loss function result as the goal, the network parameters θ π and θ Q are adjusted through the gradient descent method, and the data correlation is broken by combining experience replay. In this embodiment, a diverse set of actions are generated by training the policy network, the value network is trained for accurate evaluation, and efficiency, stability, and task requirements are dynamically weighed to balance exploration and exploitation; and the model convergence is accelerated through batch training and experience replay, reducing the training time. The optimized model can adapt to unknown dynamic environments.
[0156] The following describes and illustrates the embodiments of the present application through preferred embodiments.
[0157] Figure 3 is a flowchart of the intelligent control method for a satellite energy system according to a preferred embodiment of the present application. As Figure 3 shown, the intelligent control method for the satellite energy system includes the following steps:
[0158] Step S1: Construct an environmental model of the satellite energy system.
[0159] Construct an environmental model of the satellite energy system, including real-time operating power system data (such as battery voltage, current, temperature, etc.), load system data (such as the power consumption, working status of each device, etc.), and satellite environmental data (such as light intensity, temperature, radiation, etc.). Sensors are used to collect data to form a complete description of the operating state of the satellite energy system.
[0160] The input of the satellite energy system environmental model is the real-time operating data of the satellite energy system, which mainly includes power system data, load system data, and satellite environmental data, etc. Power system data: including but not limited to battery voltage, current, temperature, charge and discharge status, output power of solar panels, etc., which are used to describe the real-time operating state of the satellite power system. Load system data: including but not limited to the power consumption, working status, task priority of each device (such as communication devices, sensors, computing units, etc.), which are used to describe the real-time requirements of the satellite load system. Satellite environmental data: including but not limited to light intensity, temperature, radiation level, orbital position, etc., which are used to describe the external environmental conditions of the satellite operation. By collecting the above multi-dimensional data, a complete satellite energy system environmental model is constructed, providing a comprehensive data basis for subsequent intelligent monitoring and optimization.
[0161] Step S2: Establish a data preprocessing and feature extraction module.
[0162] Figure 4It is a framework diagram of the data preprocessing and feature extraction module in the intelligent control method of the satellite energy system in the preferred embodiment of this application. As Figure 4 shown, the satellite energy system data collected is cleaned, denoised, and normalized, and key features are extracted. The large language model is used to perform semantic understanding and feature correlation analysis on historical data and real-time data to construct high-quality data input.
[0163] The data preprocessing and feature extraction module is established based on the real-time operation data of the satellite energy system, mainly used to clean, denoise, and normalize the collected power system data, load system data, and satellite environment data, and extract key features to provide high-quality data input for subsequent large language model analysis.
[0164] Data cleaning: Remove invalid values, duplicate values, and outliers in the collected data to ensure the integrity and accuracy of the data. Data denoising: Eliminate noise in the data through filtering algorithms or statistical methods to improve the signal-to-noise ratio of the data. Data normalization: Standardize data with different dimensions and magnitudes uniformly for subsequent analysis and model processing. Feature extraction: Use the semantic understanding ability of the large language model to extract key features (such as the health status of the power system, the power consumption trend of the load system, the dynamic changes in environmental conditions, etc.) from the preprocessed data to provide a high-quality data basis for intelligent monitoring and optimization. Through the above processing, the data preprocessing and feature extraction module can significantly improve the usability of the data and provide reliable support for subsequent large language model analysis and optimization decisions.
[0165] Step S3: Analyze the energy system status and optimization requirements.
[0166] Figure 5 It is a framework diagram of the status analysis and optimization requirement assessment in the intelligent control method of the satellite energy system in the preferred embodiment of this application. As Figure 5 shown, based on the large language model, a deep analysis of the operation status of the satellite energy system is carried out to identify abnormal patterns, predict potential faults, and evaluate the optimization requirements of the current energy system. Combining satellite mission objectives and environmental constraints, optimization objectives and constraint conditions are formulated.
[0167] In the energy system status analysis and optimization requirement assessment, the optimization objectives are mainly to maximize energy utilization efficiency, ensure system stability, and meet satellite mission requirements. The constraint conditions consider the output capacity of the power system, the power consumption requirements of load devices, the dynamic changes in environmental conditions, and the real-time operation status of the satellite energy system, etc.
[0168] The optimization objectives are as follows:
[0169] Maximize energy utilization efficiency: By optimizing power output and load distribution, reduce energy waste and improve energy utilization rate. The objective function can be expressed as:
[0170]
[0171] Among them, is the energy utilization efficiency, is the effectively utilized energy, is the total available energy.
[0172] Ensure system stability: Ensure that the power supply system and load devices operate within a safe range to avoid system failures caused by insufficient energy or overload. The objective function can be expressed as:
[0173]
[0174] Among them, σ is the system stability index, is the actual value of the i-th system parameter, is the target value, and are the upper and lower safety limits of the parameter respectively.
[0175] Meet satellite mission requirements: Dynamically adjust energy allocation according to the priority and real-time requirements of satellite missions to ensure the smooth completion of critical missions. The objective function can be expressed as:
[0176]
[0177] Among them, is the mission completion degree, is the priority weight of the j-th mission, is the mission completion status (1 indicates completed, 0 indicates not completed).
[0178] The constraint conditions are as follows:
[0179] Output capacity of the power supply system: Considering limitations such as battery capacity and output power of solar panels, ensure that the power supply system operates within a safe range. The constraint condition can be expressed as:
[0180]
[0181] Among them, is the output power of the power supply system, and are the lower and upper limits of the output power respectively.
[0182] Power consumption requirements of load devices: Reasonably allocate energy according to the power consumption characteristics and mission priorities of each device to avoid local overload or insufficient energy. The constraint condition can be expressed as:
[0183]
[0184] Among them, is the power consumption of the th load device, and is the currently available energy.
[0185] Dynamic changes in environmental conditions: Considering the impact of environmental factors such as light intensity, temperature, and radiation on the energy system, dynamically adjust the optimization strategy. The constraint conditions can be expressed as:
[0186]
[0187] Among them, is the impact factor of environmental conditions, and are the lower and upper limits of the impact factor respectively.
[0188] Real-time operating status of the satellite energy system: Based on real-time monitoring data, evaluate the system status and adjust the optimization strategy to ensure the stability and efficiency of the system. The constraint conditions can be expressed as:
[0189]
[0190] Among them, is the current system status, and is the set of safe states.
[0191] By comprehensively considering the above optimization objectives and constraint conditions, a scientific and reasonable energy management strategy can be formulated to achieve the efficient and stable operation of the satellite energy system.
[0192] Step S4: Convert the energy system optimization problem into an intelligent decision-making problem.
[0193] Figure 6 is the intelligent decision-making framework diagram in the intelligent control method of the satellite energy system in the preferred embodiment of this application. As Figure 6 shown, convert the monitoring and optimization problems of the satellite energy system into an intelligent decision-making problem based on a large language model. Define the state space and action space, and construct an intelligent decision-making framework. The state space includes the power system state, the load system state, and environmental conditions; the action space includes power output adjustment, load mode optimization, and energy scheduling strategy.
[0194] Model the energy system optimization problem, convert the monitoring and optimization problems of the satellite energy system into an intelligent decision-making problem based on a large language model, define the state space, action space, and reward function, and construct an intelligent decision-making framework.
[0195] The state space S is used to describe the real-time operating status of the satellite energy system and is defined as:
[0196]
[0197] Among them, is the current battery power status, is the current output power of the solar panel, is the current power consumption demand of the load device, is environmental conditions such as environmental temperature and light intensity, is the priority and real-time requirements of the satellite mission, E health is the health status of the energy system (such as battery aging degree, equipment failure risk, etc.).
[0198] The action space A is used to describe the optimization control operations on the satellite energy system and is defined as:
[0199]
[0200] Among them, is the adjustment amount of the solar panel angle, is the adjustment amount of the satellite attitude, is the adjustment of the energy scheduling strategy (such as task priority adjustment, etc.).
[0201] Furthermore, according to the current lighting conditions and satellite position, calculate the optimal solar panel angle , and the adjustment angle of the solar panel is defined as:
[0202]
[0203] Among them, is the current solar panel angle.
[0204] According to the task requirements and energy optimization goals, calculate the optimal satellite attitude , and the adjustment amount of the satellite attitude is defined as:
[0205]
[0206] Among them, is the current satellite attitude.
[0207] According to the current energy system status and task requirements, calculate the optimal energy scheduling strategy , and the adjustment of the energy scheduling strategy is defined as:
[0208]
[0209] Among them, is the current energy scheduling strategy.
[0210] The reward function R is used to evaluate the effect of each action. Considering the energy utilization efficiency, system stability, and task completion degree comprehensively, the reward function is defined as:
[0211]
[0212] Among them, is the energy utilization efficiency, and its calculation formula is ; σ is the system stability index, and its calculation formula is ; is the task completion degree, and its calculation formula is ; , , are the weight coefficients, which are used to balance the priorities of different optimization objectives.
[0213] Construction of the intelligent decision-making framework: Based on the semantic understanding and reasoning capabilities of the large language model, the state space, action space, and reward function are integrated into the intelligent decision-making framework. Through the analysis of historical data and real-time data by the large language model, the optimal energy management strategy is generated. The core of the decision-making framework is to model through the Markov decision process (MDP) to find the optimal strategy that maximizes the cumulative reward:
[0214]
[0215] Among them, is the discount factor, which is used to balance the importance of current rewards and future rewards.
[0216] Step S5: Build a network structure that integrates the large language model and the optimization algorithm.
[0217] Design a network structure that integrates the large language model and optimization algorithms (such as deep reinforcement learning, heuristic algorithms, etc.) to generate the reward function. The reward function comprehensively considers the energy utilization efficiency, system stability, and task completion degree, and improves the decision-making efficiency of the optimization algorithm through the semantic understanding and reasoning capabilities of the large language model.
[0218] The design of the network structure that integrates the large language model and the optimization algorithm includes the embedding of the large language model and the optimization algorithm to achieve intelligent monitoring and optimization of the satellite energy system.
[0219] Embedding the large language model: Embed the pre-trained large language model into the optimization framework and use its powerful semantic understanding and reasoning capabilities to deeply analyze the real-time operation data of the satellite energy system. The input of the large language model includes the state space S and historical data, and the output is the semantic understanding and feature extraction results of the current system state.
[0220] The output of the large language model is defined as:
[0221]
[0222] Among them, is the i-th key feature extracted by the large language model.
[0223] Optimization algorithm: Use deep reinforcement learning to optimize the satellite energy system. The input is the features extracted by the large language model and the state space S, and the output is the optimization action in the action space A.
[0224] The goal of the optimization algorithm is to find the optimal action through the policy network π and the value network Q:
[0225]
[0226]
[0227] Among them, and are the parameters of the policy network and the value network respectively.
[0228] Step S6: Train and optimize the model in the simulation environment.
[0229] Train the model in the constructed satellite energy system simulation environment, and continuously learn the optimization strategies in the historical data and real-time data through the large language model. Collect training experience, update the model parameters, and gradually improve the monitoring accuracy and optimization ability of the model. The simulation environment simulates the dynamic changes in the space environment, including fluctuations in lighting conditions, changes in load demand, and sudden fault events, to enhance the generalization ability of the model.
[0230] In the model training and optimization process, by training the network structure that integrates the large language model and the optimization algorithm in the simulation environment, collecting empirical data and updating the model parameters, the monitoring accuracy and optimization ability of the model are gradually improved.
[0231] Simulation environment construction: Construct a high-fidelity satellite energy system simulation environment to simulate the dynamic changes of the power system, load system, and environmental conditions. The input of the simulation environment includes power system data, load system data, and environmental data, and the output is the real-time operating state and optimization results of the system.
[0232] The state transition function of the simulation environment is defined as:
[0233]
[0234] Among them, is the current state, is the current action, is the environmental random factor.
[0235] Empirical data collection: Run the model in the simulation environment to collect empirical data , where: is the current state, is the current action, is the current reward, is the next state. The experience data is stored in the experience pool D for subsequent model training and parameter update.
[0236] Model training and parameter update: Randomly sample a batch of experience data from the experience pool D for training the large language model and the optimization algorithm. The training process includes large language model training and optimization algorithm training.
[0237] First, use the states and rewards in the experience data to update the parameters of the large language model , improving its semantic understanding and feature extraction capabilities. The loss function is defined as:
[0238]
[0239] where, is the target value.
[0240] Next, use the states, actions, and rewards in the experience data to update the parameters of the optimization algorithm. The loss functions of the optimization algorithm's deep reinforcement learning policy network and value network are respectively defined as:
[0241]
[0242]
[0243] Use the gradient descent method to update the model parameters. Regularly evaluate the performance of the model during the training process, including monitoring accuracy, optimization effect, and task completion rate. Adjust the model parameters and training strategy according to the evaluation results to further improve the performance of the model. Continuously repeat the training and evaluation process until the model converges or reaches the predetermined performance metrics, capable of efficiently monitoring the operating state of the satellite energy system and generating optimal energy management strategies.
[0244] Step S7: Output and execute the optimal energy management strategy.
[0245] After training is completed, output the intelligent monitoring and optimization strategy for the satellite energy system based on the large language model. Real-time control the satellite energy system according to the strategy, such as adjusting power output, optimizing load patterns, performing energy scheduling, etc., to achieve the efficient and stable operation of the energy system.
[0246] Optimal strategy generation and execution, based on the network structure integrating the trained large language model and the optimization algorithm, generate the optimal energy management strategy for the satellite energy system.
[0247] Using the trained large language model and optimization algorithm, according to the current state of the satellite energy system and historical data, generate the optimal energy management strategy . The generation process of the optimal strategy is defined as:
[0248]
[0249] Where: is the trained optimal strategy network, are its parameters.
[0250] The large language model performs semantic understanding and feature extraction on the current state, and the optimization algorithm generates the optimal action according to the extracted features. The generated actions include but are not limited to: the angle of the solar panel, the attitude of the satellite, the energy scheduling strategy, etc.
[0251] Apply the generated energy management strategy to the satellite energy system and perform corresponding control operations. The execution process includes: adjusting the angle of the solar panel, dynamically adjusting the angle of the solar panel according to the current light conditions and energy demand to ensure the stability and efficiency of energy supply; adjusting the attitude of the satellite, by adjusting the attitude of the satellite, making the solar panel always face the sun to maximize the energy capture efficiency; dynamically adjusting the energy scheduling strategy, according to the task priority and the state of the energy system, optimizing the energy allocation and scheduling strategy to ensure the smooth completion of key tasks.
[0252] Step S8: Visualize and display the monitoring and optimization results.
[0253] Visualize and display key information such as the real-time operating state, abnormal warning, fault prediction, and optimization strategy of the satellite energy system through the human-computer interaction interface, and generate a detailed optimization report to provide decision-making support for ground control personnel. The visualized display content includes the real-time operating state, the execution effect of the optimization strategy, and key performance indicators, which are intuitively displayed in the form of charts, dashboards, and dynamic simulations.
[0254] In this preferred embodiment, an intelligent monitoring and optimization method for satellite energy systems based on large language models is adopted to achieve efficient management and optimization of satellite energy systems. This method not only improves energy utilization efficiency and system stability, but also enhances the intelligence level and adaptive ability of the system through the semantic understanding and reasoning capabilities of large language models. In addition, the energy system optimization problem is transformed into an intelligent decision-making problem, and the combination of large language models and optimization algorithms is used to further improve decision-making efficiency and optimization accuracy. The efficient network structure and reward function design accelerate the convergence speed of the algorithm, and the trained model exhibits good generalization ability, ensuring stable performance in unknown space environments. The improvements in real-time performance and accuracy, the optimized utilization of energy, and the significantly reduced failure risk are the key advantages of this application, providing strong technical support for the execution of satellite missions and demonstrating broad application potential.
[0255] In addition, in combination with the satellite energy system intelligent control method in the above embodiments, the embodiments of the present application can be implemented by providing a storage medium. A computer program is stored on this storage medium; when the computer program is executed by a processor, any one of the satellite energy system intelligent control methods in the above embodiments is implemented.
[0256] Those skilled in the art should understand that the technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0257] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A satellite energy system intelligent control method, characterized in that, Including: Obtain the real-time operation data of the satellite energy system; Input the real-time operation data into the trained satellite energy system environment model, and use the large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the real-time operation data, and extract key features; Use the intelligent decision-making framework in the satellite energy system environment model to generate state data corresponding to the state space based on the key features; generate action data corresponding to the action space based on the key features and the state data, and generate a reward prediction value based on the key features, the state data and the reward function; wherein, the intelligent decision-making framework includes the state space, the action space and the reward function; Generate an optimization strategy based on the action data and the reward prediction value; Perform real-time control operations on the satellite energy system based on the optimization strategy.
2. The intelligent control method of the satellite energy system according to claim 1, characterized in that The method further includes: Generate an optimization goal and constraint conditions based on the key features, in combination with preset satellite mission goals and environmental constraint conditions; Define the reward function based on the optimization goal; Define the state space and the action space based on the constraint conditions.
3. The intelligent control method of the satellite energy system according to claim 1, characterized in that The state space includes the current battery power state, the current output power of the solar panel, the current power consumption demand of the load device, the current environmental conditions, the priority and real-time requirements of the satellite mission, and the health state of the energy system; The action space includes the solar panel angle adjustment amount, the satellite attitude adjustment amount, and the energy scheduling strategy adjustment amount; Each sub-goal in the reward function includes energy utilization efficiency, system stability, and mission completion.
4. The intelligent control method of the satellite energy system according to claim 1, wherein The generating an optimization strategy based on the action data and the reward prediction value includes: Taking maximizing the reward prediction value as the goal, select the optimal action from the action data, and generate an optimization strategy based on the optimal action.
5. The intelligent control method of the satellite energy system according to claim 1, characterized in that The method further includes: Use the large language sub-model to deeply analyze the operation state of the satellite energy system, and judge whether the current operation state of the satellite energy system is abnormal; If the current operation state is an abnormal mode, generate a warning message and continue to use the intelligent decision-making framework to generate the state data, the action data and the reward prediction value; If the current operation state is not an abnormal mode, directly use the intelligent decision-making framework to generate the state data, the action data and the reward prediction value.
6. The intelligent control method of the satellite energy system according to claim 5, characterized in that, The judging whether the current operation state of the satellite energy system is abnormal includes: Generate a safe state set based on preset environmental constraint conditions; Judge whether the current operation state of the satellite energy system belongs to the safe state set; If the current operation state of the satellite energy system belongs to the safe state set, the current operation state is an abnormal mode; If the current operation state of the satellite energy system does not belong to the safe state set, the current operation state is not an abnormal mode.
7. The intelligent control method of the satellite energy system according to claim 1, characterized in that The training process of the satellite energy system environment model includes: Construct a simulation environment of the satellite energy system and obtain training data of the satellite energy system; Input the training data into the initial environment model and run the initial environment model in the simulation environment; wherein, the initial environment model includes an initial large language sub-model and an optimization algorithm module; Utilize the initial large language sub-model in the satellite energy system environment model to perform semantic understanding and feature correlation analysis on the training data, generate training key features, construct a large language loss function, substitute the training key features into the large language loss function for calculation, and adjust the parameters of the initial large language sub-model with the goal of minimizing the calculation result of the large language loss function to obtain the trained large language sub-model; Input the training key features into the optimization algorithm module, construct an optimization loss function, and adjust the parameters of the optimization algorithm module with the goal of minimizing the calculation result of the optimization loss function to obtain the trained intelligent decision-making framework; Based on the large language sub-model and the intelligent decision-making framework, obtain the trained satellite energy system environment model.
8. The intelligent control method of the satellite energy system according to claim 7, characterized in that The obtaining of the training data of the satellite energy system includes: Utilize the state transition function in the simulation environment to generate empirical data; the state transition function is used to describe the state change of the satellite energy system under a given current state, executed action, and environmental random factors; Based on the empirical data, obtain the training data.
9. The intelligent control method of the satellite energy system according to claim 7, characterized in that, The inputting of the training key features into the optimization algorithm module, constructing an optimization loss function, and adjusting the parameters of the optimization algorithm module with the goal of minimizing the calculation result of the optimization loss function to obtain the trained intelligent decision-making framework includes: The optimization algorithm module includes a policy network and a value network; Based on the training key features, generate training state data corresponding to the state space; via the policy network, based on the training key features and the training state data, generate an action probability distribution, and select training action data corresponding to the action space based on the action probability distribution; Via the value network, based on the training state data, the training action data, and the reward function, calculate the immediate reward value of the current training action data; According to the training state data, the training action data, the immediate reward value, and the next training state data generated by the state transition function, calculate the result of the optimization loss function, and adjust the parameters of the optimization algorithm module with the goal of minimizing the result of the optimization loss function to obtain the trained intelligent decision-making framework.
10. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is set to execute the satellite energy system intelligent control method according to any one of claims 1 to 9 when running.
Citation Information
Patent Citations
Satellite exploration control system and method based on deep reinforcement learning
CN116692027A
Large-model-driven multi-satellite collaborative perception decision-making method
CN118631314A
Deep reinforcement learning scheduling method and device for satellite multi-point target imaging
CN118709748A
Industrial textile control system and method based on large language model
CN118798486A
Logic modeling-based high-orbit satellite network dynamic optimization method and SDN (Software Defined Network) controller
CN119472447A
Cited By
Satellite communication network construction method based on big language model reasoning
CN121036829A