Bandwidth and power optimization method, optimization algorithm construction method, device and equipment

By building a comprehensive optimization model of bandwidth and power in satellite communications and using reinforcement learning algorithms for solving, the problem of low user satisfaction in the existing technology is solved, and a more efficient communication supply and user experience is achieved.

CN120128237APending Publication Date: 2025-06-10BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142491.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art only considers bandwidth or power when optimizing the satellite communication process, and cannot effectively increase the amount of communication provided by users, resulting in low user satisfaction.

Method used

By obtaining user location, target wave position, related power and bandwidth information, combining channel gain calculation, an optimization model is built, and using the optimization algorithm built by reinforcement learning to solve it, comprehensive optimization of bandwidth and power is achieved.

Benefits of technology

In a complex and changeable communication environment, improve communication supply, reduce network latency, improve signal stability and data transmission rate, thereby improving user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128237A_ABST
    Figure CN120128237A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a bandwidth and power optimization method, an optimization algorithm construction method, a device and equipment, and relates to the technical field of wireless communication. The bandwidth and power optimization method comprises the following steps: acquiring a user position, a target wave position, a first variable, a second variable, a first power and a first bandwidth; calculating a channel gain according to the user position, the target wave position and a plurality of preset channel parameters; according to the first variable, the second variable, the first power, the first bandwidth and the channel gain, calculating to obtain a communication provision amount; performing model construction according to the communication supply quantity and a preset communication demand quantity to obtain an optimization model; and solving the optimization model according to a preset optimization algorithm to obtain an optimization scheme. In the optimization process, two dimensions of bandwidth and power are considered at the same time, so that the method can more comprehensively adapt to a complex and changeable communication environment, the communication providing amount is increased, and the user satisfaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technologies, and in particular, to a method for optimizing bandwidth and power, a method for constructing an optimization algorithm, a device, and equipment. Background Art

[0002] Satellite communication is a communication method that uses artificial satellites as relay stations for signal transmission. Its main advantage lies in its wide coverage area. However, satellite communication also faces challenges such as signal delay, limited spectrum resources, and unbalanced data transmission rates. Therefore, in order to improve the efficiency and reliability of satellite communication, it is necessary to optimize its communication process.

[0003] In the prior art, satellite hopping beam technology is usually used to optimize the satellite communication process. The satellite hopping beam technology can dynamically adjust the beam direction and coverage area emitted by the satellite to adapt to the changing user needs and geographical locations. By flexibly managing the coverage range of the beam, the satellite hopping beam technology can effectively reduce signal interference, enhance the reliability of communication, and meet the needs of more users under limited spectrum resources.

[0004] However, the prior art has the problem of low user satisfaction. When optimizing the satellite communication process in the prior art, only bandwidth or power is often optimized. This single-dimensional optimization strategy cannot effectively improve the communication supply obtained by users in the face of a complex and changing communication environment. The low communication supply obtained by users often causes problems such as network delay, unstable signal, and low data transmission rate for users, resulting in low user satisfaction. Summary of the Invention

[0005] Embodiments of this application provide a method for optimizing bandwidth and power, a method for constructing an optimization algorithm, a device, and equipment to solve the problem of low user satisfaction in the prior art.

[0006] In a first aspect, embodiments of this application provide a method for optimizing bandwidth and power. The method is applied to a bandwidth and power optimization device, and the method includes:

[0007] Obtain the user location, the target wave position, a first variable, a second variable, a first power, and a first bandwidth; wherein, the target wave position refers to the position of the wave where the user is located among a preset plurality of waves, the first variable is used to represent the activation situation of the plurality of waves, the second variable is used to represent the activation situation of a preset plurality of subcarriers in the beam corresponding to the target wave, the first power is used to represent the real-time power of the subcarrier where the user is located, and the first bandwidth is used to represent the real-time bandwidth of the subcarrier where the user is located;

[0008] Calculate the channel gain according to the user location, the target wave position, and a preset plurality of channel parameters;

[0009] Calculate the communication provision based on the first variable, the second variable, the first power, the first bandwidth, and the channel gain; wherein, the communication provision refers to the communication volume obtained by the user within a preset time period.

[0010] Construct an optimization model based on the communication provision and a preset communication demand.

[0011] Solve the optimization model according to a preset optimization algorithm to obtain an optimization solution; wherein, the optimization solution is used to improve the communication provision, and the optimization algorithm is constructed through reinforcement learning.

[0012] In a possible design, the calculating the communication provision based on the first variable, the second variable, the first power, the first bandwidth, and the channel gain includes:

[0013] Calculate the signal-to-interference-plus-noise ratio based on the first power and the channel gain.

[0014] Calculate the communication rate based on the first variable, the second variable, the first bandwidth, and the signal-to-interference-plus-noise ratio.

[0015] Calculate the communication provision based on the communication rate.

[0016] In a possible design, before obtaining the user location, the target wave position, the first variable, the second variable, the first power, and the first bandwidth, further include:

[0017] Establish a hopping beam resource optimization scenario model; wherein, the hopping beam resource optimization scenario model is used to represent the multiple wave positions of the communication satellite, the beam corresponding to each wave position, and the multiple subcarriers included in each beam.

[0018] In a second aspect, an embodiment of the present application provides a method for constructing an optimization algorithm, and the method is applied to a device for constructing an optimization algorithm, then the method includes:

[0019] Obtain the communication provision, the communication demand, the first variable, the second variable, the second power, and the channel gain; wherein, the second power refers to the power of each of the preset multiple subcarriers in the beam corresponding to the target wave position, and the channel gain is used to represent the attenuation degree of the signal during transmission.

[0020] Construct an optimization framework and a neural network architecture based on the communication provision, the communication demand, the first variable, the second variable, the second power, and the channel gain.

[0021] Construct an optimization algorithm according to the optimization framework and the neural network architecture, where the optimization algorithm is used for the bandwidth and power optimization method described in the first aspect.

[0022] In a possible design, the construction of the optimization framework and the neural network architecture according to the communication supply, the communication demand, the first variable, the second variable, the second power, and the channel gain includes:

[0023] Construct a first state, a first action, and a first reward function according to the communication supply, the communication demand, the first variable, the second variable, and the second power, where the first state is used to represent the difference between the user's communication demand and the communication supply of the communication satellite, the first action is used to represent the activation status of a preset number of wave positions, the activation status of a preset number of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the first reward function is used to represent the communication volume obtained by the user after performing the first action, and the target wave position refers to the wave position where the user is located among the plurality of wave positions;

[0024] Construct the optimization framework according to the first state, the first action, and the first reward function;

[0025] Construct a second state, a second action, and a second reward function according to the communication supply, the communication demand, the first variable, the second variable, the second power, and the channel gain, where the second state is used to represent the communication demand, the communication supply, and the channel gain, the second action is used to represent the activation status of the plurality of wave positions, the activation status of a preset number of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the second reward function is used to represent the communication volume obtained by the user after performing the second action;

[0026] Construct the neural network architecture according to the second state, the second action, and the second reward function.

[0027] In a possible design, the construction of the first state, the first action, and the first reward function according to the communication supply, the communication demand, the first variable, the second variable, and the second power includes:

[0028] Construct the first state according to the communication supply and the communication demand;

[0029] Construct the first action according to the first variable, the second variable, and the second power;

[0030] Execute the first action, and take the communication volume obtained after executing the first action as the first reward function.

[0031] In a possible design, the neural network architecture includes a first neural network, a second neural network, and a third neural network. Then, constructing a second state, a second action, and a second reward function according to the communication supply volume, the communication demand volume, the first variable, the second variable, the second power, and the channel gain includes:

[0032] Input the third state into the first neural network to obtain a third action; wherein, the third state includes the communication supply volume, the communication demand volume, and the channel gain, and the first neural network is used to determine the wave position of the user.

[0033] Input the fourth state into the second neural network to obtain a fourth action; wherein, the fourth state includes the third state and the third action, and the second neural network is used to determine the subcarriers of the user.

[0034] Input the fifth state into the third neural network to obtain a fifth action; wherein, the fifth state includes the fourth state and the fourth action, and the third neural network is used to determine the power of the subcarrier where the user is located.

[0035] Generate the second state according to the third state, the fourth state, and the fifth state.

[0036] Generate the second action according to the third action, the fourth action, and the fifth action.

[0037] Execute the second action, and take the communication volume obtained after executing the second action as the second reward function.

[0038] In a third aspect, the present application provides a bandwidth and power optimization device, and the device includes:

[0039] A first acquisition module, configured to acquire the user location, the target wave position, the first variable, the second variable, the first power, and the first bandwidth; wherein, the target wave position refers to the position of the wave where the user is located among a preset plurality of waves, the first variable is used to represent the activation situation of the plurality of waves, the second variable is used to represent the activation situation of a preset plurality of subcarriers in the beam corresponding to the target wave, the first power is used to represent the real-time power of the subcarrier where the user is located, and the first bandwidth is used to represent the real-time bandwidth of the subcarrier where the user is located.

[0040] A first calculation module, configured to calculate the channel gain according to the user location, the target wave position, and a preset plurality of channel parameters.

[0041] A second calculation module, configured to calculate a communication supply amount according to the first variable, the second variable, the first power, the first bandwidth, and the channel gain; wherein, the communication supply amount refers to the communication amount obtained by a user within a preset time period;

[0042] A first construction module, configured to construct an optimization model according to the communication supply amount and a preset communication demand amount;

[0043] A model solution module, configured to solve the optimization model according to a preset optimization algorithm to obtain an optimization solution; wherein, the optimization solution is used to improve the communication supply amount, and the optimization algorithm is constructed by reinforcement learning.

[0044] In a possible design, the second calculation module includes:

[0045] A first calculation unit, configured to calculate a signal-to-interference-plus-noise ratio according to the first power and the channel gain;

[0046] A second calculation unit, configured to calculate a communication rate according to the first variable, the second variable, the first bandwidth, and the signal-to-interference-plus-noise ratio;

[0047] A third calculation unit, configured to calculate the communication supply amount according to the communication rate.

[0048] In a possible design, the bandwidth and power optimization device further includes:

[0049] A model establishment module, configured to establish a hopping beam resource optimization scenario model; wherein, the hopping beam resource optimization scenario model is used to represent the multiple wave positions of the communication satellite, the beam corresponding to each wave position, and the multiple subcarriers included in each beam.

[0050] In a fourth aspect, the present application provides a device for constructing an optimization algorithm, and the device includes:

[0051] A second acquisition module, configured to acquire a communication supply amount, a communication demand amount, a first variable, a second variable, a second power, and a channel gain; wherein, the second power refers to the power of each of the preset multiple subcarriers in the beam corresponding to the target wave position, and the channel gain is used to represent the attenuation degree of the signal during transmission;

[0052] A second construction module, configured to construct an optimization framework and a neural network architecture according to the communication supply amount, the communication demand amount, the first variable, the second variable, the second power, and the channel gain;

[0053] A third construction module, configured to construct an optimization algorithm according to the optimization framework and the neural network architecture; wherein, the optimization algorithm is used for the bandwidth and power optimization device described in the third aspect.

[0054] In a possible design, the second construction module includes:

[0055] A first construction unit, configured to construct a first state, a first action, and a first reward function according to the communication supply, the communication demand, the first variable, the second variable, and the second power; wherein, the first state is used to represent the difference between the user's communication demand and the communication supply of the communication satellite, the first action is used to represent the activation conditions of a preset plurality of wave positions, the activation conditions of a preset plurality of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the first reward function is used to represent the communication volume obtained by the user after executing the first action, and the target wave position refers to the wave position where the user is located among the plurality of wave positions;

[0056] A second construction unit, configured to construct the optimization framework according to the first state, the first action, and the first reward function;

[0057] A third construction unit, configured to construct a second state, a second action, and a second reward function according to the communication supply, the communication demand, the first variable, the second variable, the second power, and the channel gain; wherein, the second state is used to represent the communication demand, the communication supply, and the channel gain, the second action is used to represent the activation conditions of the plurality of wave positions, the activation conditions of a preset plurality of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the second reward function is used to represent the communication volume obtained by the user after executing the second action;

[0058] A fourth construction unit, configured to construct the neural network architecture according to the second state, the second action, and the second reward function.

[0059] In a possible design, the first construction unit includes:

[0060] A first construction component, configured to construct the first state according to the communication supply and the communication demand;

[0061] A second construction component, configured to construct the first action according to the first variable, the second variable, and the second power;

[0062] A first execution component, configured to execute the first action and take the communication volume obtained after executing the first action as the first reward function.

[0063] In a possible design, the neural network architecture includes a first neural network, a second neural network, and a third neural network. Then, the third construction unit includes:

[0064] A first input component, configured to input a third state into the first neural network to obtain a third action; wherein, the third state includes the communication supply amount, the communication demand amount, and the channel gain, and the first neural network is used to determine the wave position of the user;

[0065] A second input component, configured to input a fourth state into the second neural network to obtain a fourth action; wherein, the fourth state includes the third state and the third action, and the second neural network is used to determine the subcarrier of the user;

[0066] A third input component, configured to input a fifth state into the third neural network to obtain a fifth action; wherein, the fifth state includes the fourth state and the fourth action, and the third neural network is used to determine the power of the subcarrier where the user is located;

[0067] A first generation component, configured to generate the second state according to the third state, the fourth state, and the fifth state;

[0068] A second generation component, configured to generate the second action according to the third action, the fourth action, and the fifth action;

[0069] A second execution component, configured to execute the second action and take the communication amount obtained after executing the second action as the second reward function.

[0070] A bandwidth and power optimization method, a construction method of an optimization algorithm, a device and equipment provided by the present application. The bandwidth and power optimization method includes: obtaining a user location, a target wave position, a first variable, a second variable, a first power and a first bandwidth; calculating a channel gain according to the user location, the target wave position and a plurality of preset channel parameters; calculating a communication supply according to the first variable, the second variable, the first power, the first bandwidth and the channel gain; constructing an optimization model according to the communication supply and a preset communication demand; and solving the optimization model according to a preset optimization algorithm to obtain an optimization solution. The bandwidth and power optimization method of the present application calculates a channel gain by obtaining a user location, a target wave position, and related power and bandwidth information, constructs an optimization model, and solves it using a preset optimization algorithm. This method considers both the bandwidth and power dimensions during the optimization process, enabling it to more comprehensively adapt to complex and changing communication environments and improve the communication supply. Through this optimization strategy, network latency can be effectively reduced, signal stability and data transmission rate can be improved, thereby enhancing user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 It is a schematic diagram of the application scenario of the bandwidth and power optimization method provided by the embodiment of the present application;

[0073] Figure 2 It is a flowchart of the bandwidth and power optimization method provided by the embodiment of the present application Figure 1 ;

[0074] Figure 3 It is a flowchart of the bandwidth and power optimization method provided by the embodiment of the present application Figure 2 ;

[0075] Figure 4 It is a flowchart of the construction method of the optimization algorithm provided by the embodiment of the present application;

[0076] Figure 5 It is a schematic diagram of the change of the objective function of the bandwidth and power optimization method provided by the embodiment of the present application on the validation set;

[0077] Figure 6Schematic diagram of the change of the loss function of the first neural network provided by the embodiment of the present application with the number of training rounds;

[0078] Figure 7 Schematic diagram of the change of the loss function of the second neural network provided by the embodiment of the present application with the number of training rounds;

[0079] Figure 8 Schematic diagram of the change of the loss function of the third neural network provided by the embodiment of the present application with the number of training rounds;

[0080] Figure 9 Schematic diagram of the structure of the bandwidth and power optimization device provided by the embodiment of the present application;

[0081] Figure 10 Schematic diagram of the structure of the construction device of the optimization algorithm provided by the embodiment of the present application;

[0082] Figure 11 Electronic device provided by the embodiment of the present application. Detailed implementation mode

[0083] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0084] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different. It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more.

[0085] It should be noted that "when..." in the embodiments of the present application can be at the instant when a certain situation occurs, or within a period of time after a certain situation occurs. The embodiments of the present application do not make specific limitations in this regard. In addition, the bandwidth and power optimization methods, the construction methods of optimization algorithms, devices, and equipment provided in the embodiments of the present application are only examples, and the bandwidth and power optimization methods, the construction methods of optimization algorithms, devices, and equipment may also include more or less content.

[0086] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:

[0087] Beam hopping: It is a wireless communication technology that improves the transmission efficiency and anti-interference ability of signals by quickly changing the beam direction of the antenna. By switching the beam direction at different time intervals, the beam hopping technology can dynamically adapt to environmental changes, optimize signal coverage and quality, thereby enhancing the performance and reliability of the communication system.

[0088] Beam: It refers to the directional distribution of electromagnetic waves emitted or received by the antenna in space. It usually presents as a three-dimensional area with a specific shape and width, concentrating energy in a specific direction to enhance signal strength and coverage. The shape and direction of the beam can be controlled through antenna design and configuration, and are widely used in wireless communication, radar, and satellite systems to improve the efficiency and accuracy of signal transmission.

[0089] Subcarrier: In a multi-carrier communication system, the available spectrum is divided into multiple narrower frequency bands, and each frequency band is called a subcarrier. These subcarriers are orthogonal to each other and can transmit data simultaneously without interference, thereby improving spectrum utilization and data transmission efficiency.

[0090] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0091] The following uses specific embodiments to detail the technical solutions of the present invention. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present invention in conjunction with the drawings.

[0092] To clearly understand the technical solution of this application, the solutions of the prior art will be introduced in detail first. Satellite communication is a communication method that uses artificial satellites as relay stations to achieve radio signal transmission between different locations. It is widely used in fields such as broadcasting, television, and the Internet, and has the advantages of wide coverage and being unrestricted by geographical barriers. However, the long-distance signal transmission in satellite communication may cause problems such as signal delay, path loss, and interference. Therefore, in order to improve the efficiency and reliability of satellite communication, it is necessary to optimize the satellite communication process.

[0093] In the prior art, satellite hopping beam technology is usually used to optimize the satellite communication process. The satellite hopping beam technology can dynamically adjust the beam direction and coverage area emitted by the satellite to adapt to the changing user needs and geographical locations. By flexibly managing the coverage range of the beam, the satellite hopping beam technology can effectively reduce signal interference, enhance the reliability of communication, and meet the needs of more users under limited spectrum resources. However, the prior art often only optimizes the bandwidth or power when optimizing the satellite communication process. This single-dimensional optimization strategy cannot effectively improve the communication supply obtained by users in the face of a complex and changing communication environment. The low communication supply obtained by users often causes problems such as network delay, unstable signal, and low data transmission rate, resulting in low user satisfaction.

[0094] Therefore, aiming at the problem of low user satisfaction in the prior art, it is found in the research that to solve this problem, a multi-dimensional optimization strategy can be constructed, comprehensively considering factors such as bandwidth, power, beam direction, and coverage area, so as to improve the user's communication experience and satisfaction: ① A multi-objective optimization algorithm can be introduced to incorporate multiple factors such as bandwidth, power, beam direction, and coverage area into the optimization model. By weighing the relationships between different objectives, the resource allocation strategy can be dynamically adjusted. ② Adaptive beamforming technology can be adopted to dynamically adjust the shape and direction of the beam according to the real-time user distribution and demand changes, maximize the signal coverage and intensity, reduce interference, improve the signal stability and data transmission rate, and thus enhance the user's communication experience. ③ Artificial intelligence technology can be used for intelligent spectrum management to predict user needs and environmental changes, and optimize the allocation and use of spectrum resources.

[0095] Specifically:

[0096] A multi-dimensional optimization strategy can be adopted. This strategy takes into account various factors such as bandwidth, power, and network load. By using machine learning technology, the system can analyze and predict the changes in user needs in real time, and dynamically adjust the direction, coverage area, and resource allocation of the satellite beam. Through this comprehensive optimization method, the signal stability and data transmission rate can be effectively improved, and the network delay can be reduced, thereby enhancing user satisfaction.

[0097] The bandwidth and power optimization method of the embodiments of the present application obtains the user location, target wave position, and related power and bandwidth information, constructs an optimization model by combining channel gain calculations, and uses a preset optimization algorithm for solution. This method considers both the bandwidth and power dimensions during the optimization process, enabling it to more comprehensively adapt to complex and changing communication environments and improve communication throughput. Through this optimization strategy, network latency can be effectively reduced, signal stability and data transmission rate can be improved, thereby enhancing user satisfaction.

[0098] Based on the above creative findings, the technical solution of the present application is proposed.

[0099] The application scenarios of the bandwidth and power optimization method provided by the embodiments of the present invention are introduced below. Figure 1 It is a schematic diagram of the application scenario of the bandwidth and power optimization method provided by the embodiments of the present application. As Figure 1 shown, this application scenario includes a signal sending end 101, a communication area 102, multiple beams 103, and multiple signal receiving ends 104. The communication area includes multiple activated wave positions 1021 and multiple unactivated wave positions 1022. The signals sent by the signal sending end 101 are sent to multiple signal receiving ends 104 in multiple activated wave positions 1021 through multiple beams 103. For example, the signal sending end 101 can be a multi-beam satellite. Within the service range of the multi-beam satellite, there are a total of N users waiting to be served. This multi-beam satellite can generate at most K beams to meet the service needs of these users, and each beam can generate M subcarriers. In order to meet the needs of users as much as possible, the multi-beam satellite needs to determine which wave positions to light up and which subcarrier to allocate to the users under the lit wave positions.

[0100] The embodiments of the present invention are introduced below in conjunction with the accompanying drawings of the specification.

[0101] Figure 2 It is a flowchart of the bandwidth and power optimization method provided by the embodiments of the present application Figure 1 As Figure 2 shown, in this embodiment, the execution subject of the embodiments of the present invention is the signal sending end. Then the bandwidth and power optimization method provided by this embodiment includes the following steps:

[0102] S201. Obtain the user location, target wave position, first variable, second variable, first power, and first bandwidth.

[0103] Specifically, the user location, target wave position, first variable, second variable, first power, and first bandwidth can be obtained through multiple sensors and data collection modules, which are capable of monitoring and recording the geographical location of the user device, the currently connected network wave position, the network configuration status, and the power and bandwidth usage of the device in real time. The acquisition of this information is the basis of the optimization process, aiming to provide data support for subsequent channel gain calculation and communication volume calculation, so as to ensure that the optimization model can accurately reflect the actual communication environment and requirements of the user and provide reliable input for the final optimization solution. Among them, the target wave position refers to the position of the wave where the user is located among the preset multiple waves, the first variable is used to represent the activation situation of the multiple waves, the second variable is used to represent the activation situation of the preset multiple subcarriers in the beam corresponding to the target wave, the first power is used to represent the real-time power of the subcarrier where the user is located, and the first bandwidth is used to represent the real-time bandwidth of the subcarrier where the user is located.

[0104] Assume that the time for the satellite to serve this area is T time slots, and the satellite antenna has already allocated L wave positions. Assuming that the user's state and channel do not change within each time slot, the first variable at time slot t can be expressed as:

[0105]

[0106] where t is a variable used to describe which time slot. If wave position l is lit at time slot t, then Otherwise . The second variable at time slot t can be expressed as:

[0107]

[0108] where n is a variable used to represent which user, m is a variable used to represent which subcarrier, and l is a variable used to represent which wave position. When user n uses subcarrier m at time slot t, then , Otherwise .

[0109] S202. Calculate the channel gain according to the user location, target wave position, and a preset plurality of channel parameters.

[0110] Specifically, the channel gain can be calculated through a path loss model and channel modeling technology. This process involves using the geographical location of the user and the target wave position, combined with preset channel parameters such as frequency, antenna gain, and propagation environment, to calculate the attenuation and enhancement of the signal in the transmission path, so as to obtain the channel gain. The channel gain is used to evaluate the transmission quality and efficiency of the signal in a specific environment and is one of the key factors for optimizing communication volume, helping to accurately reflect the impact of channel conditions on communication performance in the optimization model.

[0111] For example, in a multi-beam low-Earth orbit satellite system, without considering Doppler frequency shift and rain attenuation, the channel model for beam k irradiating user n can be established as follows:

[0112]

[0113] where k is a variable used to represent which beam, h is the channel gain, represents the receiving antenna gain of user n, represents the transmitting antenna gain of beam k for user n, represents the free space attenuation, represents the distance from the satellite to user n, represents the wavelength.

[0114] S203. Calculate the communication supply based on the first variable, second variable, first power, first bandwidth, and channel gain.

[0115] Specifically, the performance calculation formula of the communication system can be applied to calculate the communication supply. The first variable and the second variable are respectively used to determine which wave positions and subcarriers are in the active state. The first power and the channel gain jointly determine the effective transmission power of the signal, while the first bandwidth affects the data transmission rate. By substituting these factors into Shannon's theorem or other communication capacity formulas, the maximum data transmission amount that the user can obtain within a specific time period, that is, the communication supply, can be calculated. The communication supply is used to evaluate whether the current network configuration can meet the communication needs of the user and is one of the core indicators for optimizing model construction and solution. Among them, the communication supply refers to the communication volume obtained by the user within a preset time period.

[0116] S204. Construct a model based on the communication supply and the preset communication demand to obtain an optimization model.

[0117] Specifically, a mathematical optimization problem can be established based on the communication supply and the preset communication demand. The communication supply represents the performance output of the system under the current configuration, while the preset communication demand is the minimum requirement of the user or application for communication services. By using these two as the input of the constraint conditions and the objective function, an optimization model is constructed. This optimization model aims to maximize the matching degree between the communication supply and the demand. This optimization model is used to identify and adjust system parameters to ensure that while meeting or exceeding user needs, the utilization efficiency of network resources is improved, and it is the basis for generating optimization solutions.

[0118] For example, assume that the total communication demand of user n is . In order to satisfy the user's communication needs as much as possible, the following optimization model can be constructed:

[0119]

[0120] The goal of this optimization model is to minimize the unmet communication requirements of users. Among them, constraint C1 is used to ensure that the sum of the transmission powers of all beams does not exceed the total power; constraint C2 is used to ensure that the power of subcarriers under each beam cannot exceed the maximum power; constraint C3 is the first variable, when the beam is lit at time slot t, otherwise it is 0, constraint C4 is the second variable, when user n is assigned subcarrier m at time slot t, otherwise it is 0. Constraint C5 states that each carrier of each beam will be selected by at most one user at the current wave position; constraint C6 indicates that each user selects at most one subcarrier at the same time, and constraint C7 is used to ensure that at most K beams are lit simultaneously at the same moment.

[0121] S205. Solve the optimization model according to the preset optimization algorithm to obtain an optimization solution.

[0122] Specifically, the optimization model can be solved by applying a reinforcement learning algorithm to obtain an optimization solution. The optimization algorithm will iterate and evaluate different system configurations to find a parameter combination that can increase the communication supply. By solving the optimization model, the algorithm can identify the optimization solution under the given constraints. This optimization solution is used to guide the system to adjust the bandwidth allocation and power settings to improve the overall communication performance and resource utilization efficiency, and ensure that the user requirements are effectively met. Among them, the optimization solution is used to increase the communication supply, and the optimization algorithm is constructed through reinforcement learning.

[0123] A bandwidth and power optimization method provided in this embodiment includes: obtaining the user location, the target wave position, the first variable, the second variable, the first power, and the first bandwidth; calculating the channel gain according to the user location, the target wave position, and a plurality of preset channel parameters; calculating the communication supply according to the first variable, the second variable, the first power, the first bandwidth, and the channel gain; constructing an optimization model according to the communication supply and the preset communication demand; and solving the optimization model according to the preset optimization algorithm to obtain an optimization solution. A bandwidth and power optimization method achieves the following technical effects: by obtaining the user location, the target wave position, and the relevant power and bandwidth information, combining the channel gain calculation, constructing an optimization model, and using the preset optimization algorithm for solution. This method considers both the bandwidth and power dimensions during the optimization process, so as to more comprehensively adapt to the complex and changeable communication environment and increase the communication supply. Through this optimization strategy, network latency can be effectively reduced, signal stability and data transmission rate can be improved, thereby enhancing user satisfaction.

[0124] In a possible design, S203 calculates the communication offering based on the first variable, the second variable, the first power, the first bandwidth, and the channel gain, including:

[0125] S2031 calculates the signal-to-interference-plus-noise ratio (SINR) based on the first power and the channel gain.

[0126] Specifically, the SINR is calculated by taking the ratio of the signal power to the sum of the interference and noise power. The interference and noise power can be obtained through measurement. The calculated SINR is used to evaluate the quality and reliability of the signal and is an important parameter for further calculating the communication rate and the communication offering, as it directly affects the effectiveness and efficiency of data transmission.

[0127] For example, under the same beam, orthogonal frequency division multiple access (OFDMA) technology is adopted between different subcarriers, so there is no interference between them. However, if subcarriers in the same frequency band are allocated between different beams, interference will occur. Therefore, the SINR of user n can be expressed as:

[0128]

[0129] Where, is the power spectral density of Gaussian white noise, and the term describes the interference between different beams.

[0130] S2032 calculates the communication rate based on the first variable, the second variable, the first bandwidth, and the SINR.

[0131] Specifically, the communication rate can be calculated by applying Shannon's theorem or other relevant communication rate formulas. The first variable and the second variable are used to determine which wave positions and subcarriers are in the active state, while the first bandwidth and the SINR are used to calculate the maximum data transmission rate that can be achieved on these active subcarriers. The calculated communication rate is used to evaluate the potential data transmission capacity under the current network configuration and is the basis for further calculating the communication offering.

[0132] For example, the communication rate provided to the user can be:

[0133]

[0134] Where, is the first variable, indicating whether wave position l is active, is the second variable, indicating whether user n under wave position l selects subcarrier m in time slot t, and B is the bandwidth of each subcarrier.

[0135] S2033 calculates the communication offering based on the communication rate.

[0136] Specifically, the communication rate represents the amount of data that can be transmitted per unit time. Therefore, by multiplying the communication rate by the length of a specific time period, the total data transmission amount during that time period, i.e., the communication provision, can be calculated. The calculated communication provision is used to evaluate the actual communication service level that users can obtain within a given time and is one of the important input parameters for optimizing model construction, helping to determine whether the current configuration meets the communication needs of users.

[0137] For example, the communication provision of user n during a communication process can be expressed as:

[0138]

[0139] The technical effect of this solution in this embodiment is that by calculating the signal-to-interference-plus-noise ratio and the communication rate, the performance of the communication system can be evaluated and optimized more accurately. By comprehensively considering factors such as the first power, channel gain, bandwidth, and activation status, this method can effectively calculate the communication provision of users under specific conditions. This precise calculation method improves the utilization efficiency of communication resources, ensures that users can obtain stable and efficient communication services under different network conditions, thereby enhancing the overall user experience and network performance.

[0140] Figure 3 Schematic flow of the bandwidth and power optimization method provided by the embodiments of the present application Figure 2 . In this embodiment, based on the Figure 2 embodiments provided, the bandwidth and power optimization method is further explained. The bandwidth and power optimization method includes:

[0141] S301. Establish a hopping beam resource optimization scenario model.

[0142] Specifically, a mathematical model can be constructed, which describes multiple wave positions of a communication satellite, the corresponding beam for each wave position, and multiple subcarriers included in each beam. This process involves defining and configuring the structure of the satellite communication system, including the geographical distribution of wave positions, the coverage range of beams, and the spectrum allocation of subcarriers. This model is used to simulate and analyze the impact of different beam and subcarrier configurations on communication performance, providing basic data and a scenario environment for subsequent bandwidth and power optimization, thereby supporting more accurate resource allocation and optimization decisions. Among them, the hopping beam resource optimization scenario model is used to represent multiple wave positions of a communication satellite, the corresponding beam for each wave position, and multiple subcarriers included in each beam.

[0143] S302. Obtain the user location, the target wave position, the first variable, the second variable, the first power, and the first bandwidth. Among them, the target wave position refers to the position of the wave where the user is located among the preset multiple waves. The first variable is used to represent the activation situation of the multiple waves, the second variable is used to represent the activation situation of the preset multiple subcarriers in the beam corresponding to the target wave, the first power is used to represent the real-time power of the subcarrier where the user is located, and the first bandwidth is used to represent the real-time bandwidth of the subcarrier where the user is located.

[0144] S303. Calculate the channel gain according to the user location, the target wave position, and the preset multiple channel parameters.

[0145] S304. Calculate the communication supply according to the first variable, the second variable, the first power, the first bandwidth, and the channel gain. Among them, the communication supply refers to the communication volume obtained by the user within a preset time period.

[0146] S305. Construct an optimization model according to the communication supply and the preset communication demand to obtain an optimization model.

[0147] S306. Solve the optimization model according to the preset optimization algorithm to obtain an optimization solution. Among them, the optimization solution is used to improve the communication supply, and the optimization algorithm is constructed through reinforcement learning.

[0148] S302 - S306 is similar to S201 - S205, and will not be elaborated in this embodiment.

[0149] The technical effect of this solution in this embodiment is that by establishing a hopping beam resource optimization scenario model, it can systematically describe and simulate the configuration of the waves, beams, and subcarriers of communication satellites. This model provides a detailed framework, enabling more accurate consideration of the resource allocation of different beams and subcarriers and their impact on communication performance when performing bandwidth and power optimization. In this way, the optimization process can more effectively adjust the resource allocation strategy, improve the overall efficiency and service quality of the communication system, and enhance the communication experience of users.

[0150] Figure 4 It is a flow chart of the construction method of the optimization algorithm provided by the embodiment of the present application. As Figure 4 shown, in this embodiment, the execution subject of the embodiment of the present invention is the signal sending end. Then the construction method of the optimization algorithm provided by this embodiment includes the following steps:

[0151] S401. Obtain the communication supply, the communication demand, the first variable, the second variable, the second power, and the channel gain.

[0152] Specifically, communication supply volume, communication demand volume, a first variable, a second variable, a second power, and channel gain can be obtained by collecting and analyzing data such as the network environment, user requirements, and system status. These data are used to construct an optimization framework and a neural network architecture in order to develop an optimization algorithm that can effectively allocate bandwidth and power resources in a given communication environment, thereby improving the communication supply volume to meet the communication needs of users. This process involves the evaluation of channel conditions, the prediction of user requirements, and the dynamic adjustment of system resources to achieve the optimization of overall performance. Among them, the second power refers to the power of each of the preset multiple subcarriers in the beam corresponding to the target wave position, and the channel gain is used to express the attenuation degree of the signal during transmission.

[0153] S402. Construct an optimization framework and a neural network architecture according to the communication supply volume, communication demand volume, first variable, second variable, second power, and channel gain.

[0154] Specifically, these parameters can be used as input features to design an optimization framework that can capture system dynamics and user requirements. These parameters are also used to train the neural network so that it can learn and predict resource allocation strategies under different conditions. The combination of the optimization framework and the neural network architecture aims to improve the decision-making efficiency and accuracy of the system, thereby realizing the allocation of bandwidth and power in a complex communication environment and meeting the communication needs of users.

[0155] S403. Construct an optimization algorithm according to the optimization framework and the neural network architecture.

[0156] Specifically, the optimization framework provides the objective function and constraint conditions of the algorithm, and the neural network architecture can generate a model that can predict and optimize resource allocation by learning historical data and environmental characteristics. This optimization algorithm is used to dynamically adjust the bandwidth and power allocation in the communication system to improve the communication supply volume, ensure that the communication needs of users are met under various network conditions, and improve the overall system performance. Among them, the optimization algorithm is used for Figures 2 to 3 bandwidth and power optimization method.

[0157] The technical effect of this solution in this embodiment is that by constructing an optimization algorithm that combines an optimization framework and a neural network architecture, the intelligent and dynamic allocation of bandwidth and power resources in the communication system is realized. This method can effectively adapt to different channel conditions and user requirements, improve the communication supply volume and overall performance of the system. This optimization not only improves the resource utilization efficiency but also enhances the adaptability of the system in a complex environment, thereby meeting the real-time communication needs of users.

[0158] In a possible design, S402. Construct an optimization framework and a neural network architecture based on communication supply volume, communication demand volume, a first variable, a second variable, a second power, and channel gain, including:

[0159] S4021. Construct a first state, a first action, and a first reward function based on communication supply volume, communication demand volume, a first variable, a second variable, and a second power.

[0160] Specifically, these parameters can be transformed into elements in a reinforcement learning framework. The first state represents the difference between the user's communication demand and the communication volume provided by the satellite, reflecting the current state of the system. The first action involves how to activate different wave positions and subcarriers in the beam, and adjust the power of these subcarriers to optimize resource allocation. The first reward function is used to evaluate the communication volume obtained by the user after performing these actions, helping the reinforcement learning algorithm to optimize decisions through a reward mechanism. This design can use reinforcement learning to dynamically adjust system parameters to improve communication supply volume and resource utilization efficiency while meeting user needs. Among them, the first state is used to represent the difference between the user's communication demand volume and the communication supply volume of the communication satellite. The first action is used to represent the activation situation of a preset number of wave positions, the activation situation of a preset number of subcarriers in the beam corresponding to the target wave position, and the power of each of the multiple subcarriers. The first reward function is used to represent the communication volume obtained by the user after performing the first action. The target wave position refers to the wave position where the user is located among multiple wave positions.

[0161] S4022. Construct an optimization framework based on the first state, the first action, and the first reward function.

[0162] Specifically, the first state of the optimization framework is used to describe the conditions of the current system. The first action is used to define possible operation strategies, and the first reward function is used to evaluate the effects of these strategies. Through continuous iteration and learning, this framework can identify resource allocation strategies to increase the communication volume obtained by the user. This optimization framework can efficiently adjust system parameters in a dynamic environment, thereby improving the overall communication performance and resource utilization efficiency.

[0163] S4023. Construct a second state, a second action, and a second reward function based on communication supply volume, communication demand volume, a first variable, a second variable, a second power, and channel gain.

[0164] Specifically, the second state is used to describe the current conditions of the system, including communication demand, supply, and channel gain, which reflects the comprehensive situation of the network environment and user needs; the second action defines the strategies that can be taken under these conditions, including the activation status of wave positions and subcarriers and their power allocation; the second reward function is used to evaluate the effects of these strategies, that is, the communication volume obtained by the user after executing these actions. In this way, the reinforcement learning model can continuously adjust the strategy in a dynamic network environment to optimize resource allocation and meet user needs, ultimately improving the communication efficiency and performance of the system. Among them, the second state is used to represent the communication demand, communication supply, and channel gain, the second action is used to represent the activation status of multiple wave positions, the activation status of a preset number of subcarriers in the beam corresponding to the target wave position, and the power of each of the multiple subcarriers, and the second reward function is used to represent the communication volume obtained by the user after executing the second action.

[0165] S4024. Construct a neural network architecture according to the second state, the second action, and the second reward function.

[0166] Specifically, the neural network architecture can use the second state as the input layer to capture the conditions and environmental characteristics of the current system; the second action serves as the output layer, representing the optimization strategies that can be taken under these conditions; the second reward function is used to train the network, adjusting the network weights through a feedback mechanism to improve the effectiveness of the strategy. The purpose of this neural network architecture is to automatically optimize resource allocation and power management through deep learning technology, improving the system's decision-making ability and communication efficiency in a complex dynamic environment.

[0167] The technical effect of this solution in this embodiment is: By constructing a reinforcement learning framework and a neural network architecture, the intelligent optimization allocation of bandwidth and power resources in the communication system is realized. By defining the state, action, and reward function, this method can dynamically adapt to changes in user needs and channel conditions, automatically adjusting the activation and power allocation strategies of wave positions and subcarriers. This intelligent optimization not only improves the communication supply and resource utilization efficiency but also ensures that the communication needs of users are met in a complex and dynamic network environment, thus enhancing the overall performance and reliability of the system.

[0168] In a possible design, S4021. Construct the first state, the first action, and the first reward function according to the communication supply, communication demand, first variable, second variable, and second power, including:

[0169] S40211. Construct the first state according to the communication supply and communication demand.

[0170] Specifically, the first state can be constructed by calculating the difference between the actual communication volume obtained by the user and their demand. The first state reflects whether the current system meets the user's communication needs, as well as existing deficiencies or surpluses. This state is used as an input in the reinforcement learning framework to help the algorithm understand the effectiveness of the current resource allocation and guide subsequent policy adjustments. By accurately describing the system state, the algorithm can better optimize the resource allocation strategy to meet user needs and improve the overall communication efficiency of the system.

[0171] For example, the original optimization model is modeled as a Markov decision process:

[0172]

[0173] Taking the multi-beam satellite system as the environment, the first state is the unmet communication demand of the user, that is:

[0174]

[0175] where represents the total communication volume provided to user n at time t, and the total communication demand of user n is . That is, the unmet communication volume of the user at the current moment is equal to the unmet communication volume of the user at the previous moment plus the communication volume provided to the user at the current moment.

[0176] S40212. Construct the first action according to the first variable, the second variable, and the second power.

[0177] Specifically, the first action can be constructed by defining a set of policies to determine how to activate different wave positions and subcarriers and allocate appropriate power to these subcarriers. The first variable and the second variable are used to describe the activation states of the wave positions and subcarriers, while the second power is used to determine the power allocation for each subcarrier. These parameters together constitute the first action, representing the resource allocation strategy that can be taken in the current state. This action is used to explore and evaluate different resource configuration schemes in the reinforcement learning framework to optimize communication performance and meet user needs.

[0178] For example, the first action can be expressed as:

[0179]

[0180] where is used to represent the activation situation of multiple wave positions at time slot t, is used to represent the activation situation of multiple subcarriers at time slot t, is used to represent the power allocation situation of multiple subcarriers at time slot t.

[0181] S40213. Execute the first action and take the communication volume obtained after executing the first action as the first reward function.

[0182] Specifically, executing the first action involves activating specific wave positions and subcarriers and allocating corresponding powers. Subsequently, the system records the communication volume of the user under this configuration. This communication volume is used as the first reward function, reflecting the effectiveness of the current policy. This reward function is used to evaluate and optimize policy selection in the reinforcement learning framework, guiding the algorithm to learn and adjust in the direction of improving communication efficiency and meeting user needs through the reward mechanism.

[0183] For example, the reward function is used to evaluate the quality of executing the first action for the current environment. The better the first action executed, the larger the value of the reward function. Since the goal of the optimization model is to reduce the unmet communication needs of users, the communication needs provided to the user after each action execution can be used as the reward function:

[0184]

[0185] That is, for user n, in a time slot t, the communication volume provided after executing the selected first action is used as the reward function.

[0186] The technical effect of this solution in this embodiment is: By constructing the state, action, and reward function in reinforcement learning, a method for dynamically optimizing communication resource allocation is provided. By constructing the state based on the communication provided volume and demand volume, using the activation situation of wave positions and subcarriers and power allocation to construct the action, and using the actually obtained communication volume as the reward function, the system can adaptively adjust resource configuration in a complex network environment. This method effectively improves communication efficiency, ensures that user needs are met, and at the same time optimizes the overall performance and resource utilization rate of the system.

[0187] In a possible design, S4023. Construct a second state, a second action, and a second reward function according to the communication provided volume, communication demand volume, first variable, second variable, second power, and channel gain, including:

[0188] S40231. Input the third state into the first neural network to obtain the third action.

[0189] Specifically, the first neural network takes the third state as input and outputs an optimized decision, i.e., the third action. The third state includes communication supply volume, communication demand volume, and channel gain. This information is input into the first neural network, which is trained to identify the wave position selection strategy. The third action represents the wave position to which the user should be assigned under the current network conditions. This process aims to automatically optimize wave position selection through deep learning techniques to improve communication efficiency and meet user needs. Among them, the third state includes communication supply volume, communication demand volume, and channel gain, and the first neural network is used to determine the user's wave position.

[0190] For example, for the first neural network, its state can be composed of parts such as user communication demand volume, user communication supply volume, user channel gain, and the current time slot, expressed as . Where D represents the total communication demand volume of each user; represents the demand volume that has been provided to each user up to time t, and the dimension should be consistent with the number of users N; H is the user gain matrix, which represents the channel gain of each user to the other users. Here, it is assumed that the number of wave positions is L, and the dimension of the user gain matrix is set to , where each element stores the channel gain from the user in the current wave position to the user in the target wave position; t represents the current time slot.

[0191] S40232. Input the fourth state into the second neural network to obtain the fourth action.

[0192] Specifically, the second neural network takes the fourth state as input and outputs an optimized decision, i.e., the fourth action. The fourth state includes the third state and the third action. This information comprehensively describes the current communication environment and the selected wave position. By learning these input data, the second neural network can determine the subcarrier activation strategy. The fourth action represents which subcarriers should be activated under the current conditions to optimize communication performance. This process aims to optimize subcarrier selection through deep learning techniques to improve the spectral efficiency of the system and the communication experience of users. Among them, the fourth state includes the third state and the third action, and the second neural network is used to determine the user's subcarriers.

[0193] For example, for the second neural network, in addition to including the same parts as the first neural network, its state also needs to include the action information of the first neural network. Since the first neural network is for wave position selection, the action of the first neural network can be converted into state information through the following method:

[0194]

[0195] Among them, is the state input provided by the first neural network to the second neural network, is a 0-1 variable. When wave position l is selected, all users in this wave position are set to 1, otherwise set to 0. In this way, the input of the second neural network can be obtained according to the action selected by the first neural network.

[0196] S40233. Input the fifth state into the third neural network to obtain the fifth action.

[0197] Specifically, the third neural network takes the fifth state as input and outputs an optimized decision, that is, the fifth action. The fifth state includes the fourth state and the fourth action, comprehensively describing the current communication environment, the selected wave position and subcarriers. The third neural network uses this input information and can determine the power allocation strategy after training. The fifth action indicates how to allocate power to the subcarriers where users are located under the current conditions to optimize communication performance and resource utilization efficiency. This process aims to optimize power allocation through deep learning technology to improve the energy efficiency of the system and the communication quality of users. Among them, the fifth state includes the fourth state and the fourth action, and the third neural network is used to determine the power of the subcarriers where users are located.

[0198] For example, for the third neural network, its state not only needs to include the same part as the first neural network, but also needs to include the state input brought by the action executed by the first neural network and the input brought by the action executed by the second neural network, which is expressed as:

[0199]

[0200] Since the second neural network is for subcarrier allocation, so here needs to reflect the characteristics of allocating different subcarriers to different users. Therefore, design , where . When allocating subcarrier m to user N, the corresponding value is m.

[0201] This embodiment proposes a triple neural network architecture. In this network architecture, three neural networks are used for action training. The first neural network is used to judge the selection of wave positions, the second neural network is used for subcarrier allocation, and the third neural network is used for power allocation. Therefore, the action selections of the three networks are independent of each other, and then the dimension of the action space of each network will be greatly reduced. The first network can make an independent selection only based on the environmental information. The second network needs to receive the input from the first network, which contains the information of wave position selection, and makes subcarrier allocation by means of the environmental information and the wave position selection information together. Similarly, the third network needs to receive the environmental information as well as the information of the first and second networks to perform power allocation. The three networks are trained independently, but the reward function setting of the environment is composed of the rewards of the three networks. Therefore, it can be considered that there is a cooperative relationship among the three networks. Through cooperative training, the three networks will select an action plan with a large total reward value. Since the environmental information is changing and unknown, a model-free reinforcement learning algorithm is adopted, which enables the algorithm to be used offline after training, that is, without retraining, and still make reasonable judgments when facing a new environment. After being trained, this neural network model can be deployed at the signal sending end, so that the signal sending end can make its own action selection according to the ground user information.

[0202] S40234. Generate the second state according to the third state, the fourth state, and the fifth state.

[0203] Specifically, the third state, the fourth state, and the fifth state respectively contain information such as communication supply, demand, channel gain, wave position selection, subcarrier activation, and power allocation. By comprehensively analyzing these state information, a comprehensive second state is generated, which comprehensively describes the current network environment and resource configuration. The second state is used as an input in the neural network architecture to help the optimization algorithm more accurately evaluate and adjust the resource allocation strategy to improve the communication efficiency of the system and meet user needs.

[0204] S40235. Generate the second action according to the third action, the fourth action, and the fifth action.

[0205] Specifically, the third action, the fourth action, and the fifth action respectively correspond to the optimization decisions of wave position selection, subcarrier activation, and power allocation. By combining these independent decisions together, a comprehensive second action is generated, which comprehensively indicates the overall resource allocation strategy under the current network conditions. The second action is used to be executed in the optimization framework to actually adjust the resource configuration of the system, so as to improve communication efficiency, optimize resource utilization, and meet the communication needs of users.

[0206] S40236. Perform the second action and use the communication volume obtained after performing the second action as the second reward function.

[0207] Specifically, the second action includes a comprehensive wave position selection, subcarrier activation, and power allocation scheme. After performing this action, the system measures the actual communication volume obtained by the user within a preset time period. This communication volume is used as the second reward function, which reflects the effectiveness and performance of the current strategy. The second reward function is used to evaluate and optimize the strategy in the reinforcement learning framework. Through the reward feedback mechanism, the algorithm gradually improves resource allocation to enhance the overall communication efficiency and user satisfaction.

[0208] For example, the reward function is the immediate feedback received from the environment after performing the action, which can measure the quality of the current second action selection. To reduce the unmet communication volume of the user, the second reward function is designed as the communication volume provided to the user after performing the current step, that is

[0209]

[0210] To make the reward function as large as possible, it is necessary to try to make each selected second action provide the maximum communication volume to the user, which exactly meets the goal of minimizing the unmet communication volume of the user.

[0211] The technical effect of this solution in this embodiment is as follows: This method processes different levels of decision-making problems such as wave position selection, subcarrier activation, and power allocation through the first, second, and third neural networks respectively to form a comprehensive resource allocation strategy. Through this hierarchical processing, the system can more accurately adapt to the complex and changeable network environment, improve communication efficiency and resource utilization rate. In addition, by generating the second state and the second action and evaluating their reward functions, the system can continuously optimize the decision-making process under the reinforcement learning framework, enhancing the overall communication performance and user experience.

[0212] Figure 5 It is a schematic diagram of the change of the objective function of the bandwidth and power optimization method provided by the embodiment of the present application on the validation set. As Figure 5 shown, the abscissa represents the number of training rounds, and the ordinate represents the value of the objective function. The line shows the change trend of the objective function value as the number of training rounds increases. It can be seen that as the number of training rounds increases, the objective function value of the bandwidth and power optimization method also increases continuously and finally converges. Reaching convergence means that the objective function value of the bandwidth and power optimization method gradually stabilizes during the training process and no longer shows significant changes. This indicates that the model has found a stable solution, and further training will not significantly improve the performance or change the objective function value.

[0213] Figure 6Schematic diagram of the change of the loss function of the first neural network provided by the embodiments of this application with the number of training rounds. Figure 7 Schematic diagram of the change of the loss function of the second neural network provided by the embodiments of this application with the number of training rounds. Figure 8 Schematic diagram of the change of the loss function of the third neural network provided by the embodiments of this application with the number of training rounds. As Figures 6 to 8 shown, the abscissa represents the number of training rounds or iterations, and the ordinate represents the values of the loss functions of these three neural networks respectively. The lines show the change trends of the loss function values of the three neural networks during the training process. It can be seen that as the number of training rounds increases, the loss function value first increases and then decreases, indicating that the neural network may be exploring different strategies and action spaces in the initial stage, resulting in an increase in loss; subsequently, the network gradually learns better strategies, and the loss function value begins to decrease, indicating that the model is approaching the optimal solution.

[0214] Figure 9 Schematic diagram of the structure of the bandwidth and power optimization device provided by the embodiments of this application. As Figure 9 shown, the bandwidth and power optimization device includes:

[0215] A first acquisition module 901, configured to acquire the user location, the target wave position, the first variable, the second variable, the first power, and the first bandwidth; wherein, the target wave position refers to the position of the wave where the user is located among a plurality of preset waves, the first variable is used to represent the activation situation of the plurality of waves, the second variable is used to represent the activation situation of a plurality of preset subcarriers in the beam corresponding to the target wave, the first power is used to represent the real-time power of the subcarrier where the user is located, and the first bandwidth is used to represent the real-time bandwidth of the subcarrier where the user is located.

[0216] A first calculation module 902, configured to calculate the channel gain according to the user location, the target wave position, and a plurality of preset channel parameters.

[0217] A second calculation module 903, configured to calculate the communication supply according to the first variable, the second variable, the first power, the first bandwidth, and the channel gain; wherein, the communication supply refers to the amount of communication obtained by the user within a preset time period.

[0218] A first construction module 904, configured to construct an optimization model according to the communication supply and a preset communication demand.

[0219] A model solving module 905, configured to solve the optimization model according to a preset optimization algorithm to obtain an optimization solution; wherein, the optimization solution is used to improve the communication supply, and the optimization algorithm is constructed through reinforcement learning.

[0220] In a possible design, the second calculation module 903 includes:

[0221] A first calculation unit, configured to calculate a signal-to-interference-plus-noise ratio according to a first power and a channel gain.

[0222] A second calculation unit, configured to calculate a communication rate according to a first variable, a second variable, a first bandwidth, and the signal-to-interference-plus-noise ratio.

[0223] A third calculation unit, configured to calculate a communication supply according to the communication rate.

[0224] In a possible design, the bandwidth and power optimization device further includes:

[0225] A model establishment module, configured to establish a hopping beam resource optimization scenario model; wherein, the hopping beam resource optimization scenario model is used to represent multiple wave positions of a communication satellite, a beam corresponding to each wave position, and multiple subcarriers included in each beam.

[0226] The bandwidth and power optimization device provided in this embodiment may execute Figures 2 to 3 the technical solution of the method embodiment shown, and its implementation principle and technical effect are similar to Figures 2 to 3 the method embodiment shown, and will not be elaborated herein one by one.

[0227] Figure 10 It is a schematic structural diagram of a construction device for an optimization algorithm provided in an embodiment of the present application. As Figure 10 shown, the construction device for the optimization algorithm includes:

[0228] A second acquisition module 1001, configured to acquire a communication supply, a communication demand, a first variable, a second variable, a second power, and a channel gain; wherein, the second power refers to the power of each of a plurality of preset subcarriers in the beam corresponding to the target wave position, and the channel gain is used to represent the attenuation degree of the signal during transmission.

[0229] A second construction module 1002, configured to construct an optimization framework and a neural network architecture according to the communication supply, the communication demand, the first variable, the second variable, the second power, and the channel gain.

[0230] A third construction module 1003, configured to construct an optimization algorithm according to the optimization framework and the neural network architecture; wherein, the optimization algorithm is used for the bandwidth and power optimization device in the third aspect.

[0231] In a possible design, the second construction module 1002 includes:

[0232] The first construction unit is used to construct a first state, a first action, and a first reward function according to the communication supply volume, the communication demand volume, a first variable, a second variable, and a second power; wherein, the first state is used to represent the difference between the user's communication demand volume and the communication supply volume of the communication satellite, the first action is used to represent the activation conditions of a preset plurality of wave positions, the activation conditions of a preset plurality of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the first reward function is used to represent the communication volume obtained by the user after performing the first action, and the target wave position refers to the wave position where the user is located among the plurality of wave positions.

[0233] The second construction unit is used to construct an optimization framework according to the first state, the first action, and the first reward function.

[0234] The third construction unit is used to construct a second state, a second action, and a second reward function according to the communication supply volume, the communication demand volume, the first variable, the second variable, the second power, and the channel gain; wherein, the second state is used to represent the communication demand volume, the communication supply volume, and the channel gain, the second action is used to represent the activation conditions of a plurality of wave positions, the activation conditions of a preset plurality of subcarriers in the beam corresponding to the target wave position, and the power of each of the plurality of subcarriers, and the second reward function is used to represent the communication volume obtained by the user after performing the second action.

[0235] The fourth construction unit is used to construct a neural network architecture according to the second state, the second action, and the second reward function.

[0236] In a possible design, the first construction unit includes:

[0237] The first construction component is used to construct the first state according to the communication supply volume and the communication demand volume.

[0238] The second construction component is used to construct the first action according to the first variable, the second variable, and the second power.

[0239] The first execution component is used to execute the first action and take the communication volume obtained after executing the first action as the first reward function.

[0240] In a possible design, if the neural network architecture includes a first neural network, a second neural network, and a third neural network, then the third construction unit includes:

[0241] The first input component is used to input the third state into the first neural network to obtain a third action; wherein, the third state includes the communication supply volume, the communication demand volume, and the channel gain, and the first neural network is used to determine the wave position of the user.

[0242] A second input component for inputting a fourth state into a second neural network to obtain a fourth action, where the fourth state includes a third state and a third action, and the second neural network is used to determine the subcarriers of the user.

[0243] A third input component for inputting a fifth state into a third neural network to obtain a fifth action, where the fifth state includes the fourth state and the fourth action, and the third neural network is used to determine the power of the subcarriers where the user is located.

[0244] A first generation component for generating a second state based on the third state, the fourth state, and the fifth state.

[0245] A second generation component for generating a second action based on the third action, the fourth action, and the fifth action.

[0246] A second execution component for executing the second action and taking the traffic obtained after executing the second action as the second reward function.

[0247] The construction device of the optimization algorithm provided in this embodiment can execute Figure 4 the technical solutions of the method embodiment shown, and its implementation principle and technical effects are similar to Figure 4 the method embodiment shown, and will not be elaborated here one by one.

[0248] Figure 11 This is an electronic device provided in an embodiment of the present application. As Figure 11 shown, the electronic device includes at least one processor 1110 and a memory 1120. The electronic device also includes a communication component 1130. Among them, the processor 1110, the memory 1120, and the communication component 1130 are connected through a bus 1140.

[0249] In a specific implementation process, at least one processor 1110 executes computer execution instructions stored in the memory 1120, so that at least one processor 1110 is used to implement the bandwidth and power optimization method, the construction method of the optimization algorithm, the device, and the equipment in the above embodiments.

[0250] The specific implementation process of the processor 1110 can refer to the above method embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0251] In the above embodiments, it should be understood that the processor 1110 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor.

[0252] The memory 1120 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory.

[0253] The bus 1140 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 1140 in the drawings of this application is not limited to only one bus or one type of bus.

[0254] The functions implemented for the electronic device and the main control device are described above for the solution provided by the embodiments of the present invention. It can be understood that in order for the electronic device or the main control device to implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of each example described in the embodiments disclosed in the embodiments of the present invention, the embodiments of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present invention.

[0255] So far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

Claims

1. A bandwidth and power optimization method, characterized in that: The method is applied to a bandwidth and power optimization device, and the method comprises: Obtaining a user position, a target beam position, a first variable, a second variable, a first power, and a first bandwidth; wherein the target beam position refers to the position of the beam at which the user is located among multiple preset beam positions, the first variable is used to describe the activation status of the multiple beam positions, the second variable is used to describe the activation status of multiple subcarriers preset in the beam corresponding to the target beam position, the first power is used to describe the real-time power of the subcarrier at which the user is located, and the first bandwidth is used to describe the real-time bandwidth of the subcarrier at which the user is located; Calculating a channel gain according to the user position, the target wave position and a plurality of preset channel parameters; The communication provision amount is calculated according to the first variable, the second variable, the first power, the first bandwidth and the channel gain; wherein the communication provision amount refers to the communication amount obtained by the user within a preset time period; A model is constructed according to the communication provision amount and the preset communication demand amount to obtain an optimization model; According to a preset optimization algorithm, the optimization model is solved to obtain an optimization solution; wherein the optimization solution is used to improve the communication provision amount, and the optimization algorithm is constructed through reinforcement learning.

2. The bandwidth and power optimization method according to claim 1, characterized in that: The calculating the communication provision amount according to the first variable, the second variable, the first power, the first bandwidth and the channel gain includes: Calculating a signal to interference and noise ratio according to the first power and the channel gain; Calculate a communication rate according to the first variable, the second variable, the first bandwidth and the signal to interference and noise ratio; The communication provision amount is calculated according to the communication rate.

3. The bandwidth and power optimization method according to claim 1, characterized in that: Before obtaining the user position, the target wave position, the first variable, the second variable, the first power and the first bandwidth, the method further includes: A beam-hopping resource optimization scenario model is established; wherein the beam-hopping resource optimization scenario model is used to describe the multiple wave positions of the communication satellite, the beam corresponding to each wave position, and the multiple subcarriers included in each beam.

4. A method for constructing an optimization algorithm, characterized in that: The method is applied to a construction device of an optimization algorithm, and the method comprises: Obtaining the communication provision, the communication demand, the first variable, the second variable, the second power and the channel gain; wherein the second power refers to the power of each of the plurality of subcarriers preset in the beam corresponding to the target wave position, and the channel gain is used to express the attenuation degree of the signal during the transmission process; Constructing an optimization framework and a neural network architecture according to the communication provided amount, the communication demand amount, the first variable, the second variable, the second power and the channel gain; An optimization algorithm is constructed according to the optimization framework and the neural network architecture; wherein the optimization algorithm is used for the bandwidth and power optimization method described in any one of claims 1 to 3.

5. The method for constructing the optimization algorithm according to claim 4, characterized in that: The constructing of an optimization framework and a neural network architecture according to the communication provided amount, the communication demand amount, the first variable, the second variable, the second power and the channel gain includes: A first state, a first action and a first reward function are constructed according to the communication provided amount, the communication demand amount, the first variable, the second variable and the second power; wherein the first state is used to express the difference between the user's communication demand amount and the communication provided amount of the communication satellite, the first action is used to express the activation conditions of the preset multiple wave positions, the activation conditions of the preset multiple subcarriers in the beam corresponding to the target wave position and the power of each of the multiple subcarriers, the first reward function is used to express the communication amount obtained by the user after executing the first action, and the target wave position refers to the wave position where the user is located among the multiple wave positions; Constructing the optimization framework according to the first state, the first action and the first reward function; A second state, a second action, and a second reward function are constructed according to the communication provided amount, the communication demand amount, the first variable, the second variable, the second power, and the channel gain; wherein the second state is used to describe the communication demand amount, the communication provided amount, and the channel gain; the second action is used to describe the activation status of the multiple wave positions, the activation status of the multiple subcarriers preset in the beam corresponding to the target wave position, and the power of each of the multiple subcarriers; and the second reward function is used to describe the communication volume obtained by the user after executing the second action; The neural network architecture is constructed according to the second state, the second action and the second reward function.

6. The method for constructing the optimization algorithm according to claim 5, characterized in that: The step of constructing a first state, a first action and a first reward function according to the communication provided amount, the communication required amount, the first variable, the second variable and the second power includes: constructing the first state according to the communication provision amount and the communication demand amount; constructing the first action according to the first variable, the second variable and the second power; Execute the first action, and take the communication volume obtained after executing the first action as the first reward function.

7. The method for constructing the optimization algorithm according to claim 5, characterized in that: The neural network architecture includes a first neural network, a second neural network and a third neural network, and the second state, the second action and the second reward function are constructed according to the communication provided amount, the communication required amount, the first variable, the second variable, the second power and the channel gain, including: Inputting a third state into the first neural network to obtain a third action; wherein the third state includes the communication provision amount, the communication demand amount and the channel gain, and the first neural network is used to determine the wave position of the user; Inputting a fourth state into the second neural network to obtain a fourth action; wherein the fourth state includes the third state and the third action, and the second neural network is used to determine the subcarrier of the user; Inputting the fifth state into the third neural network to obtain a fifth action; wherein the fifth state includes the fourth state and the fourth action, and the third neural network is used to determine the power of the subcarrier where the user is located; generating the second state according to the third state, the fourth state and the fifth state; generating the second action according to the third action, the fourth action and the fifth action; Execute the second action, and take the communication volume obtained after executing the second action as the second reward function.

8. A bandwidth and power optimization device, characterized in that: include: A first acquisition module is used to acquire a user position, a target beam position, a first variable, a second variable, a first power, and a first bandwidth; wherein the target beam position refers to the position of the beam position where the user is located among multiple preset beam positions, the first variable is used to describe the activation status of the multiple beam positions, the second variable is used to describe the activation status of multiple subcarriers preset in the beam corresponding to the target beam position, the first power is used to describe the real-time power of the subcarrier where the user is located, and the first bandwidth is used to describe the real-time bandwidth of the subcarrier where the user is located; A first calculation module, configured to calculate a channel gain according to the user position, the target wave position and a plurality of preset channel parameters; A second calculation module is used to calculate the communication provision amount according to the first variable, the second variable, the first power, the first bandwidth and the channel gain; wherein the communication provision amount refers to the communication volume obtained by the user within a preset time period; A first construction module is used to construct a model according to the communication provision amount and the preset communication demand amount to obtain an optimization model; The model solving module is used to solve the optimization model according to a preset optimization algorithm to obtain an optimization solution; wherein the optimization solution is used to improve the communication provision amount, and the optimization algorithm is constructed through reinforcement learning.

9. A device for constructing an optimization algorithm, characterized in that: include: The second acquisition module is used to obtain the communication provision amount, the communication demand amount, the first variable, the second variable, the second power and the channel gain; wherein the second power refers to the power of each of the plurality of subcarriers preset in the beam corresponding to the target wave position, and the channel gain is used to express the attenuation degree of the signal during the transmission process; A second construction module is used to construct an optimization framework and a neural network architecture according to the communication provided amount, the communication required amount, the first variable, the second variable, the second power and the channel gain; The third building module is used to build an optimization algorithm based on the optimization framework and the neural network architecture; wherein the optimization algorithm is used for the bandwidth and power optimization device described in claim 8.

10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; When the processor executes the computer-executable instructions stored in the memory, it is used to implement the bandwidth and power optimization method as described in any one of claims 1 to 3, or to implement the method for constructing the optimization algorithm as described in any one of claims 4 to 7.