Adaptive reconstruction method and system for giant constellation network laser communication link

By collecting satellite node information in a giant constellation network to calculate link parameters and interference levels, and combining deep reinforcement learning to generate a reconstruction priority list, the dynamic adjustment of laser communication links and unbalanced resource utilization problems are solved, efficient adaptive reconstruction is achieved, and communication quality and network stability are improved.

CN120454861AInactive Publication Date: 2025-08-08ZKICME SUZHOU MICROELECTRONICS CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510830784.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The laser communication links in giant constellation networks face problems such as insufficient dynamic adjustment capabilities, unbalanced resource utilization and high computational complexity, resulting in degraded network performance and unstable communication quality.

Method used

By collecting the position and trajectory information of satellite nodes, calculating link parameters, generating resource state matrix, performing interference analysis, and generating reconstruction priority lists using deep reinforcement learning methods to achieve adaptive reconstruction.

Benefits of technology

It improves the reliability and stability of the network, optimizes resource utilization, extends the service life of the satellite network, and improves communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454861A_ABST
    Figure CN120454861A_ABST
Patent Text Reader

Abstract

The invention provides a giant constellation network laser communication link adaptive reconstruction method and system, and relates to the technical field of satellite communication, and the method comprises the steps: calculating link parameters through collecting satellite node positions and track information, generating a resource state matrix, executing interference analysis to determine an interference level, and generating a reconstruction priority list; and performing adaptive reconstruction on the target link by using deep reinforcement learning. Interference between satellites can be reduced, the network communication quality is improved, the energy utilization efficiency is optimized, and the communication reliability and stability of the constellation network are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite communication technology, and in particular to a method and system for adaptively reconfiguring laser communication links in a giant constellation network. Background Art

[0002] With the rapid development of space technology, mega-constellation networks have become a key development trend in next-generation satellite communication systems. These networks consist of hundreds or even thousands of low-orbit satellites, using laser communication links to achieve high-speed data transmission between satellites. Laser communication technology, with its high bandwidth, low latency, and high security, has become the preferred method for inter-satellite communication in mega-constellation networks. These constellation networks must maintain stable and efficient communication capabilities in a dynamically changing space environment. Therefore, adaptive reconfiguration of laser communication links is a key technology to ensure system performance.

[0003] Currently, laser communication links in giant constellation networks face a number of technical challenges. First, traditional laser communication link configuration schemes usually adopt static or semi-static topologies, which cannot dynamically adjust the link configuration according to the real-time network status and communication needs, resulting in a significant decrease in network performance when the satellite orbit changes or communication needs fluctuate. Secondly, existing link reconstruction methods lack a comprehensive consideration of the status of satellite resources and fail to fully balance the relationship between communication performance and energy consumption, causing some satellite nodes to fail prematurely due to resource exhaustion, affecting the service life of the entire constellation network. In addition, traditional reconstruction algorithms have high computational complexity and are difficult to adapt to the real-time decision-making needs of large-scale constellation networks. They cannot respond to interference events and changes in communication needs in a rapidly changing space environment in a timely manner, resulting in unstable network communication quality.

[0004] With the continuous expansion of giant constellation networks and the increasing complexity of application scenarios, there is an urgent need to develop an adaptive laser communication link reconstruction method that can comprehensively consider communication performance, resource consumption and interference factors, so as to improve the communication reliability, resource utilization efficiency and system robustness of the constellation network. Summary of the Invention

[0005] The embodiments of the present invention provide a method and system for adaptively reconfiguring a giant constellation network laser communication link, which can solve the problems in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a method for adaptively reconfiguring a laser communication link in a giant constellation network, comprising: Collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; Calculating the communication load and channel utilization of each satellite node based on the communication link parameters, monitoring the energy consumption status of each satellite node, and generating a resource status matrix of the satellite node; Perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with the preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and network communication requirements; According to the reconstruction priority list, a target laser communication link that needs to be reconstructed is selected. Using a deep reinforcement learning method, the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link are used as state inputs. The degree of communication quality improvement and energy consumption are used as reward functions. An optimization result is obtained through iterative optimization, and adaptive reconstruction is performed on the target laser communication link based on the optimization result.

[0007] According to the communication link parameters, the communication load and channel utilization of each satellite node are calculated, and the energy consumption status of each satellite node is monitored. The resource status matrix of the satellite node is generated, including: The ratio of the data transmission volume to the link capacity of each communication link is weighted and summed according to the corresponding link quality parameter to obtain a link load value. At the same time, the data queue length is multiplied by a preset queue weight coefficient to obtain a queue load value. The link load value and the queue load value are superimposed to obtain the communication load of the satellite node; Dynamically monitor the satellite node's spectrum resources within a preset time window, record the occupancy status of each frequency point at each sampling moment, mark the frequency point occupancy status according to the frequency point occupancy indicator, calculate the proportion of the bandwidth occupied by each frequency point to the total bandwidth, and accumulate and average the proportion in the time dimension and frequency dimension to obtain the channel utilization rate of the satellite node; Calculating energy consumption of satellite nodes based on the communication load and the channel utilization rate, and generating energy consumption status data; The communication load, the channel utilization and the energy consumption status data are respectively compared with their corresponding performance thresholds, the degree of deviation of each dimension is calculated, the degree of deviation is mapped into a deviation weight coefficient through an exponential function, and the communication load, the channel utilization and the energy consumption status data are integrated into a resource status matrix according to the deviation weight coefficient.

[0008] Generating a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication requirements includes: A triangular membership function is established for each dimension of the resource state matrix, and the state value of each dimension is mapped to the corresponding fuzzy linguistic variable space to obtain the fuzzified resource state matrix; Based on the fuzzified resource status matrix, a fuzzy rule base is constructed in the form of IF-THEN, the rule weight is obtained by calculating the product of the membership of the input variables in each rule, and the initial weight coefficient is obtained by combining the rule weight with the Mamdani fuzzy inference method; Encoding the initial weight coefficient and its corresponding membership function parameter and rule weight parameter into a chromosome, calculating the control error value and the smoothness value respectively and performing a weighted combination to generate a fitness evaluation value, and iteratively optimizing the chromosome based on the fitness evaluation value through a genetic algorithm to obtain an optimized weight parameter; Based on the optimized weight parameters, the resource status matrix, the interference level data and the network demand data are weighted and combined to obtain a link reconstruction priority index; The link reconstruction priority index is compared with the preset target priority threshold, the deviation value of each dimension is calculated, and the deviation value is multiplied by the corresponding compensation coefficient to obtain the compensation amount, and the compensation amount is added to the link reconstruction priority index to obtain the laser communication link reconstruction priority list.

[0009] The initial weight coefficient and its corresponding membership function parameter and rule weight parameter are encoded into a chromosome, the control error value and the smoothness value are calculated and weightedly combined to generate a fitness evaluation value, and based on the fitness evaluation value, the chromosome is iteratively optimized by a genetic algorithm to obtain the optimized weight parameters including: Arranging the initial weight coefficients, membership function parameters, and rule weight parameters into a weight vector, a parameter matrix, and a rule vector respectively, arranging the weight vector, the parameter matrix, and the rule vector in a preset order and converting them into a real number form to obtain a chromosome; Collecting the expected output value and actual output value of the control system at multiple consecutive sampling moments, calculating the root mean square error between the expected output value and the actual output value to obtain a control error value, recording the control quantities corresponding to the multiple consecutive sampling moments, and calculating the root mean square value of the difference between the control quantities at adjacent sampling moments to obtain a smoothness value; Multiplying the control error value by the inverse of the first fitness weight coefficient to obtain an error evaluation component, multiplying the smoothness value by the inverse of the second fitness weight coefficient to obtain a smooth evaluation component, and adding the error evaluation component and the smooth evaluation component to obtain a fitness evaluation value; Based on the fitness evaluation value, a parent chromosome is selected by a roulette wheel method, an arithmetic crossover operation and a Gaussian mutation operation are performed on the parent chromosome to generate a child chromosome, and the chromosome with the best fitness evaluation value is selected from the parent chromosome and the child chromosome as the next generation population, until the difference between the optimal fitness evaluation values of two adjacent generations is less than a preset optimal fitness evaluation value threshold, thereby obtaining the optimized weight parameter.

[0010] Selecting a parent chromosome by a roulette wheel method based on the fitness evaluation value, and performing an arithmetic crossover operation and a Gaussian mutation operation on the parent chromosome to generate a child chromosome includes: Calculate the selection probability of each chromosome according to the fitness evaluation value, and accumulate the selection probabilities in the order of chromosomes to obtain a cumulative probability sequence; Generate a random number and compare it with the cumulative probability sequence, select the chromosome whose cumulative probability is first greater than the random number as the parent chromosome, and repeat the selection process until two parent chromosomes are selected; Calculating a crossover coefficient based on a ratio of a current evolutionary generation to a maximum evolutionary generation, adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a first daughter chromosome, and adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a second daughter chromosome; The variable step length is calculated based on the ratio of the current evolutionary generation to the maximum evolutionary generation, the variable step length is multiplied by a random number that obeys the standard normal distribution to obtain the variation, and the variation is added to the first daughter chromosome and the second daughter chromosome respectively to obtain the mutated daughter chromosome.

[0011] Using the deep reinforcement learning method, the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link are used as state inputs, and the degree of communication quality improvement and energy consumption are used as reward functions. The optimization results obtained through iterative optimization include: Constructing a state vector based on link state parameters, resource state matrix, and historical reconstruction decision data; The local performance bonus value of a single link is obtained by subtracting the ratio of the current energy consumption to the maximum energy consumption from the ratio of the communication quality change at adjacent moments to the maximum communication quality. Multiplying the network connectivity index by the first global weight coefficient to obtain a connectivity evaluation component, multiplying the network throughput index by the second global weight coefficient to obtain a throughput evaluation component, and adding the connectivity evaluation component and the throughput evaluation component to obtain a network-level global reward value; Dynamically adjusting a dynamic weight coefficient during the training process, and performing a weighted combination of the local performance reward value and the network-level global reward value based on the dynamic weight coefficient to obtain a mixed reward value; A state-action value function is established using a deep neural network, the state vector is input into the deep neural network, the discounted sum of the mixed reward value and the maximum state-action value predicted by the target network is used as a temporal difference target, and the mean square error between the temporal difference target and the current state-action value is calculated to obtain a loss function value; Calculating the gradient of the network parameters based on the loss function value, updating the network parameters of the deep neural network by multiplying the gradient by the learning rate, and simultaneously updating the network parameters of the target network using the soft update coefficient; Repeat the above steps until the difference between two adjacent state action values is less than a preset difference threshold, and obtain the optimized link reconstruction strategy.

[0012] A second aspect of an embodiment of the present invention provides a system for adaptively reconfiguring a laser communication link in a giant constellation network, including: The first unit is used to collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; The second unit is configured to calculate the communication load and channel utilization of each satellite node based on the communication link parameters, monitor the energy consumption status of each satellite node, and generate a resource status matrix of the satellite node; The third unit is configured to perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with a preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication demand; The fourth unit is used to select a target laser communication link that needs to be reconstructed according to the reconstruction priority list, use a deep reinforcement learning method, take the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link as state input, use the degree of improvement in communication quality and energy consumption as reward functions, obtain an optimization result through iterative optimization, and perform adaptive reconstruction on the target laser communication link according to the optimization result.

[0013] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0014] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0015] The beneficial effects of this application are as follows: The method for adaptive reconstruction of laser communication links in a giant constellation network provided by the present invention achieves accurate calculation and dynamic monitoring of communication link parameters by acquiring satellite node positions and trajectory information in real time, effectively solving the problem that traditional reconstruction methods are difficult to adapt to highly dynamic space environments, and improving the reliability and stability of the network.

[0016] This method performs interference analysis in parallel in the edge computing unit and generates a reconstruction priority list based on the resource status matrix and network requirements. It achieves the rational allocation and efficient utilization of communication resources, reduces system energy consumption, extends the service life of the satellite network, and improves the communication quality of the overall network.

[0017] The deep reinforcement learning method is used to adaptively reconstruct the target laser communication link, enabling the system to continuously learn and optimize from historical decision data, forming an adaptive evolutionary network topology structure, which significantly enhances the adaptability and robustness of the giant constellation network in the face of complex space environments and changes in communication needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 1. A schematic diagram of a flow chart of a method for adaptively reconfiguring a laser communication link in a giant constellation network according to an embodiment of the present invention; Figure 2 This is a schematic diagram comparing the efficiency of resource utilization in the satellite-ground integrated network; Figure 3 Schematic diagram of the performance comparison of the fuzzy control system before and after genetic algorithm optimization. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0021] Figure 1 FIG. 1 is a flow chart of a method for adaptively reconfiguring a laser communication link in a giant constellation network according to an embodiment of the present invention. Figure 1 As shown, the method includes: Collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; Calculating the communication load and channel utilization of each satellite node based on the communication link parameters, monitoring the energy consumption status of each satellite node, and generating a resource status matrix of the satellite node; Perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with the preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and network communication requirements; According to the reconstruction priority list, a target laser communication link that needs to be reconstructed is selected. Using a deep reinforcement learning method, the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link are used as state inputs. The degree of communication quality improvement and energy consumption are used as reward functions. An optimization result is obtained through iterative optimization, and adaptive reconstruction is performed on the target laser communication link based on the optimization result.

[0022] In an optional embodiment, calculating the communication load and channel utilization of each satellite node based on the communication link parameters, and monitoring the energy consumption status of each satellite node to generate a resource status matrix of the satellite node includes: The ratio of the data transmission volume to the link capacity of each communication link is weighted and summed according to the corresponding link quality parameter to obtain a link load value. At the same time, the data queue length is multiplied by a preset queue weight coefficient to obtain a queue load value. The link load value and the queue load value are superimposed to obtain the communication load of the satellite node; Dynamically monitor the satellite node's spectrum resources within a preset time window, record the occupancy status of each frequency point at each sampling moment, mark the frequency point occupancy status according to the frequency point occupancy indicator, calculate the proportion of the bandwidth occupied by each frequency point to the total bandwidth, and accumulate and average the proportion in the time dimension and frequency dimension to obtain the channel utilization rate of the satellite node; Calculating energy consumption of satellite nodes based on the communication load and the channel utilization rate, and generating energy consumption status data; The communication load, the channel utilization and the energy consumption status data are respectively compared with their corresponding performance thresholds, the degree of deviation of each dimension is calculated, the degree of deviation is mapped into a deviation weight coefficient through an exponential function, and the communication load, the channel utilization and the energy consumption status data are integrated into a resource status matrix according to the deviation weight coefficient.

[0023] In a satellite-ground integrated network resource collaborative management system, it is necessary to accurately calculate the resource status matrix of satellite nodes to optimize resource allocation. Based on satellite communication link parameters, the system uses the following technical means to accurately assess the resource status of satellite nodes.

[0024] When calculating the satellite node's communication load, the system first calculates the ratio of the data transmission rate to the link capacity of each communication link. For example, a satellite node has three communication links with data transmission rates of 120 Mbps, 80 Mbps, and 60 Mbps, respectively, and link capacities of 200 Mbps, 150 Mbps, and 100 Mbps, respectively. The ratios are 0.6, 0.53, and 0.6, respectively. The link quality parameters are 0.9, 0.85, and 0.75, respectively. The system then weights these ratios with the link quality parameters to obtain the link load value: 0.6 × 0.9 + 0.53 × 0.85 + 0.6 × 0.75 = 1.35. The system also monitors the data queue lengths of the satellite node's three communication queues, which are 75, 50, and 25 packets, respectively. The preset queue weight coefficients are 0.4, 0.3, and 0.3, respectively. The sum of these factors yields the queue load value: 75 × 0.4 + 50 × 0.3 + 25 × 0.3 = 52.5. By adding the link load value of 1.35 and the queue load value of 52.5, the communication load of the satellite node is 53.85.

[0025] The system monitors the spectrum resource status of satellite nodes within a preset time window to ensure efficient resource utilization. The system sets a 10-second monitoring window and samples the spectrum every 0.5 seconds for a total of 20 times. Assume that the total bandwidth of the satellite node is 100 MHz, divided into 10 frequency bands, each with a bandwidth of 10 MHz. The system records the occupancy status of each frequency band using a frequency occupancy indicator, where 1 indicates occupied and 0 indicates idle. In one sampling period, the recorded frequency occupancy status is [1, 1, 1, 0, 1, 0, 0, 1, 1, 0], indicating that 6 frequencies are occupied and the occupied bandwidth accounts for 0.6% of the total bandwidth. The system averages the 20 sampling results over time. Assuming the average occupancy of each frequency band is [0.85, 0.75, 0.65, 0.45, 0.55, 0.4, 0.3, 0.7, 0.6, 0.25], and then averages the frequency band, yielding a channel utilization of 0.55 for this satellite node.

[0026] Based on the communication load and channel utilization, the system calculates the energy consumption status of the satellite node. Using an energy consumption model, the system substitutes a communication load of 53.85 and a channel utilization of 0.55. Assuming the satellite node's base energy consumption is 10W, each unit increase in communication load increases energy consumption by 0.15W, and each 0.1 increase in channel utilization increases energy consumption by 1.2W. The total energy consumption is 10 + 53.85 × 0.15 + 0.55 × 10 × 1.2 = 24.08W. The system compares this energy consumption value with the satellite battery capacity of 150W·h and estimates the remaining operating time to be 150 / 24.08, which is approximately 6.23 hours. The resulting energy consumption status data is [24.08W, 6.23h].

[0027] The system compares communication load, channel utilization, and energy consumption status data with performance thresholds to calculate the degree of deviation. Assuming a communication load threshold of 50, a channel utilization threshold of 0.6, and an energy consumption threshold of 20W, the degrees of deviation are (53.85 - 50) / 50 = 0.077, (0.6 - 0.55) / 0.6 = 0.083, and (24.08 - 20) / 20 = 0.204, respectively. The system uses an exponential function to map the degrees of deviation, specifically using the formula exp(degree of deviation × 10) - 1, to convert the degrees of deviation into a deviation weight coefficient, resulting in the values of 0.85, 0.92, and 2.26, respectively. The system multiplies the communication load 53.85 by the deviation weight coefficient 0.85 to get 45.77, the channel utilization 0.55 by the deviation weight coefficient 0.92 to get 0.506, and the energy consumption 24.08W by the deviation weight coefficient 2.26 to get 54.42, which are integrated into the resource state matrix [45.77, 0.506, 54.42].

[0028] For a satellite-ground integrated network with multiple communication nodes, the system repeats the above calculation process for each satellite node to generate a complete set of resource state matrices. For example, for a network consisting of four satellite nodes, the system generates the following resource state matrices: [[45.77, 0.506, 54.42], [38.25, 0.621, 45.36], [56.92, 0.483,62.14], [42.15, 0.572, 48.75]]. The system uses these resource state matrices to make resource allocation and scheduling decisions, optimizing the utilization of network resources.

[0029] Through these technical means, the system accurately assesses satellite node communication load, channel utilization, and energy consumption, generating a matrix that comprehensively reflects the resource status of satellite nodes, effectively supporting the coordinated resource management of integrated satellite-ground networks. This method adapts to the dynamic nature of satellite networks, ensuring communication quality, improving spectrum resource utilization efficiency, and extending the operating life of satellite nodes.

[0030] Figure 2 This is a schematic diagram comparing the efficiency of satellite-ground integrated network resource utilization: This figure compares the resource utilization efficiency of three different approaches in an integrated satellite-ground network. The horizontal axis represents constellation network size (1-8 satellite nodes), and the vertical axis represents resource utilization efficiency (0-60%). The present invention (circles) employs an adaptive laser communication link reconfiguration method based on deep reinforcement learning. By accurately calculating the communication load, channel utilization, and energy consumption of satellite nodes, it generates a resource state matrix to optimize resource allocation. Its efficiency steadily improves from approximately 36% to approximately 55% as the network scale expands, demonstrating excellent scalability. Traditional static allocation methods (squares) lack dynamic adjustment capabilities, and resource utilization efficiency steadily decreases with the addition of nodes, from approximately 24% to approximately 15%. Simple adaptive methods (triangles), while offering some dynamic adjustment capabilities, offer moderate efficiency, improving from approximately 30% to approximately 35%. The present invention performs optimally across various network sizes, with particularly significant advantages in large-scale deployments, demonstrating its superior adaptability and resource optimization capabilities in complex constellation network environments.

[0031] In an optional embodiment, generating a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication requirements includes: A triangular membership function is established for each dimension of the resource state matrix, and the state value of each dimension is mapped to the corresponding fuzzy linguistic variable space to obtain the fuzzified resource state matrix; Based on the fuzzified resource status matrix, a fuzzy rule base is constructed in the form of IF-THEN, the rule weight is obtained by calculating the product of the membership of the input variables in each rule, and the initial weight coefficient is obtained by combining the rule weight with the Mamdani fuzzy inference method; Encoding the initial weight coefficient and its corresponding membership function parameter and rule weight parameter into a chromosome, calculating the control error value and the smoothness value respectively and performing a weighted combination to generate a fitness evaluation value, and iteratively optimizing the chromosome based on the fitness evaluation value through a genetic algorithm to obtain an optimized weight parameter; Based on the optimized weight parameters, the resource status matrix, the interference level data and the network demand data are weighted and combined to obtain a link reconstruction priority index; The link reconstruction priority index is compared with the preset target priority threshold, the deviation value of each dimension is calculated, and the deviation value is multiplied by the corresponding compensation coefficient to obtain the compensation amount, and the compensation amount is added to the link reconstruction priority index to obtain the laser communication link reconstruction priority list.

[0032] In this implementation, to achieve intelligent reconstruction of laser communication links, the system generates a priority list based on interference levels, resource status, and network communication requirements. This method first acquires a resource status matrix, which contains multiple dimensions of resource status information, such as link availability, power headroom, and signal quality. The system also collects interference level data, which uses photoelectric sensors to monitor the intensity of interference factors such as atmospheric turbulence and cloud cover in real time. Furthermore, the system collects network communication demand data, including parameters such as service priority, bandwidth requirements, and latency requirements.

[0033] To process the resource status matrix, the system establishes a triangular membership function for each dimension. For example, when the link available time is 300 seconds, the system maps it to a "medium" linguistic variable with a corresponding membership of 0.75; when the power headroom is 5dB, it is mapped to a "low" linguistic variable with a membership of 0.62; and when the signal quality is 85%, it is mapped to a "high" linguistic variable with a membership of 0.83. In this way, the system completes the fuzzification of the resource status matrix, resulting in the fuzzified resource status matrix.

[0034] Based on the fuzzified resource status matrix, the system constructs a fuzzy rule base using an if-then approach. For example, Rule 1 can be formulated as "If the link availability time is long, the power headroom is high, and the signal quality is excellent, then the link reconstruction priority is low"; Rule 2 can be formulated as "If the link availability time is short, the power headroom is low, and the signal quality is poor, then the link reconstruction priority is high." The system calculates rule weights by multiplying the membership degrees of the input variables in each rule. For Rule 1, assuming the membership degrees of link availability time, power headroom, and signal quality are 0.8, 0.7, and 0.9, respectively, the rule weight is 0.8 × 0.7 × 0.9 = 0.504. The system uses these rule weights and the Mamdani fuzzy inference method to calculate the initial weight coefficients. Specifically, the system combines the conclusions of each rule according to their corresponding weights to obtain the final fuzzy output set. This is then defuzzified using the centroid method to obtain an initial weight coefficient, for example, 0.65.

[0035] To optimize the weight parameters, the system encodes the initial weight coefficients and their corresponding membership function parameters and rule weight parameters into chromosomes. A typical chromosome structure consists of the three parameters of the triangular membership function (left endpoint, center, and right endpoint) and the weights of each rule. The system calculates the control error and smoothness values separately and then weights them to generate a fitness evaluation. The control error measures the difference between the system output and the expected output, while the smoothness value reflects the stability of the system response. Assuming the system calculates a control error of 0.15 and a smoothness of 0.08, with weights of 0.7 and 0.3, respectively, the fitness evaluation is 0.15 × 0.7 + 0.08 × 0.3 = 0.129. Based on this fitness evaluation, the system iteratively optimizes the chromosome using a genetic algorithm. During the optimization process, the system uses a roulette wheel selection algorithm to select outstanding individuals, employs single-point crossover for gene exchange, and performs mutation with a probability of 0.05 to achieve population evolution. After 100 generations of iterative optimization, the system obtained optimized weight parameters, such as membership function parameters [0.2, 0.5, 0.8] and rule weights [0.35, 0.42, 0.23].

[0036] Based on the optimized weight parameters, the system performs a weighted combination of the resource status matrix, interference level data, and network demand data to generate a link reconstruction priority index. Specifically, the system multiplies the resource status value by the corresponding weight and sums the results. This sum is then combined with the weighted values of the interference level and network demand to generate the priority index. For example, if a link has a resource status score of 0.72, an interference level score of 0.85, and a network demand score of 0.63, and the corresponding weights are [0.4, 0.35, 0.25], the reconstruction priority index for this link is 0.72 × 0.4 + 0.85 × 0.35 + 0.63 × 0.25 = 0.7425.

[0037] The system then compares the link reconstruction priority index with the preset target priority threshold and calculates the deviation value of each dimension. Assuming the target priority threshold is 0.8, the deviation value of the above link is 0.8-0.7425=0.0575. The system multiplies this deviation value with the corresponding compensation coefficient (such as 1.2) to obtain a compensation amount of 0.069, and adds this compensation amount to the link reconstruction priority index to obtain a final priority value of 0.7425+0.069=0.8115. In this way, the system performs the same calculation process on all links and sorts them according to the size of the final priority value to form a laser communication link reconstruction priority list. This list intuitively displays the reconstruction priority of each link, making it easier for the system to perform link reconstruction operations according to the priority order, ensuring that the communication quality of high-priority links is prioritized under resource-constrained conditions.

[0038] Practice has shown that this method, through the combination of fuzzy logic and genetic algorithms, can effectively cope with the uncertainties in laser communication environments, realize intelligent scheduling and optimal configuration of link resources, and significantly improve the communication reliability and service quality of the system in complex environments.

[0039] In an optional embodiment, the initial weight coefficient and its corresponding membership function parameter and rule weight parameter are encoded as a chromosome, the control error value and the smoothness value are calculated and weighted together to generate a fitness evaluation value, and based on the fitness evaluation value, the chromosome is iteratively optimized by a genetic algorithm to obtain the optimized weight parameters including: Arranging the initial weight coefficients, membership function parameters, and rule weight parameters into a weight vector, a parameter matrix, and a rule vector respectively, arranging the weight vector, the parameter matrix, and the rule vector in a preset order and converting them into a real number form to obtain a chromosome; Collecting the expected output value and actual output value of the control system at multiple consecutive sampling moments, calculating the root mean square error between the expected output value and the actual output value to obtain a control error value, recording the control quantities corresponding to the multiple consecutive sampling moments, and calculating the root mean square value of the difference between the control quantities at adjacent sampling moments to obtain a smoothness value; Multiplying the control error value by the inverse of the first fitness weight coefficient to obtain an error evaluation component, multiplying the smoothness value by the inverse of the second fitness weight coefficient to obtain a smooth evaluation component, and adding the error evaluation component and the smooth evaluation component to obtain a fitness evaluation value; Based on the fitness evaluation value, a parent chromosome is selected by a roulette wheel method, an arithmetic crossover operation and a Gaussian mutation operation are performed on the parent chromosome to generate a child chromosome, and the chromosome with the best fitness evaluation value is selected from the parent chromosome and the child chromosome as the next generation population, until the difference between the optimal fitness evaluation values of two adjacent generations is less than a preset optimal fitness evaluation value threshold, thereby obtaining the optimized weight parameter.

[0040] In order to implement the method of optimizing fuzzy control weight parameters based on genetic algorithm described in the present invention, a description will be given below in conjunction with specific embodiments.

[0041] This embodiment uses a genetic algorithm to optimize the weight coefficients, membership function parameters, and rule weight parameters of a fuzzy control system. During the chromosome encoding phase, the initial weight coefficients are represented as a weight vector W = [w1, w2, ..., wn], where n is the number of weight coefficients. The membership function parameters are organized into a parameter matrix M, where each element represents the parameter of a specific membership function for a specific input or output variable. The rule weight parameters are represented as a rule vector R = [r1, r2, ..., rm], where m is the number of fuzzy rules. During chromosome encoding, W, M, and R are arranged in the order of "weight vector - parameter matrix - rule vector" and converted into real numbers. For example, assuming there are 3 weight coefficients, 2 input variables each with 3 membership functions, each with 2 parameters, 5 output variables with 2 parameters each, and 9 fuzzy rules, then the chromosome can be represented as [w1, w2, w3, m11, m12, ..., m(2×3×2+5×2), r1, r2, ..., r9], with a total length of 3+(2×3×2+5×2)+9=34 real numbers.

[0042] The fitness evaluation phase involves two aspects of control system performance: control error and control smoothness. The control error reflects control accuracy, while the smoothness reflects the severity of changes in the controlled variable. To calculate the control error, the expected output and actual output values are collected at multiple consecutive sampling points during system operation. Assuming 100 sampling points are collected, the expected output value sequence is [y1d, y2d, ..., y100d], and the actual output value sequence is [y1, y2, ..., y100]. The control error value Ec is obtained by calculating the sum of the squares of the differences between the expected and actual output values, averaging them, and taking the square root. Specifically, the squared errors (y1d-y1)², (y2d-y2)², ..., (y100d-y100)² at each sampling point are calculated, the sum is divided by 100, and the square root is taken to obtain the control error value Ec.

[0043] At the same time, record the control quantity sequence [u1, u2, ..., u100] corresponding to these 100 sampling moments, calculate the square values of the control quantity differences between adjacent sampling moments (u2-u1)², (u3-u2)², ..., (u100-u99)², sum them up and divide by 99, then take the square root to obtain the smoothness value Es.

[0044] The calculation of the fitness evaluation value takes into account both control error and smoothness. Assuming the first fitness weight coefficient is λ1 and the second fitness weight coefficient is λ2, the error evaluation component is multiplied by 1 / λ1, and the smoothness evaluation component is multiplied by 1 / λ2. The smoothness evaluation component is then added to obtain the fitness evaluation value J. For example, when λ1 = 0.7, λ2 = 0.3, the control error value Ec = 0.05, and the smoothness evaluation value Es = 0.02, the error evaluation component is 0.05 / 0.7 ≈ 0.0714, the smoothness evaluation component is 0.02 / 0.3 ≈ 0.0667, and the fitness evaluation value J is ≈ 0.1381. A smaller fitness evaluation value indicates better control performance.

[0045] The optimization phase of the genetic algorithm uses a population iteration method to search for the optimal solution. Assuming the initial population size is 50, each individual is a chromosome, representing a set of weight parameters. The selection operation uses the roulette wheel method, and the individual with the smaller fitness evaluation value has a greater probability of being selected. Specifically, the selection probability of each individual is calculated, and the selection probability is proportional to the inverse of its fitness evaluation value. For example, if the fitness evaluation value of individual i is J i , then its selection probability can be expressed as (1 / J i ) / ∑(1 / J j ), where j = 1, 2, ..., 50. In this way, 50 pairs of parent chromosomes are selected.

[0046] Perform an arithmetic crossover on the selected parent chromosomes. Set the crossover probability to 0.8. For each pair of parent chromosomes, if the randomly generated number is less than 0.8, a crossover is performed. This arithmetic crossover generates two daughter chromosomes, each containing a weighted average of the genes at the corresponding position in the parent chromosome. For example, if the gene values at a certain position on the two parent chromosomes are a and b, respectively, then after crossover, the gene values at the corresponding positions in the two daughter chromosomes will be 0.3a + 0.7b and 0.7a + 0.3b, respectively, where 0.3 and 0.7 are the crossover weights.

[0047] Perform a Gaussian mutation on the offspring chromosomes after the crossover. Set the mutation probability to 0.1. For each gene position on the chromosome, if the randomly generated value is less than 0.1, the mutation is performed. Gaussian mutation is achieved by adding a random perturbation that follows a Gaussian distribution to the original gene value. For example, if the value of a gene position is v, the mutated value is v + δ, where δ is a Gaussian random number with mean 0 and standard deviation 10% of the original range.

[0048] The 50 chromosomes with the best fitness values are selected from the parent and offspring chromosomes to form the next generation population. The iterative process continues until the termination condition is met: the difference between the best fitness values of two consecutive generations is less than a preset threshold (such as 0.0001). The optimal chromosomes are decoded to obtain the optimized weight coefficients, membership function parameters, and rule weight parameters, which are used in the fuzzy control system.

[0049] In this example, through actual application testing, the initially randomly generated parameters resulted in a control error of 0.15, a smoothness of 0.08, and a fitness evaluation of approximately 0.48. After optimization using the genetic algorithm described above, the control error dropped to 0.02, the smoothness dropped to 0.01, and the fitness evaluation dropped to approximately 0.06, significantly improving control performance.

[0050] Figure 3 This is a schematic diagram comparing the performance of the fuzzy control system before and after genetic algorithm optimization: The figure provides a quantitative comparison of three key dimensions: control error, smoothness, and fitness evaluation. In terms of control error, the error was 0.15 under the initial parameter configuration, which was reduced to 0.06 using conventional optimization methods and further reduced to 0.02 after genetic algorithm optimization, indicating a significant improvement in control accuracy. The smoothness value, which assesses the stability of the control system, was 0.08 under the initial parameter configuration, but dropped to 0.04 after conventional optimization and 0.01 after genetic algorithm optimization, demonstrating smoother system operation. The fitness evaluation value, a comprehensive indicator that takes into account both control accuracy and smoothness, was 0.48 under the initial configuration, but dropped to 0.18 after conventional optimization and 0.06 after genetic algorithm optimization, a decrease of 87.5%. This data intuitively demonstrates the superior performance of genetic algorithms in parameter optimization of fuzzy control systems.

[0051] In an optional embodiment, selecting a parent chromosome by a roulette wheel method based on the fitness evaluation value, and performing an arithmetic crossover operation and a Gaussian mutation operation on the parent chromosome to generate a child chromosome includes: Calculate the selection probability of each chromosome according to the fitness evaluation value, and accumulate the selection probabilities in the order of chromosomes to obtain a cumulative probability sequence; Generate a random number and compare it with the cumulative probability sequence, select the chromosome whose cumulative probability is first greater than the random number as the parent chromosome, and repeat the selection process until two parent chromosomes are selected; Calculating a crossover coefficient based on a ratio of a current evolutionary generation to a maximum evolutionary generation, adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a first daughter chromosome, and adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a second daughter chromosome; The variable step length is calculated based on the ratio of the current evolutionary generation to the maximum evolutionary generation, the variable step length is multiplied by a random number that obeys the standard normal distribution to obtain the variation, and the variation is added to the first daughter chromosome and the second daughter chromosome respectively to obtain the mutated daughter chromosome.

[0052] In this implementation, an adaptive optimization method based on a genetic algorithm is proposed. This method uses a roulette wheel method to select parent chromosomes and generates offspring chromosomes through arithmetic crossover and Gaussian mutation operations. The implementation process and key technical points are described in detail below.

[0053] The core of genetic algorithms lies in solving optimization problems by simulating natural selection and heredity. After initializing the population, the fitness of each chromosome needs to be evaluated, and high-quality individuals are selected for reproduction based on their fitness values. This implementation uses a roulette wheel selection mechanism, which determines the probability of a chromosome being selected based on the proportion of its fitness to the total fitness.

[0054] In the specific implementation, assume that the population size is 100 and the fitness evaluation value of each chromosome has been calculated by the objective function. To perform roulette selection, first calculate the selection probability of each chromosome. For example, if the fitness value of the i-th chromosome is F i , the total fitness is the sum of all chromosome fitness Sum F , then the selection probability of the chromosome P i F i Divide by Sum F For a minimization problem, the probability of selection can be converted by taking the inverse of the fitness value or the maximum fitness minus the current fitness plus one.

[0055] Next, we construct a cumulative probability sequence, which is the accumulation of selection probabilities according to the order of chromosomes. For example, the cumulative probability C1 of the first chromosome is equal to its selection probability P1, the cumulative probability C2 of the second chromosome is equal to C1 plus P2, and so on. The cumulative probability of the last chromosome should be 1.

[0056] To select a parent chromosome, generate a random number R between 0 and 1 and compare it to the cumulative probability sequence. The chromosome whose cumulative probability first exceeds R is selected as the parent. For example, if the generated random number R is 0.35, and C5 and C6 in the cumulative probability sequence are 0.32 and 0.38, the sixth chromosome is selected as the parent. Repeat this process to select two parent chromosomes, designated Parent1 and Parent2.

[0057] After obtaining the parent chromosome, an arithmetic crossover operation is performed. The arithmetic crossover coefficient α is dynamically adjusted to allow for greater exploration capabilities in the early stages of evolution and more refined local search capabilities in the later stages. The calculation of α is based on the current evolutionary generation number current gen and the maximum evolutionary generation max gen The ratio of: α is equal to 0.9 minus 0.8 times the current gen Divide by max gen For example, if current gen is 50, max gen If α is 200, then α is 0.9 minus 0.8 times 0.25, which is 0.7.

[0058] Using the calculated crossover coefficient α, two daughter chromosomes are generated. The first daughter chromosome, Child1, is calculated as α multiplied by Parent1 plus (1-α) multiplied by Parent2. The second daughter chromosome, Child2, is calculated as (1-α) multiplied by Parent1 plus α multiplied by Parent2. For example, if Parent1 is [0.5, 0.8, 0.3], Parent2 is [0.2, 0.6, 0.9], and α is 0.7, then Child1 is [0.5×0.7+0.2×0.3, 0.8×0.7+0.6×0.3, 0.3×0.7+0.9×0.3], that is, [0.41, 0.74, 0.48]; Child2 is [0.5×0.3+0.2×0.7, 0.8×0.3+0.6×0.7,0.3×0.3+0.9×0.7], that is, [0.29, 0.66, 0.72].

[0059] Next, the Gaussian mutation operation is performed to prevent the algorithm from falling into the local optimum and to balance the global exploration and local development capabilities. The variable step length σ is also dynamically adjusted and is calculated as 0.5 multiplied by (1-current gen / max gen ). When current gen is 50, max gen When σ is 200, σ is 0.5 multiplied by (1-0.25), which is 0.375.

[0060] For each gene on each daughter chromosome, generate a random number N(0,1) from a standard normal distribution and calculate the mutation as σ multiplied by N(0,1). Add the mutation to the gene value to obtain the mutated gene value. For example, if the first gene of Child1 is 0.41, the generated random number is 0.5, and σ is 0.375, then the mutation is 0.375 × 0.5 = 0.1875, and the mutated gene value is 0.41 + 0.1875 = 0.5975. Perform the same operation on all genes in Child1 and Child2 to obtain the mutated daughter chromosomes.

[0061] To ensure that the mutated gene value is within the valid range, it is necessary to check and adjust the out-of-bounds values. Assuming the valid range of gene values is [0,1], if the mutated value is less than 0, it is set to 0; if it is greater than 1, it is set to 1.

[0062] Through the above roulette wheel selection, arithmetic crossover, and Gaussian mutation operations, the evolutionary process from parent to offspring is completed. This process is repeated until the population size reaches a preset value, such as 100 individuals. In each evolutionary generation, all chromosomes with higher fitness have a chance to be selected for reproduction, while chromosomes with lower fitness have a lower probability of being selected, thus implementing the natural selection principle of survival of the fittest. Furthermore, the dynamically adjusted crossover coefficient and variable step length ensure a good balance between exploration and exploitation during the evolutionary process, effectively improving the algorithm's convergence performance and solution quality.

[0063] In an optional embodiment, a deep reinforcement learning method is used, with the link state parameters, resource state matrix, and historical reconstruction decision data of the target laser communication link as state inputs, and the degree of communication quality improvement and energy consumption as reward functions, and the optimization results obtained through iterative optimization include: Constructing a state vector based on link state parameters, resource state matrix, and historical reconstruction decision data; The local performance bonus value of a single link is obtained by subtracting the ratio of the current energy consumption to the maximum energy consumption from the ratio of the communication quality change at adjacent moments to the maximum communication quality. Multiplying the network connectivity index by the first global weight coefficient to obtain a connectivity evaluation component, multiplying the network throughput index by the second global weight coefficient to obtain a throughput evaluation component, and adding the connectivity evaluation component and the throughput evaluation component to obtain a network-level global reward value; Dynamically adjusting a dynamic weight coefficient during the training process, and performing a weighted combination of the local performance reward value and the network-level global reward value based on the dynamic weight coefficient to obtain a mixed reward value; A state-action value function is established using a deep neural network, the state vector is input into the deep neural network, the discounted sum of the mixed reward value and the maximum state-action value predicted by the target network is used as a temporal difference target, and the mean square error between the temporal difference target and the current state-action value is calculated to obtain a loss function value; Calculating the gradient of the network parameters based on the loss function value, updating the network parameters of the deep neural network by multiplying the gradient by the learning rate, and simultaneously updating the network parameters of the target network using the soft update coefficient; Repeat the above steps until the difference between two adjacent state action values is less than a preset difference threshold, and obtain the optimized link reconstruction strategy.

[0064] The present invention discloses a laser communication link reconstruction optimization method based on deep reinforcement learning. In this method, the link state parameters, resource state matrix, and historical reconstruction decision data of the target laser communication link are first processed to construct a state vector. The link state parameters include indicators such as the current link's signal-to-noise ratio, atmospheric attenuation index, and jitter amplitude, represented as 8-bit floating-point numbers; the resource state matrix contains resource information such as the number of available wavelengths, power level, and modulation mode, using a 16×16 matrix structure; and the historical reconstruction decision data records the link reconstruction decisions of the previous five time steps, forming a historical decision vector of length 15. These data are combined to form a state vector with a dimension of 256, which serves as the input of the deep neural network.

[0065] In the reward function design, the local performance reward for a single link is calculated using the difference between the change in communication quality and the energy consumption. Specifically, the local performance reward is calculated by taking the ratio of the change in communication quality between two adjacent moments to the maximum communication quality, and then subtracting the ratio of the current energy consumption to the maximum energy consumption. For example, if the communication quality of a link improves from 30Mbps to 45Mbps, the maximum communication quality is 100Mbps, the current energy consumption is 2W, and the maximum energy consumption is 10W, the local performance reward is (45-30) / 100-2 / 10=0.15-0.2=-0.05.

[0066] The calculation of the network-level global reward combines two key metrics: network connectivity and throughput. The network connectivity metric is calculated by calculating the connectivity ratio between pairs of nodes, ranging from 0 to 1. The network throughput metric is the ratio of the current total network throughput to the theoretical maximum throughput. The connectivity metric is multiplied by the first global weight coefficient of 0.6 to obtain the connectivity evaluation component, and the throughput metric is multiplied by the second global weight coefficient of 0.4 to obtain the throughput evaluation component. The two are added together to obtain the network-level global reward. For example, when the network connectivity is 0.85 and the network throughput is 0.7 of the current theoretical maximum, the network-level global reward is 0.85 × 0.6 + 0.7 × 0.4 = 0.79.

[0067] To balance local and global optimization goals, this method introduces dynamic weight coefficients for reward blending. Initially, the local reward weight is set to 0.8, and the global reward weight is 0.2. As training epochs increase, the local reward weight decreases linearly to 0.3, while the global reward weight increases to 0.7. The blended reward is calculated as: local performance reward × local reward weight + network-level global reward × global reward weight. At the 500th epoch of training, if the local reward weight is 0.6, the global reward weight is 0.4, the local performance reward is 0.25, and the global reward is 0.8, then the blended reward is 0.25 × 0.6 + 0.8 × 0.4 = 0.47.

[0068] The deep neural network architecture uses a four-layer fully connected network to implement the state-action-value function. The input layer receives a 256-dimensional state vector. Hidden layer 1 contains 512 neurons, hidden layer 2 contains 256 neurons, and hidden layer 3 contains 128 neurons. The output layer outputs the expected value of each action, corresponding to the number of possible actions. The activation function uses the Reluctant Unified Unit (ReLU) function. During training, the temporal difference (TD) target is calculated by adding the current reward value to a discount factor (set to 0.95) and multiplying it by the target network's predicted maximum action value for the next state. For example, if the current mixed reward value is 0.47 and the target network's predicted maximum action value for the next state is 0.82, the TD target is 0.47 + 0.95 × 0.82 = 1.249.

[0069] The loss function uses the mean squared error (MSE) between the temporal difference target and the state-action value predicted by the current network. If the state-action value predicted by the current network is 1.08 and the temporal difference target is 1.249, the loss function value is (1.249 - 1.08)² = 0.028561. Based on this loss value, the Adam optimizer is used to calculate the network parameter gradients, with a learning rate set to 0.0001. The product of the gradient and the learning rate is used to update the deep neural network parameters. Simultaneously, a soft update mechanism is used for the target network, with a soft update coefficient set to 0.01. This means that the target network parameters are updated as follows: target network parameters × 0.99 + current network parameters × 0.01.

[0070] During the training iteration process, the model performance is evaluated every 50 rounds, and the average difference in the state action values of two adjacent evaluations is calculated. When the difference is less than the preset threshold of 0.001, the model is determined to have converged, and the optimized link reconstruction strategy is output. In a laser communication network containing 25 nodes, after 2000 rounds of training using this method, the average network throughput increased by 37% and the energy efficiency increased by 25%. Compared with the traditional fixed threshold switching strategy and rule-based dynamic adjustment method, the overall performance is significantly improved. The final output link reconstruction strategy can give the optimal link reconstruction decision in real time based on the input link state parameters, resource state matrix and historical decision data, including parameter configuration such as wavelength selection, power allocation, modulation mode and coding rate.

[0071] A system for adaptively reconfiguring laser communication links in a giant constellation network according to an embodiment of the present invention includes: The first unit is used to collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; The second unit is configured to calculate the communication load and channel utilization of each satellite node based on the communication link parameters, monitor the energy consumption status of each satellite node, and generate a resource status matrix of the satellite node; The third unit is configured to perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with a preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication demand; The fourth unit is used to select a target laser communication link that needs to be reconstructed according to the reconstruction priority list, use a deep reinforcement learning method, take the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link as state input, use the degree of improvement in communication quality and energy consumption as reward functions, obtain an optimization result through iterative optimization, and perform adaptive reconstruction on the target laser communication link according to the optimization result.

[0072] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0073] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0074] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for adaptively reconfiguring laser communication links in a giant constellation network, characterized in that: include: Collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; Calculating the communication load and channel utilization of each satellite node based on the communication link parameters, while monitoring the energy consumption status of each satellite node to generate a resource status matrix of the satellite node; Perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with the preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and network communication requirements; According to the reconstruction priority list, a target laser communication link that needs to be reconstructed is selected. Using a deep reinforcement learning method, the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link are used as state inputs. The degree of communication quality improvement and energy consumption are used as reward functions. An optimization result is obtained through iterative optimization, and adaptive reconstruction is performed on the target laser communication link based on the optimization result.

2. The method according to claim 1, characterized in that According to the communication link parameters, the communication load and channel utilization of each satellite node are calculated, and the energy consumption status of each satellite node is monitored. The resource status matrix of the satellite node is generated, including: The ratio of the data transmission volume to the link capacity of each communication link is weighted and summed according to the corresponding link quality parameter to obtain a link load value. At the same time, the data queue length is multiplied by a preset queue weight coefficient to obtain a queue load value. The link load value and the queue load value are superimposed to obtain the communication load of the satellite node; Dynamically monitor the satellite node's spectrum resources within a preset time window, record the occupancy status of each frequency point at each sampling moment, mark the frequency point occupancy status according to the frequency point occupancy indicator, calculate the proportion of the bandwidth occupied by each frequency point to the total bandwidth, and accumulate and average the proportion in the time dimension and frequency dimension to obtain the channel utilization rate of the satellite node; Calculating energy consumption of satellite nodes based on the communication load and the channel utilization rate, and generating energy consumption status data; The communication load, the channel utilization and the energy consumption status data are respectively compared with their corresponding performance thresholds, the degree of deviation of each dimension is calculated, the degree of deviation is mapped into a deviation weight coefficient through an exponential function, and the communication load, the channel utilization and the energy consumption status data are integrated into a resource status matrix according to the deviation weight coefficient.

3. The method according to claim 1, characterized in that Generating a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication requirements includes: A triangular membership function is established for each dimension of the resource state matrix, and the state value of each dimension is mapped to the corresponding fuzzy linguistic variable space to obtain the fuzzified resource state matrix; Based on the fuzzified resource status matrix, a fuzzy rule base is constructed in the form of IF-THEN, the rule weight is obtained by calculating the product of the membership of the input variables in each rule, and the initial weight coefficient is obtained by combining the rule weight with the Mamdani fuzzy inference method; Encoding the initial weight coefficient and its corresponding membership function parameter and rule weight parameter into a chromosome, calculating the control error value and the smoothness value respectively and performing a weighted combination to generate a fitness evaluation value, and iteratively optimizing the chromosome based on the fitness evaluation value through a genetic algorithm to obtain an optimized weight parameter; Based on the optimized weight parameters, the resource status matrix, the interference level data and the network demand data are weighted and combined to obtain a link reconstruction priority index; The link reconstruction priority index is compared with the preset target priority threshold, the deviation value of each dimension is calculated, and the deviation value is multiplied by the corresponding compensation coefficient to obtain the compensation amount, and the compensation amount is added to the link reconstruction priority index to obtain the laser communication link reconstruction priority list.

4. The method according to claim 3, characterized in that The initial weight coefficient and its corresponding membership function parameter and rule weight parameter are encoded into a chromosome, the control error value and the smoothness value are calculated and weightedly combined to generate a fitness evaluation value, and based on the fitness evaluation value, the chromosome is iteratively optimized by a genetic algorithm to obtain the optimized weight parameters including: Arranging the initial weight coefficients, membership function parameters, and rule weight parameters into a weight vector, a parameter matrix, and a rule vector respectively, arranging the weight vector, the parameter matrix, and the rule vector in a preset order and converting them into a real number form to obtain a chromosome; Collecting the expected output value and actual output value of the control system at multiple consecutive sampling moments, calculating the root mean square error between the expected output value and the actual output value to obtain a control error value, recording the control quantities corresponding to the multiple consecutive sampling moments, and calculating the root mean square value of the difference between the control quantities at adjacent sampling moments to obtain a smoothness value; Multiplying the control error value by the inverse of the first fitness weight coefficient to obtain an error evaluation component, multiplying the smoothness value by the inverse of the second fitness weight coefficient to obtain a smooth evaluation component, and adding the error evaluation component and the smooth evaluation component to obtain a fitness evaluation value; Based on the fitness evaluation value, a parent chromosome is selected by a roulette wheel method, an arithmetic crossover operation and a Gaussian mutation operation are performed on the parent chromosome to generate a child chromosome, and the chromosome with the best fitness evaluation value is selected from the parent chromosome and the child chromosome as the next generation population, until the difference between the optimal fitness evaluation values of two adjacent generations is less than a preset optimal fitness evaluation value threshold, thereby obtaining the optimized weight parameter.

5. The method according to claim 4, characterized in that Selecting a parent chromosome by a roulette wheel method based on the fitness evaluation value, and performing an arithmetic crossover operation and a Gaussian mutation operation on the parent chromosome to generate a child chromosome includes: Calculate the selection probability of each chromosome according to the fitness evaluation value, and accumulate the selection probabilities according to the chromosome order to obtain a cumulative probability sequence; Generate a random number and compare it with the cumulative probability sequence, select the chromosome whose cumulative probability is first greater than the random number as the parent chromosome, and repeat the selection process until two parent chromosomes are selected; Calculating a crossover coefficient based on a ratio of a current evolutionary generation to a maximum evolutionary generation, adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a first daughter chromosome, and adding one to the product of the crossover coefficient and the first parent chromosome minus the product of the crossover coefficient and the second parent chromosome to obtain a second daughter chromosome; The variable step length is calculated based on the ratio of the current evolutionary generation to the maximum evolutionary generation, the variable step length is multiplied by a random number that obeys the standard normal distribution to obtain the variation, and the variation is added to the first daughter chromosome and the second daughter chromosome respectively to obtain the mutated daughter chromosome.

6. The method according to claim 1, characterized in that Using the deep reinforcement learning method, the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link are used as state inputs, and the degree of communication quality improvement and energy consumption are used as reward functions. The optimization results obtained through iterative optimization include: Constructing a state vector based on link state parameters, resource state matrix, and historical reconstruction decision data; The local performance bonus value of a single link is obtained by subtracting the ratio of the current energy consumption to the maximum energy consumption from the ratio of the communication quality change at adjacent moments to the maximum communication quality. Multiplying the network connectivity index by the first global weight coefficient to obtain a connectivity evaluation component, multiplying the network throughput index by the second global weight coefficient to obtain a throughput evaluation component, and adding the connectivity evaluation component and the throughput evaluation component to obtain a network-level global reward value; Dynamically adjusting a dynamic weight coefficient during the training process, and performing a weighted combination of the local performance reward value and the network-level global reward value based on the dynamic weight coefficient to obtain a mixed reward value; A state-action value function is established using a deep neural network, the state vector is input into the deep neural network, the discounted sum of the mixed reward value and the maximum state-action value predicted by the target network is used as a temporal difference target, and the mean square error between the temporal difference target and the current state-action value is calculated to obtain a loss function value; Calculating the gradient of the network parameters based on the loss function value, updating the network parameters of the deep neural network by multiplying the gradient by the learning rate, and simultaneously updating the network parameters of the target network using the soft update coefficient; Repeat the above steps until the difference between two adjacent state action values is less than a preset difference threshold, and obtain the optimized link reconstruction strategy.

7. A system for adaptively reconfiguring laser communication links in a giant constellation network, configured to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to collect and use the real-time position information and motion trajectory information of each satellite node in the giant constellation network to calculate the communication link parameters between each satellite node; The second unit is configured to calculate the communication load and channel utilization of each satellite node based on the communication link parameters, monitor the energy consumption status of each satellite node, and generate a resource status matrix of the satellite node; The third unit is configured to perform interference analysis in parallel in the edge computing unit of each satellite node, determine the interference level of each laser communication link by comparing the actual interference value of each laser communication link with a preset inter-satellite interference threshold, and generate a laser communication link reconstruction priority list based on the interference level, the resource status matrix, and the network communication demand; The fourth unit is used to select a target laser communication link that needs to be reconstructed according to the reconstruction priority list, use a deep reinforcement learning method, take the link state parameters, resource state matrix and historical reconstruction decision data of the target laser communication link as state input, use the degree of improvement in communication quality and energy consumption as reward functions, obtain an optimization result through iterative optimization, and perform adaptive reconstruction on the target laser communication link according to the optimization result.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Resource scheduling optimization method for giant constellation

    CN121239592A

  • Low-orbit satellite cooperative routing and resource scheduling method and system

    CN122533639A