Spectrum decision method and device based on double graph structure and reconstruction cost, and equipment

By employing a spectrum decision-making method based on a dual-graph structure and reconstruction cost, and utilizing a reinforcement learning architecture to optimize spectrum decision-making for low-altitude intelligent networks, this approach addresses the issues of slow response speed and resource waste in existing technologies. It achieves efficient and stable spectrum management, adapting to the intelligent and large-scale needs of low-altitude intelligent networks.

CN122269290APending Publication Date: 2026-06-23湖南工商大学
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
湖南工商大学
Filing Date
2026-05-26
Publication Date
2026-06-23

Smart Images

  • Figure CN122269290A_ABST
    Figure CN122269290A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network resource management and the technical field of reinforcement learning, and discloses a spectrum decision method, device and equipment based on a double-graph structure and reconstruction cost, which comprises the following steps: inputting original state data of low-altitude intelligent networking into an actor network of a reinforcement learning architecture, generating a predicted spectrum decision scheme of the low-altitude intelligent networking through the actor network, and obtaining an expected reward value output by a critic network based on the original state data and the predicted spectrum decision scheme; determining a trained actor network based on the expected reward value and an actual reward value; and generating a target spectrum decision scheme of the low-altitude intelligent networking based on a current double-graph structure and the trained actor network. The application can improve the generation efficiency of the target spectrum decision scheme of the low-altitude intelligent networking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of network resource management technology and reinforcement learning technology, and in particular to spectrum decision-making methods, apparatus and devices based on dual-graph structure and reconstruction cost. Background Technology

[0002] The low-altitude environment faced by the Low-Altitude Intelligent Network is highly dynamic and complex. The real-time changes in aircraft positions and the dense access of airspace users make it easy for problems such as imbalance between spectrum resource supply and demand, mutual signal interference, and unstable channel quality to occur. Therefore, it is necessary to make efficient, real-time, and reasonable decisions and dynamically schedule spectrum resources.

[0003] However, existing low-altitude intelligent networks (LAHICs) rely heavily on manual planning and experience-based configuration for spectrum decisions. This approach is not only slow and time-consuming, but also ill-suited to the real-time changes in low-altitude environments. Furthermore, it is prone to scheduling delays, resource waste, and interference control failures, failing to meet the practical needs of large-scale, intelligent operation of LHAICs. Therefore, generating target spectrum decision-making schemes for LHAICs is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a spectrum decision-making method, apparatus, and device based on a dual-graph structure and reconstruction cost to solve the aforementioned technical problem of how to generate a target spectrum decision-making scheme for low-altitude intelligent networks.

[0005] In a first aspect, embodiments of this application provide a spectrum decision-making method based on a dual-graph structure and reconstruction cost, applied to electronic devices, the spectrum decision-making method comprising: Obtain the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network. The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the low-altitude intelligent network. The raw state data and the prediction spectrum decision scheme are input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. Based on a predefined reconstruction cost generation model, a predefined performance benefit generation model, and a predefined reward value generation model, the actual reward value of the low-altitude intelligent network is determined. Based on the expected reward value and the actual reward value, the trained actor network is determined. Based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for the low-altitude intelligent network is generated.

[0006] In one possible implementation of the first aspect, the step of concatenating the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude intelligent network includes: Obtain the location information, resource information, and historical bandwidth information of the original candidate neighbor graph, and combine the location information, resource information, and historical bandwidth information of the original candidate neighbor graph to form the feature information of the original candidate neighbor graph. Obtain the location information, resource information, and historical bandwidth information of the original link conflict graph, and combine the location information, resource information, and historical bandwidth information of the original link conflict graph to form the feature information of the original link conflict graph. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network.

[0007] In one possible implementation of the first aspect, determining the actual reward value of the low-altitude intelligent network based on a predefined reconstruction cost generation model, a predefined performance gain generation model, and a predefined reward value generation model includes: The bandwidth value of the link between each drone and each access base station in the low-altitude intelligent network is generated by using a predefined bandwidth generation model. Based on the bandwidth value of the link between each UAV and each access base station of the low-altitude intelligent network and the predefined reconstruction cost generation model, the reconstruction cost of the low-altitude intelligent network is generated. Based on the performance benefit generation model, generate the performance benefits of the low-altitude intelligent network. Based on the performance benefits of the Low-Altitude Intelligent Network, the reconstruction cost of the Low-Altitude Intelligent Network, and the predefined reward value generation model, the actual reward value of the Low-Altitude Intelligent Network is generated.

[0008] In one possible implementation of the first aspect, determining the trained actor network based on the expected reward value and the actual reward value, and generating a target spectrum decision scheme for the low-altitude intelligent network based on the current dual-graph structure and the trained actor network, includes: The difference between the expected reward value and the actual reward value is obtained, and the training objective of the critic network is to reduce the difference. When the difference is less than the preset value, the training of the critic network is stopped, and the trained critic network is obtained. The state value is output through the trained critic network, and the gradient update training of the actor network is performed with the goal of maximizing the state value, thus generating the trained actor network. Obtain the current dual-graph structure of the Low-Altitude Intelligent Network, which includes the current candidate neighbor graph and the current link conflict graph of the Low-Altitude Intelligent Network; The system acquires the location, resource, and historical bandwidth information of the current candidate neighbor graph, and combines these information to form the feature information of the current candidate neighbor graph. It also acquires the location, resource, and historical bandwidth information of the current link conflict graph, and combines these information to form the feature information of the current link conflict graph. The system then concatenates the feature information of the current candidate neighbor graph and the feature information of the current link conflict graph to generate the current state data of the Low-Altitude Intelligent Network. This current state data is then input into the trained actor network, which generates the target spectrum decision scheme for the Low-Altitude Intelligent Network.

[0009] In one possible implementation of the first aspect, the raw state data of the Low-Altitude Intelligent Network (LAI) is input into the actor network of a reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the LAI. The raw state data and the prediction spectrum decision scheme are then input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme, including: The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture to obtain the preference information output by the actor network based on the raw state data. The preference information includes the preference score of each UAV in the low-altitude intelligent network for each access base station. By using the preference scores of each UAV to each access base station in the low-altitude intelligent network and the predefined bandwidth generation model, the bandwidth value of the link between each UAV and each access base station in the low-altitude intelligent network is generated. The bandwidth value of the link between each UAV and each access base station is used to form a predictive spectrum decision scheme. The original state data and the predictive spectrum decision scheme are input into the critic network to obtain the expected reward value output by the critic network based on the original state data and the predictive spectrum decision scheme. The bandwidth generation model is defined as follows: ; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the total bandwidth budget for the i-th drone in the low-altitude intelligent network. Let represent the preference score of the i-th UAV in the low-altitude intelligent network for the b-th access base station at time t; This indicates that at time t, the i-th drone in the low-altitude intelligent network and the i-th drone... The bandwidth value of the link between candidate base stations; Let represent the set of candidate base stations for the i-th UAV in the low-altitude intelligent network at time t; This represents the temperature coefficient in the Softmax operation.

[0010] In one possible implementation of the first aspect, the reconstruction cost generation model is defined as follows: ; This represents the cost of reconstructing the low-altitude intelligent network at time t; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the time preceding time t. This represents the set of links in the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at the time preceding time t. This indicates the number of links in the low-altitude intelligent network. express and The union of .

[0011] In one possible implementation of the first aspect, The performance gain generation model is defined as follows: ; This represents the performance gains of the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at time t; This represents the amount of bandwidth allocated to the i-th link in the link set at time t; This represents the signal-to-noise-interference ratio (SNR) at the receiver of the i-th link in the link set at time t.

[0012] In one possible implementation of the first aspect, the reward value generation model is defined as follows: ; This represents the actual reward value of the low-altitude intelligent network at time t; This represents the performance gains of the low-altitude intelligent network at time t; This represents the cost of reconstructing the low-altitude intelligent network at time t; Indicates the trade-off coefficient. The tradeoff coefficient is used to adjust the strength of the tradeoff between performance gains and reconstruction costs.

[0013] Secondly, embodiments of this application provide a spectrum decision-making device based on a dual-graph structure and reconstruction cost, applied to electronic devices, including: The acquisition module is used to acquire the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The component module is used to splice the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude intelligent network. The input module is used to input the raw state data of the low-altitude intelligent network into the actor network of the reinforcement learning architecture, generate the prediction spectrum decision scheme of the low-altitude intelligent network through the actor network, input the raw state data and the prediction spectrum decision scheme into the critic network of the reinforcement learning architecture, and obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. The determination module is used to determine the actual reward value of the low-altitude intelligent network based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model. The generation module is used to determine the trained actor network based on the expected reward value and the actual reward value, and to generate a target spectrum decision scheme for the low-altitude intelligent network based on the current dual-graph structure and the trained actor network.

[0014] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the spectrum decision method described in the first aspect above.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the spectrum decision method described in the first aspect above.

[0016] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the spectrum decision method described in the first aspect.

[0017] The beneficial effects of the embodiments of this application are as follows: Firstly, since the actual reward value of the low-altitude intelligent network is determined based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model, the trained actor network is determined based on the expected reward value and the actual reward value, and the target spectrum decision scheme of the low-altitude intelligent network is generated based on the current dual-graph structure and the trained actor network, the generation time of the target spectrum decision scheme of the low-altitude intelligent network is reduced, which is conducive to improving the generation efficiency of the target spectrum decision scheme of the low-altitude intelligent network. Secondly, based on the current dual-graph structure and the trained actor network, the target spectrum decision scheme of the low-altitude intelligent network is automatically generated, which not only greatly improves the generation efficiency of the target spectrum decision scheme, but also reduces resource waste and signal conflicts, improves the stability and reliability of the low-altitude intelligent network operation, and better meets the actual needs of low-altitude logistics, emergency rescue, urban air traffic and other scenarios for efficient, safe and intelligent management. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a diagram illustrating the application scenario of the spectrum decision-making method provided in the embodiments of this application. Figure 2 This is a flowchart illustrating the spectrum decision-making method provided in an embodiment of this application; Figure 3 A flowchart illustrating the implementation of S202 provided in this application embodiment; Figure 4 A schematic block diagram of a spectrum decision-making device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0021] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0022] It should be understood that in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0023] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0024] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0025] The spectrum decision-making method provided in this application can be applied to electronic devices, including but not limited to servers, mobile phones, tablets, wearable devices, vehicle-mounted devices, and laptops. This application does not impose any restrictions on the specific type of electronic device.

[0026] Please see Figure 1 , Figure 1 The application scenario diagram of the spectrum decision method provided in the embodiments of this application is described in detail below: Electronic devices access the database to obtain the original dual-graph structure of the low-altitude intelligent network. The original dual-graph structure includes the original candidate neighbor graph and the original link conflict graph of the low-altitude intelligent network.

[0027] In this embodiment, the electronic device accesses the database to obtain the original dual-map structure of the low-altitude intelligent network, which can reduce the acquisition time of the original dual-map structure and improve the acquisition efficiency of the original dual-map structure.

[0028] Please see Figure 2 , Figure 2 This is a flowchart illustrating the spectrum decision-making method provided in an embodiment of this application, which can be applied to electronic devices.

[0029] As Figure 2 shown, the spectrum decision-making method provided by the embodiment of this application includes the following steps, which are described in detail as follows: S201. Obtain the original dual-graph structure of the low-altitude aerial intelligent network. The original dual-graph structure includes the original candidate neighbor graph of the low-altitude aerial intelligent network and the original link conflict graph of the low-altitude aerial intelligent network; Among them, the low-altitude aerial intelligent network (Low-Altitude Aerial Intelligent Network, LAIN) is a new type of intelligent information infrastructure facing the low-altitude airspace of 1000 - 3000 meters, integrating the coordinated operation of five networks including low-altitude communication network, sensing network, navigation network, meteorological network, and computing power network, and integrating global sensing, high-reliability communication, real-time computing, and intelligent management and control.

[0030] S202. Concatenate the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude aerial intelligent network; Among them, the original candidate neighbor graph is used to represent the candidate connection relationship between the unmanned aerial vehicle and the access base station, and is used to define the actionable action space of the Actor network. The original link conflict graph is used to represent the co-frequency interference coupling relationship between the candidate links, and is used to provide a structured value evaluation basis for the Critic value network.

[0031] Among them, the original candidate neighbor graph and the original link conflict graph respectively depict the underlying network structure from the perspectives of neighbor connectivity relationship and link conflict interference. A strict entity correspondence association is established between the two graphs, binding the topological relationship and conflict constraints to each other and联动配合 (cooperating in a coordinated manner), so as to jointly support the dual-graph structured collaborative modeling.

[0032] S203. Input the original state data of the low-altitude aerial intelligent network into the actor network of the reinforcement learning architecture. Through the actor network, generate the predicted spectrum decision-making scheme of the low-altitude aerial intelligent network. Input the original state data and the predicted spectrum decision-making scheme into the critic network of the reinforcement learning architecture, and obtain the expected reward value output by the critic network based on the original state data and the predicted spectrum decision-making scheme; Among them, the Chinese name of the actor network is: Policy Network.

[0033] Among them, the Chinese name of the critic network is: Evaluation Network.

[0034] Specifically, the raw state data of the Low-Altitude Intelligent Network (LAI) is input into the actor network of a reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the LAI. The raw state data and the prediction spectrum decision scheme are then input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. This reward value includes: The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture to obtain the preference information output by the actor network based on the raw state data. The preference information includes the preference score of each UAV in the low-altitude intelligent network for each access base station. By using the preference scores of each UAV to each access base station in the low-altitude intelligent network and the predefined bandwidth generation model, the bandwidth value of the link between each UAV and each access base station in the low-altitude intelligent network is generated. The bandwidth value of the link between each UAV and each access base station is used to form a predictive spectrum decision scheme. The original state data and the predictive spectrum decision scheme are input into the critic network to obtain the expected reward value output by the critic network based on the original state data and the predictive spectrum decision scheme. The bandwidth generation model is defined as follows: ; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the total bandwidth budget for the i-th drone in the low-altitude intelligent network. Let represent the preference score of the i-th UAV in the low-altitude intelligent network for the b-th access base station at time t; This indicates that at time t, the i-th drone in the low-altitude intelligent network and the i-th drone... The bandwidth value of the link between candidate base stations; Let represent the set of candidate base stations for the i-th UAV in the low-altitude intelligent network at time t; This represents the temperature coefficient in the Softmax operation.

[0035] S204. Based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model, the actual reward value of the low-altitude intelligent network is determined. The determination of the actual reward value of the low-altitude intelligent network based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model includes: Based on the bandwidth value of the link between each UAV and each access base station of the low-altitude intelligent network and the predefined reconstruction cost generation model, the reconstruction cost of the low-altitude intelligent network is generated. Based on the performance benefit generation model, generate the performance benefits of the low-altitude intelligent network. Based on the performance benefits of the Low-Altitude Intelligent Network, the reconstruction cost of the Low-Altitude Intelligent Network, and the predefined reward value generation model, the actual reward value of the Low-Altitude Intelligent Network is generated.

[0036] Using a predefined bandwidth generation model, the bandwidth values ​​of the link between each drone and each access base station in the low-altitude intelligent network are generated, including: The reconstruction cost generation model is defined as follows: ; This represents the cost of reconstructing the low-altitude intelligent network at time t; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the time preceding time t. This represents the set of links in the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at the time preceding time t. This indicates the number of links in the low-altitude intelligent network. express and The union of .

[0037] The performance benefit generation model is defined as follows: ; This represents the performance gains of the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at time t; This represents the amount of bandwidth allocated to the i-th link in the link set at time t; This represents the signal-to-noise-interference ratio (SNR) at the receiver of the i-th link in the link set at time t.

[0038] The reward value generation model is defined as follows: ; This represents the actual reward value of the low-altitude intelligent network at time t; This represents the performance gains of the low-altitude intelligent network at time t; This represents the cost of reconstructing the low-altitude intelligent network at time t; Indicates the trade-off coefficient. The tradeoff coefficient is used to adjust the strength of the tradeoff between performance gains and reconstruction costs.

[0039] S205. Based on the expected reward value and the actual reward value, determine the trained actor network, and based on the current dual-graph structure and the trained actor network, generate a target spectrum decision scheme for the low-altitude intelligent network.

[0040] In this process, the current state data of the low-altitude intelligent network is input into the trained actor network, and the target spectrum decision scheme of the low-altitude intelligent network is generated through the trained actor network. This is not affected by human intervention and helps to improve the reliability of the target spectrum decision scheme of the low-altitude intelligent network.

[0041] Specifically, based on the expected and actual reward values, the trained actor network is determined. Based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for the Low-Altitude Intelligent Network is generated, including: The difference between the expected reward value and the actual reward value is obtained, and the training objective of the critic network is to reduce the difference. When the difference is less than the preset value, the training of the critic network is stopped, and the trained critic network is obtained. The state value is output through the trained critic network, and the gradient update training of the actor network is performed with the goal of maximizing the state value, thus generating the trained actor network. Obtain the current dual-graph structure of the Low-Altitude Intelligent Network, which includes the current candidate neighbor graph and the current link conflict graph of the Low-Altitude Intelligent Network; The system acquires the location, resource, and historical bandwidth information of the current candidate neighbor graph, and combines these information to form the feature information of the current candidate neighbor graph. It also acquires the location, resource, and historical bandwidth information of the current link conflict graph, and combines these information to form the feature information of the current link conflict graph. The system then concatenates the feature information of the current candidate neighbor graph and the feature information of the current link conflict graph to generate the current state data of the Low-Altitude Intelligent Network. This current state data is then input into the trained actor network, which generates the target spectrum decision scheme for the Low-Altitude Intelligent Network.

[0042] Among them, the current candidate neighbor graph is used to depict the dynamic candidate connection relationship between the drone and the base station, and to limit the current action space; The current link conflict graph is a topology diagram with links as nodes, used to explicitly depict the co-channel interference coupling relationships between links. In the current link conflict graph, each link corresponds to one node. If two links have non-negligible interference coupling due to using the same frequency, a conflict edge is established between the corresponding nodes. The weight of the conflict edge is determined by the frequency domain overlap ratio and the power gain of the interfering channel, used to explicitly depict the interference coupling strength. This graphical modeling method clearly expresses the degree of conflict and interference range between links. Based on the current link conflict graph, link value assessment can be further carried out. During the assessment process, the interference coupling strength of the links is comprehensively considered to support spectrum allocation decisions.

[0043] The location information of the current link conflict graph refers to the actual spatial coordinate data of all nodes in the link conflict graph at the current moment. It clarifies the specific location of each node and is used to determine the feasibility of communication between nodes and avoid link conflicts. It is the basic data to ensure the stability of link connection and can accurately reflect the spatial distribution of each node, providing a location basis for subsequent link scheduling.

[0044] The resource information of the current link conflict graph refers to the resource status of each node and corresponding link in the link conflict graph at the current moment. The core includes key resource parameters such as node computing power, communication bandwidth, and energy reserves. It clarifies the available resource capabilities of each node, provides data support for the rational allocation of resources and avoidance of resource waste, and ensures the efficient operation of the link.

[0045] The historical bandwidth information of the current link conflict diagram refers to the bandwidth usage records of each communication link in the link conflict diagram over a period of time, including bandwidth usage peaks, fluctuation patterns, interference, and usage stability. It can be used to predict current bandwidth demand, optimize bandwidth allocation strategies, reduce link congestion, and ensure smooth communication transmission.

[0046] The location information of the current candidate neighbor graph refers to the actual spatial coordinates of each node in the candidate neighbor graph at the current moment. It clarifies the real-time distribution of each node, is used to determine the communication reachability between nodes, avoid link connection failures caused by location deviations, and provide a basic location basis for connection planning between nodes.

[0047] The resource information of the current candidate neighbor graph refers to the resource capabilities that each node in the candidate neighbor graph can provide at the current moment, including computing power, communication bandwidth, energy reserves, etc. It clarifies the resource redundancy of each node, provides data support for resource scheduling and avoidance of resource waste, and adapts to the dynamically changing network environment.

[0048] The historical bandwidth information of the current candidate neighbor graph refers to the bandwidth usage records of each link in the candidate neighbor graph over a period of time, including the average bandwidth usage, fluctuation range, and abnormal losses. It can be used to analyze bandwidth usage patterns, predict future bandwidth demand, and provide a reference for link bandwidth optimization and conflict avoidance.

[0049] Among them, the target spectrum decision scheme of low-altitude intelligent network is a systematic decision scheme that takes into account the spectrum utilization benefits and implementation costs. Its core is to achieve a balance between performance benefits and reconfiguration costs while ensuring communication quality and meeting application needs, so as to give full play to the value of spectrum resources and control implementation costs.

[0050] Among them, based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for low-altitude intelligent network is generated. It can perceive the bandwidth requirements, communication priorities and link interference of each terminal in real time, and dynamically allocate available spectrum. This can not only avoid the idleness and waste of spectrum resources, but also improve the operational stability and reliability of low-altitude intelligent network.

[0051] For ease of explanation, the following example is provided: For example, when drones perform emergency rescue missions, the target spectrum decision-making scheme of the low-altitude intelligent network can allocate stable bandwidth to the drones to ensure efficient transmission of rescue commands. At the same time, it can rationally schedule the spectrum usage of other terminals, ensuring the communication needs of critical missions while maximizing the use of limited spectrum resources and solving the communication lag and interruption problems caused by unreasonable spectrum allocation.

[0052] For example, in large-scale drone swarm operation scenarios, the target spectrum decision-making scheme of the low-altitude intelligent network can adapt to the flight trajectory and communication needs of each drone in real time, automatically adjust the spectrum configuration, avoid signal conflicts among drones in the swarm, ensure the stability of swarm collaborative operation, promote the upgrading of the low-altitude intelligent network to intelligence and scale, and meet the needs of high-quality development of the low-altitude industry.

[0053] The beneficial effects of the embodiments of this application are as follows: Firstly, since the actual reward value of the low-altitude intelligent network is determined based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model, the trained actor network is determined based on the expected reward value and the actual reward value, and the target spectrum decision scheme of the low-altitude intelligent network is generated based on the current dual-graph structure and the trained actor network, the generation time of the target spectrum decision scheme of the low-altitude intelligent network is reduced, which is conducive to improving the generation efficiency of the target spectrum decision scheme of the low-altitude intelligent network. Secondly, based on the current dual-graph structure and the trained actor network, the target spectrum decision scheme of the low-altitude intelligent network is automatically generated, which not only greatly improves the generation efficiency of the target spectrum decision scheme, but also reduces resource waste and signal conflicts, improves the stability and reliability of the low-altitude intelligent network operation, and better meets the actual needs of low-altitude logistics, emergency rescue, urban air traffic and other scenarios for efficient, safe and intelligent management.

[0054] Please see Figure 3 , Figure 3 The implementation flowchart of S202 provided in the embodiments of this application is described in detail below: S301, obtain the location information, resource information, and historical bandwidth information of the original candidate neighbor graph, and combine the location information, resource information, and historical bandwidth information of the original candidate neighbor graph to form the feature information of the original candidate neighbor graph; obtain the location information, resource information, and historical bandwidth information of the original link conflict graph, and combine the location information, resource information, and historical bandwidth information of the original link conflict graph to form the feature information of the original link conflict graph. The location information of the original link conflict diagram is the theoretical location data of each node in the preset link conflict diagram. It clarifies the initial layout distribution of each node, provides a basic location basis for subsequent resource allocation and conflict avoidance, and reflects the initial state of the original network topology.

[0055] The resource information of the original link conflict diagram is the resource configuration parameters of each node and link in the preset link conflict diagram, including the core resources such as the initially available bandwidth and computing power. It clarifies the resource reserve status in the original state and provides a benchmark reference for subsequent resource scheduling.

[0056] The historical bandwidth information in the original link conflict diagram is based on a preset scenario and pre-set historical reference data for link bandwidth usage, including historical bandwidth averages, fluctuation ranges, and records of abnormal situations, providing a reference for initial decision-making and resource planning.

[0057] The location information of the original candidate neighbor graph is the theoretical location parameters of each node in the preset candidate neighbor graph. It clarifies the initial layout plan of each node, provides a basis for subsequent node connection and topology design, and reflects the topological framework of the original network.

[0058] The resource information of the original candidate neighbor graph is the initial resource configuration of each node in the preset candidate neighbor graph, including the initial available computing power, communication bandwidth, etc., which clarifies the resource supply capacity in the original state and provides an initial basis for subsequent resource optimization and allocation.

[0059] The historical bandwidth information of the original candidate neighbor graph is a preset historical reference data of bandwidth usage corresponding to the original scenario, including the historical bandwidth usage range, fluctuation patterns and preset abnormal situations, which provides data support for the formulation of the original decision-making scheme.

[0060] S302, the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network.

[0061] In this embodiment, the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are concatenated to generate the original state data of the Low-Altitude Intelligent Network (LAI). This data comprehensively reflects the available link resources, node connection relationships, and interference and conflict status between links in the LAI, enabling the actor network to directly learn the correlation patterns and joint representation methods between different features during the training phase. This allows the model to better understand the complex interaction between link resources and conflict constraints in the LAI, thus exhibiting stronger generalization adaptability and robustness when facing unseen network topologies, spectrum environments, or flight scenarios, effectively improving the overall performance and reliability of the decision-making model.

[0062] For the spectrum decision method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of a spectrum decision-making device provided in an embodiment of this application. Figure 4 The spectrum decision device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The spectrum decision-making device 400 shown will be described in detail. The spectrum decision-making device 400 may include an acquisition module 401, a composition module 402, an input module 403, a determination module 404, and a generation module 405.

[0063] The acquisition module 401 is used to acquire the original dual-graph structure of the low-altitude intelligent network, which includes the original candidate neighbor graph of the low-altitude intelligent network and the original link conflict graph of the low-altitude intelligent network. Module 402 is used to splice the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude intelligent network. The input module 403 is used to input the raw state data of the low-altitude intelligent network into the actor network of the reinforcement learning architecture, generate the prediction spectrum decision scheme of the low-altitude intelligent network through the actor network, input the raw state data and the prediction spectrum decision scheme into the critic network of the reinforcement learning architecture, and obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. Module 404 is used to determine the actual reward value of the low-altitude intelligent network based on a predefined reconstruction cost generation model, a predefined performance benefit generation model, and a predefined reward value generation model. The generation module 405 is used to determine the trained actor network based on the expected reward value and the actual reward value, and to generate a target spectrum decision scheme for the low-altitude intelligent network based on the current dual-graph structure and the trained actor network.

[0064] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0065] The beneficial effects of the embodiments of this application are as follows: Firstly, since the actual reward value of the low-altitude intelligent network is determined based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model, the trained actor network is determined based on the expected reward value and the actual reward value, and the target spectrum decision scheme of the low-altitude intelligent network is generated based on the current dual-graph structure and the trained actor network, the generation time of the target spectrum decision scheme of the low-altitude intelligent network is reduced, which is conducive to improving the generation efficiency of the target spectrum decision scheme of the low-altitude intelligent network. Secondly, based on the current dual-graph structure and the trained actor network, the target spectrum decision scheme of the low-altitude intelligent network is automatically generated, which not only greatly improves the generation efficiency of the target spectrum decision scheme, but also reduces resource waste and signal conflicts, improves the stability and reliability of the low-altitude intelligent network operation, and better meets the actual needs of low-altitude logistics, emergency rescue, urban air traffic and other scenarios for efficient, safe and intelligent management.

[0066] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0067] like Figure 5 As shown, Figure 5 The electronic device includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0068] The electronic device may include, but is not limited to, processor 20 and memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of an electronic device and does not constitute a limitation on electronic devices. It may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0069] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22: Obtain the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network. The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the low-altitude intelligent network. The raw state data and the prediction spectrum decision scheme are input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. Based on a predefined reconstruction cost generation model, a predefined performance benefit generation model, and a predefined reward value generation model, the actual reward value of the low-altitude intelligent network is determined. Based on the expected reward value and the actual reward value, the trained actor network is determined. Based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for the low-altitude intelligent network is generated.

[0070] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0071] In some embodiments, the memory 21 may be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory 21 may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device.

[0072] Furthermore, the memory 21 may include both internal storage units and external storage devices of the electronic device. The memory 21 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0073] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0074] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0075] The computer-readable storage medium may also be an external storage device of the spectrum decision-making device or electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, or non-transitory computer-readable storage medium equipped on the spectrum decision-making device or electronic device.

[0076] Since the computer program stored in the computer-readable storage medium can execute any of the spectrum decision-making methods based on dual-graph structure and reconstruction cost provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the spectrum decision-making methods based on dual-graph structure and reconstruction cost provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0077] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the aforementioned spectrum decision method.

[0078] When a computer program is loaded into an electronic device, it can perform the following steps: Obtain the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network. The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the low-altitude intelligent network. The raw state data and the prediction spectrum decision scheme are input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. Based on a predefined reconstruction cost generation model, a predefined performance benefit generation model, and a predefined reward value generation model, the actual reward value of the low-altitude intelligent network is determined. Based on the expected reward value and the actual reward value, the trained actor network is determined. Based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for the low-altitude intelligent network is generated.

[0079] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0080] Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium includes: an entity or device for carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium.

[0081] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0082] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A spectrum decision-making method based on dual-graph structure and reconstruction cost, characterized in that, The spectrum decision method, applied to electronic devices, includes: Obtain the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network. The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the low-altitude intelligent network. The raw state data and the prediction spectrum decision scheme are input into the critic network of the reinforcement learning architecture to obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. Based on a predefined reconstruction cost generation model, a predefined performance benefit generation model, and a predefined reward value generation model, the actual reward value of the low-altitude intelligent network is determined. Based on the expected reward value and the actual reward value, the trained actor network is determined. Based on the current dual-graph structure and the trained actor network, a target spectrum decision scheme for the low-altitude intelligent network is generated.

2. The spectrum decision method according to claim 1, characterized in that, The step of concatenating the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude intelligent network includes: Obtain the location information, resource information, and historical bandwidth information of the original candidate neighbor graph, and combine the location information, resource information, and historical bandwidth information of the original candidate neighbor graph to form the feature information of the original candidate neighbor graph. Obtain the location information, resource information, and historical bandwidth information of the original link conflict graph, and combine the location information, resource information, and historical bandwidth information of the original link conflict graph to form the feature information of the original link conflict graph. The feature information of the original candidate neighbor graph and the feature information of the original link conflict graph are spliced ​​together to generate the original state data of the low-altitude intelligent network.

3. The spectrum decision method according to claim 1, characterized in that, The actual reward value of the low-altitude intelligent network is determined based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model, including: The bandwidth value of the link between each drone and each access base station in the low-altitude intelligent network is generated by using a predefined bandwidth generation model. Based on the bandwidth value of the link between each UAV and each access base station of the low-altitude intelligent network and the predefined reconstruction cost generation model, the reconstruction cost of the low-altitude intelligent network is generated. Based on the performance benefit generation model, generate the performance benefits of the low-altitude intelligent network. Based on the performance benefits of the Low-Altitude Intelligent Network, the reconstruction cost of the Low-Altitude Intelligent Network, and the predefined reward value generation model, the actual reward value of the Low-Altitude Intelligent Network is generated.

4. The spectrum decision method according to claim 1, characterized in that, The process involves determining the trained actor network based on the expected and actual reward values, and generating a target spectrum decision scheme for the low-altitude intelligent network based on the current dual-graph structure and the trained actor network, including: The difference between the expected reward value and the actual reward value is obtained, and the training objective of the critic network is to reduce the difference. When the difference is less than the preset value, the training of the critic network is stopped, and the trained critic network is obtained. The state value is output through the trained critic network, and the gradient update training of the actor network is performed with the goal of maximizing the state value, thus generating the trained actor network. Obtain the current dual-graph structure of the Low-Altitude Intelligent Network, which includes the current candidate neighbor graph and the current link conflict graph of the Low-Altitude Intelligent Network; The system acquires the location, resource, and historical bandwidth information of the current candidate neighbor graph, and combines these information to form the feature information of the current candidate neighbor graph. It also acquires the location, resource, and historical bandwidth information of the current link conflict graph, and combines these information to form the feature information of the current link conflict graph. The system then concatenates the feature information of the current candidate neighbor graph and the feature information of the current link conflict graph to generate the current state data of the Low-Altitude Intelligent Network. This current state data is then input into the trained actor network, which generates the target spectrum decision scheme for the Low-Altitude Intelligent Network.

5. The spectrum decision method according to claim 1, characterized in that, The raw state data of the Low-Altitude Intelligent Network (LAI) is input into the actor network of a reinforcement learning architecture. The actor network generates a prediction spectrum decision scheme for the LAI. The raw state data and the prediction spectrum decision scheme are then input into the critic network of the reinforcement learning architecture. The expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme is obtained, including: The raw state data of the low-altitude intelligent network is input into the actor network of the reinforcement learning architecture to obtain the preference information output by the actor network based on the raw state data. The preference information includes the preference score of each UAV in the low-altitude intelligent network for each access base station. By using the preference scores of each UAV to each access base station in the low-altitude intelligent network and the predefined bandwidth generation model, the bandwidth value of the link between each UAV and each access base station in the low-altitude intelligent network is generated. The bandwidth value of the link between each UAV and each access base station is used to form a predictive spectrum decision scheme. The original state data and the predictive spectrum decision scheme are input into the critic network to obtain the expected reward value output by the critic network based on the original state data and the predictive spectrum decision scheme. The bandwidth generation model is defined as follows: ; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the total bandwidth budget for the i-th drone in the low-altitude intelligent network. Let represent the preference score of the i-th UAV in the low-altitude intelligent network for the b-th access base station at time t; This indicates that at time t, the i-th drone in the low-altitude intelligent network and the i-th drone... The bandwidth value of the link between candidate base stations; Let represent the set of candidate base stations for the i-th UAV in the low-altitude intelligent network at time t; This represents the temperature coefficient in the Softmax operation.

6. The spectrum decision method according to claim 1, characterized in that, The reconstruction cost generation model is defined as follows: ; This represents the cost of reconstructing the low-altitude intelligent network at time t; This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the low-altitude intelligent network at time t. This represents the bandwidth value of the link between the i-th UAV and the b-th access base station in the time preceding time t. This represents the set of links in the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at the time preceding time t. This indicates the number of links in the low-altitude intelligent network. express and The union of .

7. The spectrum decision method according to claim 1, characterized in that, The performance gain generation model is defined as follows: ; This represents the performance gains of the low-altitude intelligent network at time t; This represents the set of links in the low-altitude intelligent network at time t; This represents the amount of bandwidth allocated to the i-th link in the link set at time t; This represents the signal-to-noise-interference ratio (SNR) at the receiver of the i-th link in the link set at time t.

8. The spectrum decision method according to claim 1, characterized in that, The reward value generation model is defined as follows: ; This represents the actual reward value of the low-altitude intelligent network at time t; This represents the performance gains of the low-altitude intelligent network at time t; This represents the cost of reconstructing the low-altitude intelligent network at time t; Indicates the trade-off coefficient. The tradeoff coefficient is used to adjust the strength of the tradeoff between performance gains and reconstruction costs.

9. A spectrum decision-making device based on a dual-graph structure and reconstruction cost, characterized in that, Applied to electronic devices, including: The acquisition module is used to acquire the original dual-graph structure of the Low Altitude Intelligent Network, which includes the original candidate neighbor graph and the original link conflict graph of the Low Altitude Intelligent Network. The component module is used to splice the feature information of the original candidate neighbor graph and the feature information of the original link conflict graph to generate the original state data of the low-altitude intelligent network. The input module is used to input the raw state data of the low-altitude intelligent network into the actor network of the reinforcement learning architecture, generate the prediction spectrum decision scheme of the low-altitude intelligent network through the actor network, input the raw state data and the prediction spectrum decision scheme into the critic network of the reinforcement learning architecture, and obtain the expected reward value output by the critic network based on the raw state data and the prediction spectrum decision scheme. The determination module is used to determine the actual reward value of the low-altitude intelligent network based on the predefined reconstruction cost generation model, the predefined performance benefit generation model, and the predefined reward value generation model. The generation module is used to determine the trained actor network based on the expected reward value and the actual reward value, and to generate a target spectrum decision scheme for the low-altitude intelligent network based on the current dual-graph structure and the trained actor network.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the spectrum decision method as described in any one of claims 1 to 8.