User requirement analysis-based on-demand service policy system for space-air-ground scenarios

By designing a collaborative architecture between the user end and the network end, and combining multi-level data processing and dynamic link selection algorithms, the problem of low resource allocation efficiency in the integrated air-space-ground network is solved, achieving efficient resource scheduling and user demand response, and improving the service quality and efficiency of the system.

WO2026102802A1PCT designated stage Publication Date: 2026-05-21XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2024-11-22
Publication Date
2026-05-21

Smart Images

  • Figure CN2024133840_21052026_PF_FP_ABST
    Figure CN2024133840_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a user demand analysis-based on-demand service policy system for space-air-ground scenarios, which solves the problem in the prior art of lacking in-depth analysis of personalized user demands and being unable to achieve accurate resource allocation on the basis of dynamically changing demands. The solution comprises: a user end comprises a data acquisition module, a data processing module and a demand upload module which are successively connected; the user end is used for data acquisition, data preprocessing and demand uploading; a network end comprises a network state acquisition module, a demand analysis module, a policy calculation module and a scheduling module which are successively connected, the network state acquisition module being used for acquiring the current network resource in real time; the network end analyzes demands, and, on the basis of an analysis result, obtains a network resource allocation policy, so as to achieve efficient data transmission and dynamic resource configuration.
Need to check novelty before this filing date? Find Prior Art

Description

A system for on-demand service strategies in air-space-ground scenarios based on user demand analysis Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) technology, and in particular to an on-demand service strategy system for air-space-ground scenarios based on user demand analysis. Background Technology

[0002] With the development of integrated air-space-ground networks, the diversity of Internet of Things (IoT) devices and user needs is increasing. These devices and user needs are distributed across different spatial environments and exhibit different service requirements at different times. Existing resource management mechanisms mostly adopt static resource configuration, which is difficult to adapt to the constantly changing needs and resource status in air-space-ground scenarios in real time.

[0003] With the rapid development of communication networks and information processing technologies, user demands for services are becoming increasingly diversified and personalized. To meet these dynamic needs within limited network resources, on-demand service strategies have become crucial for improving user experience and resource utilization. However, existing architectures lack effective demand resolution mechanisms and dynamic resource allocation algorithms, resulting in low efficiency and accuracy in resource allocation. Summary of the Invention

[0004] This invention provides an on-demand service strategy system for air, space, and ground scenarios based on user demand analysis. This solves the problem in existing technologies that lack in-depth analysis of personalized user needs and cannot achieve accurate resource allocation based on dynamically changing needs, thus realizing efficient data transmission and dynamic resource configuration.

[0005] This invention provides an on-demand service strategy system for air-space-ground scenarios based on user demand analysis. The system includes a user terminal and a network terminal, wherein the user terminal and the network terminal are connected through a data processing module and a demand analysis module.

[0006] The user terminal includes a data acquisition module, a data processing module, and a demand upload module connected in sequence. The data acquisition module is used to acquire environmental data collected in real time by sensors on each user device. The data processing module is used to preprocess the environmental data to obtain preprocessed environmental data. The demand upload module is used to propose business demands based on the preprocessed environmental data to obtain user device demand data.

[0007] The network terminal includes a network status acquisition module, a demand parsing module, a policy calculation module, and a scheduling module connected in sequence. The network status acquisition module is used to acquire current network resources in real time. The demand parsing module performs demand parsing on the user equipment demand data to obtain the parsing results. The policy calculation module is used to calculate a network resource allocation policy based on the parsing results and the preprocessed environmental data. The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation policy. The network resources include: terrestrial network, air-based network, and space-based network.

[0008] In one possible implementation, the data acquisition module, in the step of preprocessing the environmental data to obtain preprocessed environmental data, includes: performing data filtering, temporal redundancy elimination, and spatial redundancy elimination on the environmental data to obtain preprocessed environmental data.

[0009] In one possible implementation, the requirement parsing module, in the process of parsing the user equipment requirement data, includes: classifying the user equipment requirement data and assigning data priorities using a hierarchical architecture.

[0010] In one possible implementation, the policy calculation module, which calculates the network resource allocation policy based on the parsing result and the preprocessed environmental data, includes:

[0011] The first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network are determined respectively.

[0012] The third interference plus noise ratio of the signal between the land-based network and the user terminal is determined, and the third data transmission rate is calculated based on the third interference plus noise ratio.

[0013] Determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio;

[0014] Determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio;

[0015] The actual time consumed by the parsing result is calculated based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate.

[0016] A temporary reward function for each user device is determined based on the actual time consumed, and a long-term reward function for each user device is determined based on the temporary reward function.

[0017] Based on the access volume constraint function of the land-based network and the storage volume constraint function of the land-based network and the air-based network, the first network environment state at time slot t is determined;

[0018] The long-term reward function is optimized based on the first network environment state at time slot t to obtain the optimized reward function;

[0019] Based on the first network environment state at time slot t and the second network environment state at time slot t+1, determine the optimal strategy for the current user equipment and the optimal strategies for the other user equipment.

[0020] The optimization reward function is optimized based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. The final reward function is then solved to obtain the network resource allocation policy for each user equipment.

[0021] In one possible implementation, the temporary reward function is expressed as:

[0022] in, Indicates the maximum acceptable latency when a user device requests a service; T i (t) represents the actual time consumed by the task; θ1 represents the cost weight; θ2 represents the weight of the overall service rate (Rate); C B Indicates the cost of using terrestrial networks; C V This represents the cost of using the space-based network; Rate represents the overall service rate. Indicates; N V express; Indicates; N B express.

[0023] In one possible implementation, the long-term reward function is expressed as:

[0024] Where λ represents the first constant; τ represents the slot exponent of the time step; r i (t+τ+1) represents; t represents the time step; v i (t) represents the long-term reward of the i-th user device in time slot t.

[0025] In one possible implementation, the optimized reward function is expressed as:

[0026] Where Ε[·] represents the expectation operation; si Indicates the current environmental state; v i (s i ,π) represents the state in s i The optimized reward is calculated using action π in the state; λ represents the first constant; τ represents the slot exponent of the time step; t represents the time step; s i (t) represents the state function; r i (t+τ+1) represents the temporary reward in time slot t+τ+1.

[0027] In one possible implementation, the final reward function is expressed as:

[0028] Where Ε[·] represents the expectation operation; λ represents the first constant; r i (·) represents the temporary reward corresponding to the i-th user device; τ represents the slot exponent of the time step; t represents the time step; s i (t+τ) represents; π i Indicates the current user device's policy; π -i Indicates the policies of other user devices; s i (t) represents the state function.

[0029] In one possible implementation, the network terminal further includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.

[0030] One or more technical solutions provided in this invention have at least the following technical effects or advantages:

[0031] (1) This invention adopts a collaborative architecture design of user-end module and network-end module. The user end is responsible for efficient data preprocessing and optimization, while the network end performs in-depth demand analysis and resource allocation. The two form an end-to-end closed-loop system through close interaction, realizing efficient data transmission and dynamic resource allocation; (2) The user-end module of this invention has a built-in multi-level data preprocessing algorithm library, which can realize preprocessing functions such as time redundancy elimination and spatial redundancy optimization. By optimizing before data transmission, the amount of transmitted data is significantly reduced and the network burden is reduced; (3) This invention deploys a demand analysis module on the network end, which uses data mining and machine learning algorithms to accurately identify demands based on user data and historical demand information. The on-demand service submodule can dynamically adjust the resource allocation strategy according to the demand analysis results, realizing efficient collaboration of cross-domain (air, ground, space-based) resources; (4) The on-demand service submodule of this invention integrates a dynamic link selection algorithm and a resource allocation algorithm library, which supports real-time calculation of the best link and dynamic adjustment according to the availability of network resources. This design ensures that the service quality and efficiency of the system can be maximized even when resources are scarce. Attached Figure Description

[0032] Figure 1 is a schematic diagram of an on-demand service strategy system for air-space-ground scenarios based on user demand analysis provided by an embodiment of the present invention;

[0033] [Correction 03.01.2025 based on Rule 91] Figure 2a is a schematic diagram of the entity network of the present invention provided in an embodiment of the present invention; [0033.1] [Correction 03.01.2025 according to Article 91] Figure 2b is a schematic diagram of the virtual network of the present invention provided in an embodiment of the present invention; [0033.2] [Correction 03.01.2025 according to detailed rules 91] Figure 2c is a schematic diagram of the scheduling mechanism of the present invention provided in an embodiment of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0035] A system for on-demand service strategy in air-space-ground scenarios based on user demand analysis is shown in Figure 1. The system includes a user terminal and a network terminal, which are connected through a data processing module and a demand analysis module.

[0036] The user terminal includes a data acquisition module, a data processing module, and a request upload module connected in sequence.

[0037] The data acquisition module is used to acquire environmental data collected in real time by sensors on each user device;

[0038] For example, the user-side module is deployed on user devices, such as in-vehicle terminals or smartphones. The module initializes upon startup, including hardware resource detection and network connection configuration, ensuring the device has the capability for data acquisition and preprocessing. The network-side module consists of multiple functional sub-modules and is deployed on edge computing nodes or cloud platforms.

[0039] By calling the local sensor interface, the user-end module collects various user data, such as location information, environmental images, and device status. Based on preset acquisition frequency and data volume control strategies, this module dynamically adjusts the intensity of data acquisition to ensure energy savings for the device.

[0040] User devices (such as smart vehicles, mobile terminals, and IoT devices) collect environmental data in real time through built-in sensors, such as location, speed, road conditions, and information about nearby vehicles. For example, mobile terminals (such as smartphones and tablets) use built-in sensors (GPS, accelerometers, cameras, etc.) to collect data on user location, movement status, network connectivity, and application usage in real time. Smart vehicles collect data on road conditions, traffic flow, vehicle speed, and location through onboard sensors (such as radar, cameras, and LiDAR), and assess the vehicle's perception capabilities and network connectivity.

[0041] The data processing module is used to preprocess environmental data to obtain preprocessed environmental data. Here, preprocessing environmental data to obtain preprocessed environmental data includes: data filtering, time redundancy elimination, and spatial redundancy elimination of environmental data to obtain preprocessed environmental data.

[0042] For example, the built-in data preprocessing mechanism is optimized at the architectural level, including data filtering, redundancy elimination, and data compression. This design simplifies the data processing process and ensures data structure optimization before transmission through modular preprocessing, thereby reducing communication bandwidth usage. The user-side module establishes a secure communication channel with the network-side module, employing a lightweight protocol for data transmission. The architectural design ensures that user data can be transmitted quickly and stably to the network-side module in different network environments, reducing transmission latency.

[0043] The requirement upload module is used to propose business requirements based on preprocessed environmental data and obtain user equipment requirement data.

[0044] For example, user equipment (UE) proposes its own network service requirements based on real-time collected data. These requirements include data transmission bandwidth, accuracy of perceived data, latency tolerance, and service continuity.

[0045] In a specific embodiment of this invention, within a smart city system, each intelligent connected vehicle user (Ui) is treated as an intelligent agent, and base stations, drones, and satellites are referred to as service providers, corresponding to terrestrial networks, airborne networks, and space-based networks, respectively. When a user requests services from a service provider, they need to upload user equipment requirement data and pre-processed environmental data to the service provider.

[0046] The network-side module comprises a network status acquisition module, a demand analysis module, a policy calculation module, and a scheduling module, connected sequentially. This network-side module consists of multiple functional sub-modules deployed on edge computing nodes or cloud platforms. This architecture design enables distributed management of computing resources, automatically selecting the optimal service node based on user location and data traffic, thereby improving overall service efficiency. The network-side module integrates status awareness capabilities for air, ground, and space-based resources, collecting real-time availability and performance data. This functionality is implemented through a distributed monitoring architecture, dynamically sensing resource load and remaining capacity, and providing real-time data for resource allocation.

[0047] The network status acquisition module is used to obtain the current network resources in real time.

[0048] The requirement parsing module performs requirement parsing on user equipment requirement data and obtains the parsing results. Here, the requirement parsing of user equipment requirement data includes: using a layered architecture to classify user equipment requirement data and assign data priorities.

[0049] Here, the requirements parsing module uses a layered architecture to analyze user requirements. The first layer performs initial data classification, and the second layer performs detailed analysis and requirements prioritization. Through this layered architecture design, the system can quickly respond to high-priority requirements while ensuring that low-priority tasks are handled appropriately.

[0050] Depending on the characteristics of different devices, requirements can be categorized into several types, such as:

[0051] (1) For mobile terminals, the requirements include high-definition video streaming, real-time navigation, emergency calls, etc.

[0052] (2) For IoT devices, the requirements include low-latency device status uploading and real-time environmental monitoring.

[0053] (3) For intelligent vehicles, the requirements include high-precision environmental perception and real-time communication.

[0054] Assume there is user U in the intelligent connected vehicle, and the i-th user U i In the j-th time slot, a request is made to the network to watch a high-definition movie. Through the user demand analysis submodule, user U... i The requirement is expressed as: Where, β i,j T represents the amount of data required by the requested service. i,j This indicates the actual time cost required to complete the task. U i The maximum acceptable latency when requesting this service.

[0055] The strategy calculation module is used to calculate network resource allocation strategies based on the parsing results and preprocessed environmental data. Given the high cost and long latency of satellite communication, the intervention of space-based networks is only sought when base stations and drones cannot meet user needs. [0055.1] [Correction 03.01.2025 according to Article 91] As shown in Figure 2a, it is a schematic diagram of the physical network of the present invention provided in an embodiment of the present invention; Figure 2b shows a schematic diagram of the virtual network of the present invention provided in an embodiment of the present invention; Figure 2c shows a schematic diagram of the scheduling mechanism of the present invention provided in an embodiment of the present invention.

[0056] [Corrected according to Rule 91 03.01.2025] Here, the network resource allocation strategy is calculated based on the parsing results and preprocessed environmental data, including the following steps S1 to S9.

[0057] S1, determine the first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network respectively;

[0058] Here, the first data transmission rate is expressed as:

[0059] w i UE It is the bandwidth of UI, p i UE It is the transmit power of Ui, PL i,V For path loss between UI and drone, White noise power

[0060] The second data transmission rate is expressed as:

[0061] Among them, PL i,B For path loss between UI and drone, represents the Gaussian white noise power in the network.

[0062] S2, determine the third interference plus noise ratio of the signal between the terrestrial network and the user terminal, and calculate the third data transmission rate based on the third interference plus noise ratio; determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio; determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio;

[0063] Here, the interference plus noise ratio of the signal between the terrestrial network and the user terminal is expressed as:

[0064] Where, p Bχ represents the base station's transmit power. B,i θ represents the distance between the base station and the user. B,i Indicates base station and U i The path loss index of the link between them. This represents additive white Gaussian noise in the network.

[0065] Therefore, the third data transmission rate between the terrestrial network and user equipment can be derived:

[0066] Where, N B This indicates the number of users accessing the base station in the same time slot. This indicates that the bandwidth obtained by users accessing the base station is related to the number of users accessing the station, and an average allocation strategy is implemented.

[0067] When the resources of terrestrial and airborne networks cannot meet the needs of ground users, these networks need to seek assistance from satellite networks, using R... L,B The fourth data transmission rate, representing the data transfer rate between space-based and terrestrial networks, is expressed as:

[0068] Among them, W L PL represents the satellite's communication bandwidth, h represents the satellite's transmit power, and h represents the channel gain. This represents additive white Gaussian noise in the network.

[0069] Similarly, the interference-to-noise ratio and the fifth data transmission rate between the space-based network and the user terminal are expressed as:

[0070] Wherein, PL' represents path loss. This represents the power of white noise.

[0071] S3, calculate the actual time consumed by parsing the results based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate;

[0072] S4, determine the temporary reward function for each user device based on the actual time consumed, and determine the long-term reward function for each user device based on the temporary reward function;

[0073] S5. Based on the access volume constraint function and the storage volume constraint function of the land-based network and the air-based network, determine the first network environment state at time slot t.

[0074] S6. Optimize the long-term reward function based on the first network environment state at time slot t to obtain the optimized reward function;

[0075] S7. Based on the first network environment state at time slot t and the second network environment state at time slot t+1, determine the optimal strategy for the current user equipment and the optimal strategies for the other user equipment.

[0076] S8. Optimize the reward function based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. Solve the final reward function to obtain the network resource allocation policy for each user equipment.

[0077] For example, to model the uncertainty of a stochastic environment, the problem of vehicle users choosing resource providers is formulated as a stochastic game. Addressing the diverse service needs of users (UI) in vehicle-to-everything (V2X) systems, the goal of this invention is to select suitable service providers within an integrated air-space-ground network, providing services to as many users as possible under constraints of latency and economic cost, thereby maximizing the overall system service rate. In the method provided by this invention, the overall system service rate is the most critical issue. In time slot t, user (UI) selects its respective service provider based on the current environmental state, its own service needs, and latency requirements. It then receives a reward to evaluate the performance of its chosen action. Therefore, the design of the reward function directly guides the learning process. In the method, the temporary reward function for user (UI) is defined as:

[0078] in, Indicates the maximum acceptable latency when a user device requests a service; T i (t) represents the actual time consumed by the task; θ1 represents the cost weight; θ2 represents the weight of the overall service rate (Rate); C B Indicates the cost of using terrestrial networks; C V This represents the cost of using the space-based network; Rate represents the overall service rate. This represents the probability that a space-based network provides services. or N V Indicates the number of airborne network accesses; This indicates the probability that a land-based network will provide services. or N B This represents the number of terrestrial network accesses; i represents the i-th user device.

[0079] The overall service rate of the system is expressed as:

[0080] Among them, T i (t) represents the actual time consumed when a user requests the current service and the service provider provides the service.

[0081] The immediate reward for user Ui in any time slot t depends primarily on the information it observes.

[0082] 1) Observed information: the type of service currently requested (M), the action taken (the selected service provider), the number of Ui accesses in the previous time slot of the base station / drone, the storage capacity of the base station / drone, and the current system service rate.

[0083] 2) Unobserved information: Actions taken by other users Ui in the network and their benefits.

[0084] Next, consider maximizing the long-term reward v by selecting appropriate actions in each time slot. i (t). Specifically, at a certain time slot in the process, the discounted reward is the sum of the rewards in the current time slot, plus the sum of future rewards multiplied by a constant factor. Therefore, the user's long-term reward function can be expressed as:

[0085] Where λ represents the first constant; τ represents the slot exponent of the time step; r i (t+τ+1) represents the temporary reward for time slot t+τ+1; t represents the time step; v i (t) represents the long-term reward of the i-th user device in time slot t.

[0086] Specifically, the value of λ reflects the impact of future rewards on the optimal decision: if λ is close to 0, it means the decision emphasizes short-term gains; conversely, if λ is close to 1, it gives more weight to future rewards, and the decision is considered farsighted. Each user's action space A i ∈{0,1,2} represent: Ui abandons the service request, chooses a terrestrial network for service, and chooses an airborne network for service, respectively. In any time slot t, each user's goal is to take the optimal action a. * i (t)∈A i This maximizes the long-term reward in formula (9). Therefore, the optimal matching problem for end users Ui, i∈I can be formulated as:

[0087] Among them, a i This represents the action corresponding to the i-th user device; v i (t) represents the long-term reward of the i-th user device; t represents the time step; a * i (t) represents the optimal action corresponding to the i-th user device; A i It represents a set of actions.

[0088] Note that the resource-on-demand matching problem under consideration consists of I subproblems, corresponding to I distinct users. Furthermore, each user has no information about other users, such as actions taken or rewards received. In the network under consideration, it is assumed that all users Ui are selfish and rational. Therefore, at any time slot t, all users are informed by observing the current environmental state s. i (t)∈S i They choose their actions non-cooperatively. i (t)∈A i To maximize the long-term reward in (9).

[0089] Random games are a generalization of Markov decision processes in the multi-agent scenario, also known as Markov games, and are represented by tuples:<I,S,A,P,R> I is the number of users; S is the state set including the state of each user; A is the action set A = A1 × A2 × … × A U A i It is the action set of user Ui; P: S×A×S∈[0,1] is the state transition probability function; R={R1,R2,…,R U This includes rewards for all users.

[0090] Since user UIs access base stations or drones, service provider resources are allocated evenly. The resources each user receives are related to the number of accessing users and the service provider's capacity limitations. Changes in environmental conditions affect user UI selection. Because competing user UIs do not cooperate, the environmental condition is defined based on each user UI's local observation. In time slot t, the environmental condition observed by the user UI is given by the following formula: s i (t)=(M i (t),Ρ(t),Q(t)) (11)

[0091] Where M i (t) = {0, 1, 2} represents the service type requested by Ui in time slot t, which respectively represent Ui requesting data upload service, computing service, and storage service.

[0092] Because the resource provider employs a resource allocation strategy, the computing and bandwidth resources allocated to each user are related to the number of Ui connections. Too many Ui choosing the same resource provider will reduce the resources they acquire, significantly extending computation and communication time and impacting overall efficiency. Therefore, it is necessary to set a maximum allowed number of base stations / drones to connect. Let P(t)∈{0,1,2} represent whether the number of Ui connections to the base station or drone has reached its peak, expressed as:

[0093] Where, N Bmax N V max These represent the maximum number of users allowed to access the base station and the drone, respectively. P(t) = 0 indicates that neither the base station nor the drone has reached its access peak capacity, P(t) = 1 indicates that the base station's access capacity is full, and P(t) = 2 indicates that the drone's access capacity is full.

[0094] Q(t)∈{0,1,2} represents the storage capacity limit of the base station and the drone:

[0095] Where, β i (t) represents the amount of data uploaded by user Ui in time slot t, λ B ,λ V These represent the cache capacity of the base station and the drone, respectively.

[0096] π i (s i ,a i π is the mapping from state to action. i (s i ,a i )=Pr(a t =a|s t =s)∈[0,1] gives the state s)∈[0,1] i Take action a i The probability of being in state s. In other words, for each state s i ∈S i Ui has many different strategy choices, which can be represented as: π i (s i )={π i (s i ,a i )|a i ∈A i In random games, the joint strategy of I players is defined as the strategy vector π = (π1(s1), π2(s2), ..., π). I (s I Based on the above discussion, in the formalized stochastic game, the optimization objective for each user Ui is to maximize its expected return over time. Therefore, the optimization objective in formula (7) can be restated as:

[0097] Where Ε[·] represents the expectation operation; s i Indicates the current environmental state; v i (s i ,π) represents the state in s i The optimized reward is calculated using action π in the state; λ represents the first constant; τ represents the slot exponent of the time step; t represents the time step; si (t) represents the state function; r i (t+τ+1) represents the temporary reward in time slot t+τ+1.

[0098] From environmental state s i (t) to the new state s i The state transition at (t+1) is determined by the joint policy of all users. Furthermore, in non-cooperative games, at each time slot t, each user is in state s. i (t) independently chooses its strategy to maximize the optimal reward v i (s i ,π i The user device then receives its current personal reward based on the joint policy (action set) π. Therefore, it cannot be simply expected that the user device will maximize the corresponding expected reward.

[0099] Here, each user's goal is to start from any state s i ∈S i Learning the optimal strategy π i * The optimal strategy for other users is learned as π. -i * =(π1) * ,π2 * ,…,π i * ,…,π U * If the expected return can be restated as Equation (15), then the Nash equilibrium is used to describe a solution of the stochastic game, as defined in Definition 1.

[0100] The final reward function is expressed as:

[0101] Let s i (t)=s i s i (t+1)=s i Formula (15) can be broken down and transformed into:

[0102] Where Ε[·] represents the expectation operation; λ represents the first constant; r i (·) represents the temporary reward corresponding to the i-th user device; τ represents the slot exponent of the time step; t represents the time step; s i (t+τ) represents; π i Indicates the current user device's policy; π -i Indicates the policies of other user devices; s i (t) represents the state function.

[0103] Nash equilibrium is a set of I-optimal strategies, (π1) * ,π2 * ,…,π I * In this case, no user can gain any higher reward by simply changing their own strategy. That is, for each user Ui∈I, in each state s... i ∈S i ,have:

[0104] Among them, Π i It is a set of strategies that the user UI can adopt.

[0105] This means that in the formalized on-demand matching game problem, there always exists a NE strategy. In the network, each user obtains their own optimal strategy, and no end-user can obtain a better strategy by changing their own. Therefore, in this method, the goal of each user device Ui is to achieve optimal strategy in any state s. i Find an optimal strategy for NE (Non- ...

[0106] Each user observes their local information. i (t) and receive their own reward r i Since the future state S(t+1) depends only on the current state S(t) and the action A(t) taken, this dynamic multi-agent reinforcement learning process exhibits Markov properties and is therefore formulated as a stochastic game. Specifically, the stochastic game of a single user is modeled as a Markov decision process (MDP).

[0107] This paper applies Q-learning to solve the independent MDP problem and proposes a multi-agent q-learning (MA-Q) algorithm based on independent learning to address the resource on-demand matching problem among multiple users and services in the SAGIN network. In the proposed algorithm, each user runs an independent Q-learning algorithm and simultaneously learns a separate optimal policy for their MDP. Specifically, the choice of the optimal action depends on the Q function. According to the definition of Q-learning, the Q value is updated according to formula (18):

[0108] Where α is the step size (learning rate), 0 < α ≤ 1; s i ' is the next state observed after the current state executes the behavior policy.

[0109] To ensure the convergence of Q-learning, the learning rate is set according to the following equation (17):

[0110] Where, α begin αend These are the initial and final values ​​of α, respectively, and iterations is the maximum number of iterations for the learning algorithm.

[0111] The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation strategy; the network resources include: land-based network, air-based network and space-based network.

[0112] The network side also includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.

[0113] During system operation, the feedback module continuously optimizes resource allocation strategies through a feedback mechanism to ensure maximum task priority adjustment and resource utilization. If improper resource utilization or execution failures are detected, adjustments are made through the feedback mechanism. Based on real-time feedback information, the resource allocation strategy is dynamically adjusted. For example, if a resource is overloaded, some tasks are reassigned to other resources to optimize system performance.

[0114] The method provided by this invention realizes an on-demand service architecture that integrates air, space, and ground through multi-level demand analysis and resource allocation strategies. The system can not only analyze user needs in real time and quantify them into specific KPIs, but also intelligently allocate resources through advanced algorithms to ensure efficient utilization of network bandwidth, computing resources, and device endurance, thereby improving the overall traffic management system.

[0115] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0116] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present invention.

Claims

1.A space-air-ground scene on-demand service policy system based on user demand analysis, characterized in that, It includes: a user terminal and a network terminal, wherein the user terminal and the network terminal are connected through a data processing module and a demand parsing module; The user terminal includes a data acquisition module, a data processing module, and a demand upload module connected in sequence; the data acquisition module is used to acquire environmental data collected in real time by sensors on each user device, and the data processing module is used to preprocess the environmental data to obtain preprocessed environmental data; The demand upload module is used to propose business demands based on the preprocessed environmental data and obtain user equipment demand data. The network side includes a network status acquisition module, a demand parsing module, a strategy calculation module, and a scheduling module connected in sequence; The network status acquisition module is used to acquire current network resources in real time; The demand parsing module is used to parse the user equipment demand data to obtain the parsing result; the strategy calculation module is used to calculate the network resource allocation strategy based on the parsing result and the preprocessed environmental data. The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation strategy; wherein, the network resources include: land-based network, air-based network and space-based network. 2.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 1, wherein, The data acquisition module performs data preprocessing on the environmental data to obtain preprocessed environmental data, including: data filtering, time redundancy elimination, and spatial redundancy elimination on the environmental data to obtain preprocessed environmental data. 3.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 1, wherein, The requirement parsing module performs requirement parsing on the user equipment requirement data, including: classifying the user equipment requirement data and assigning data priorities using a hierarchical architecture. 4.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 1, wherein, The policy calculation module calculates a network resource allocation policy based on the parsing results and the preprocessed environmental data, including: The first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network are determined respectively. The third interference plus noise ratio of the signal between the land-based network and the user terminal is determined, and the third data transmission rate is calculated based on the third interference plus noise ratio. Determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio; Determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio; The actual time consumed by the parsing result is calculated based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate. A temporary reward function for each user device is determined based on the actual time consumed, and a long-term reward function for each user device is determined based on the temporary reward function. Based on the access volume constraint function of the land-based network and the storage volume constraint function of the land-based network and the air-based network, the first network environment state at time slot t is determined; The long-term reward function is optimized based on the first network environment state at time slot t to obtain the optimized reward function; Based on the first network environment state at time slot t and the second network environment state at time slot t+1, determine the optimal strategy for the current user equipment and the optimal strategies for the other user equipment. The optimization reward function is optimized based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. The final reward function is then solved to obtain the network resource allocation policy for each user equipment. 5.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 4, wherein, The temporary reward function is represented as: wherein denotes the maximum delay acceptable by the user equipment when requesting a service; T i (t) denotes the time consumed by the actual task; θ1 denotes the cost weight; θ2 denotes the weight of the overall service rate Rate; C B denotes the usage cost of the land-based network; C V denotes the usage cost of the air-based network; Rate denotes the overall service rate; denotes the probability of an aerial base network providing service; N V denotes the number of aerial base network accesses; denotes the probability that a terrestrial network provides service; N B denotes the number of terrestrial network accesses; i denotes the i-th user equipment. 6.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 5, wherein, The long-term reward function is represented as: where λ denotes a first constant; τ denotes a time slot index of a time step; r i (t + τ + 1) denotes a temporary reward of time slot t + τ + 1; t denotes a time step; v i (t) denotes a long-term reward of the i-th user equipment of time slot t. 7.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 4, wherein, The optimization reward function is represented as: where E[·] denotes an expectation operation; s i denotes a current environment state; v i (s i , π) denotes an optimized reward using an action π in a state s i ; λ denotes a first constant; τ denotes a time step slot index; t denotes a time step; s i (t) denotes a state function; r i (t+τ+1) denotes a temporary reward at a time slot t+τ+1. 8.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 4, wherein, The final reward function is represented as: wherein E[·] represents an expectation operation; λ represents a first constant; r i (·) represents a temporary reward corresponding to the i-th user equipment; τ represents a time step slot index; t represents a time step; s i (t+τ) represents; π i represents a strategy of the current user equipment; π -i represents a strategy of the remaining user equipment; s i (t) represents a state function. 9.The space-air-ground scene on-demand service policy system based on user demand analysis according to claim 1, wherein, The network terminal also includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.