A system for on-demand service strategies in air-space-ground scenarios based on user demand analysis
By designing a collaborative architecture between the user end and the network end, and combining multi-level data preprocessing and dynamic link selection algorithms, the problem of low resource allocation efficiency in air-space-ground scenarios is solved, achieving efficient resource allocation and improved user experience.
Patent Information
- Application Number
- CN202411647319.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing resource management mechanisms are unable to adapt in real time to the ever-changing user needs and resource status in air-space-ground scenarios, resulting in low efficiency and accuracy in resource allocation.
The system adopts an on-demand service strategy for air-space-ground scenarios based on user demand analysis. Through the collaborative architecture design of the user end and the network end, it realizes data preprocessing, demand analysis and resource allocation. It uses a multi-level data preprocessing algorithm library and dynamic link selection algorithm to perform precise resource allocation.
It achieves efficient data transmission and dynamic resource allocation, improves resource utilization and user experience, and ensures that the system's service quality and efficiency are maximized when resources are scarce.
Smart Images

Figure CN119676851B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Internet of Things (IoT) technology, and in particular to an on-demand service strategy system for air-space-ground scenarios based on user demand analysis. Background Technology
[0002] With the development of integrated air-space-ground networks, the diversity of Internet of Things (IoT) devices and user needs is increasing. These devices and user needs are distributed across different spatial environments and exhibit different service requirements at different times. Existing resource management mechanisms mostly adopt static resource configuration, which is difficult to adapt to the constantly changing needs and resource status in air-space-ground scenarios in real time.
[0003] With the rapid development of communication networks and information processing technologies, user demands for services are becoming increasingly diversified and personalized. To meet these dynamic needs within limited network resources, on-demand service strategies have become crucial for improving user experience and resource utilization. However, existing architectures lack effective demand resolution mechanisms and dynamic resource allocation algorithms, resulting in low efficiency and accuracy in resource allocation. Summary of the Invention
[0004] This invention provides an on-demand service strategy system for air, space, and ground scenarios based on user demand analysis. This solves the problem in existing technologies that lack in-depth analysis of personalized user needs and cannot achieve accurate resource allocation based on dynamically changing needs, thus realizing efficient data transmission and dynamic resource configuration.
[0005] This invention provides an on-demand service strategy system for air-space-ground scenarios based on user demand analysis. The system includes a user terminal and a network terminal, wherein the user terminal and the network terminal are connected through a data processing module and a demand analysis module.
[0006] The user terminal includes a data acquisition module, a data processing module, and a demand upload module connected in sequence. The data acquisition module is used to acquire environmental data collected in real time by sensors on each user device. The data processing module is used to preprocess the environmental data to obtain preprocessed environmental data. The demand upload module is used to propose business demands based on the preprocessed environmental data to obtain user device demand data.
[0007] The network terminal includes a network status acquisition module, a demand parsing module, a policy calculation module, and a scheduling module connected in sequence. The network status acquisition module is used to acquire current network resources in real time. The demand parsing module performs demand parsing on the user equipment demand data to obtain the parsing results. The policy calculation module is used to calculate a network resource allocation policy based on the parsing results and the preprocessed environmental data. The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation policy. The network resources include: terrestrial network, air-based network, and space-based network.
[0008] In one possible implementation, the data acquisition module, in the step of preprocessing the environmental data to obtain preprocessed environmental data, includes: performing data filtering, temporal redundancy elimination, and spatial redundancy elimination on the environmental data to obtain preprocessed environmental data.
[0009] In one possible implementation, the requirement parsing module, in the process of parsing the user equipment requirement data, includes: classifying the user equipment requirement data and assigning data priorities using a hierarchical architecture.
[0010] In one possible implementation, the policy calculation module, which calculates the network resource allocation policy based on the parsing result and the preprocessed environmental data, includes:
[0011] The first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network are determined respectively.
[0012] The third interference plus noise ratio of the signal between the land-based network and the user terminal is determined, and the third data transmission rate is calculated based on the third interference plus noise ratio.
[0013] Determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio;
[0014] Determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio;
[0015] The actual time consumed by the parsing result is calculated based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate.
[0016] A temporary reward function for each user device is determined based on the actual time consumed, and a long-term reward function for each user device is determined based on the temporary reward function.
[0017] Based on the access constraint function and storage constraint function of the land-based network and the air-based network, time slots are determined. The initial network environment status at that time;
[0018] According to time slot The long-term reward function is optimized based on the first network environment state to obtain the optimized reward function.
[0019] According to time slot The first network environment state and time slot The second network environment state at that time determines the optimal policy for the current user equipment and the optimal policies for the other user equipment.
[0020] The optimization reward function is optimized based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. The final reward function is then solved to obtain the network resource allocation policy for each user equipment.
[0021] In one possible implementation, the temporary reward function is expressed as:
[0022] ;
[0023] in, This indicates the maximum acceptable latency when a user device requests a service. This indicates the actual time consumed by the task; Indicates cost weight; Indicates the overall service rate The weights; This indicates the cost of using a land-based network; This indicates the cost of using a space-based network; Indicates the overall service rate; Indicates the number of airborne network accesses; Indicates the number of airborne network accesses; This indicates the probability that a terrestrial network will provide services. This indicates the number of terrestrial network connections.
[0024] In one possible implementation, the long-term reward function is expressed as:
[0025] ;
[0026] in, Denotes the first constant; The slot index represents the time step. Indicates time slot Temporary reward; Indicates the time step; Indicates time slot No. Long-term rewards for individual user devices.
[0027] In one possible implementation, the optimized reward function is expressed as:
[0028] ;
[0029] in, This represents the expectation operation; Indicates the current environmental state; Indicates in Take action in state Optimization rewards; Denotes the first constant; The slot index represents the time step. Indicates the time step; Represents the state function; Indicates time slot A temporary reward.
[0030] In one possible implementation, the final reward function is expressed as:
[0031] ;
[0032] in, This represents the expectation operation; Denotes the first constant; Indicates the first Temporary rewards corresponding to each user device; The slot index represents the time step. Indicates the time step; Indicates time slot State function; Indicates the current policy of the user device; Indicates the policies of other user devices; Represents the state function.
[0033] In one possible implementation, the network terminal further includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.
[0034] One or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0035] (1) This invention adopts a collaborative architecture design of user-end module and network-end module. The user end is responsible for efficient data preprocessing and optimization, while the network end performs in-depth demand analysis and resource allocation. The two form an end-to-end closed-loop system through close interaction, realizing efficient data transmission and dynamic resource allocation; (2) The user-end module of this invention has a built-in multi-level data preprocessing algorithm library, which can realize preprocessing functions such as time redundancy elimination and spatial redundancy optimization. By optimizing before data transmission, the amount of transmitted data is significantly reduced and the network burden is reduced; (3) This invention deploys a demand analysis module on the network end, which uses data mining and machine learning algorithms to accurately identify demands based on user data and historical demand information. The on-demand service submodule can dynamically adjust the resource allocation strategy according to the demand analysis results, realizing efficient collaboration of cross-domain (air, ground, space-based) resources; (4) The on-demand service submodule of this invention integrates a dynamic link selection algorithm and a resource allocation algorithm library, which supports real-time calculation of the best link and dynamic adjustment according to the availability of network resources. This design ensures that the service quality and efficiency of the system can be maximized even when resources are scarce. Attached Figure Description
[0036] Figure 1 A schematic diagram of an on-demand service strategy system for air-space-ground scenarios based on user demand analysis is provided in an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of the entity network-virtual network-scheduling mechanism of the present invention, provided for an embodiment of the present invention. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0039] A system for on-demand service strategies in air-space-ground scenarios based on user demand analysis, such as... Figure 1 As shown, the system includes a user terminal and a network terminal, which are connected through a data processing module and a demand analysis module.
[0040] The user terminal includes a data acquisition module, a data processing module, and a request upload module connected in sequence.
[0041] The data acquisition module is used to acquire environmental data collected in real time by sensors on each user device;
[0042] For example, the user-side module is deployed on user devices, such as in-vehicle terminals or smartphones. The module initializes upon startup, including hardware resource detection and network connection configuration, ensuring the device has the capability for data acquisition and preprocessing. The network-side module consists of multiple functional sub-modules and is deployed on edge computing nodes or cloud platforms.
[0043] By calling the local sensor interface, the user-end module collects various user data, such as location information, environmental images, and device status. Based on preset acquisition frequency and data volume control strategies, this module dynamically adjusts the intensity of data acquisition to ensure energy savings for the device.
[0044] User devices (such as smart vehicles, mobile terminals, and IoT devices) collect environmental data in real time through built-in sensors, such as location, speed, road conditions, and information about nearby vehicles. For example, mobile terminals (such as smartphones and tablets) use built-in sensors (GPS, accelerometers, cameras, etc.) to collect data on user location, movement status, network connectivity, and application usage in real time. Smart vehicles collect data on road conditions, traffic flow, vehicle speed, and location through onboard sensors (such as radar, cameras, and LiDAR), and assess the vehicle's perception capabilities and network connectivity.
[0045] The data processing module is used to preprocess environmental data to obtain preprocessed environmental data. Here, preprocessing environmental data to obtain preprocessed environmental data includes: data filtering, time redundancy elimination, and spatial redundancy elimination of environmental data to obtain preprocessed environmental data.
[0046] For example, the built-in data preprocessing mechanism is optimized at the architectural level, including data filtering, redundancy elimination, and data compression. This design simplifies the data processing process and ensures data structure optimization before transmission through modular preprocessing, thereby reducing communication bandwidth usage. The user-side module establishes a secure communication channel with the network-side module, employing a lightweight protocol for data transmission. The architectural design ensures that user data can be transmitted quickly and stably to the network-side module in different network environments, reducing transmission latency.
[0047] The requirement upload module is used to propose business requirements based on preprocessed environmental data and obtain user equipment requirement data.
[0048] For example, user equipment (UE) proposes its own network service requirements based on real-time collected data. These requirements include data transmission bandwidth, accuracy of perceived data, latency tolerance, and service continuity.
[0049] In a specific embodiment provided by the present invention, in a smart city system, each intelligent connected vehicle user... As an intelligent agent, base stations, drones, and satellites are referred to as service providers, corresponding to terrestrial networks, airborne networks, and space-based networks, respectively. When a user requests services from a service provider, they need to upload user equipment demand data and pre-processed environmental data to the service provider.
[0050] The network-side module comprises a network status acquisition module, a demand analysis module, a policy calculation module, and a scheduling module, connected sequentially. This network-side module consists of multiple functional sub-modules deployed on edge computing nodes or cloud platforms. This architecture design enables distributed management of computing resources, automatically selecting the optimal service node based on user location and data traffic, thereby improving overall service efficiency. The network-side module integrates status awareness capabilities for air, ground, and space-based resources, collecting real-time availability and performance data. This functionality is implemented through a distributed monitoring architecture, dynamically sensing resource load and remaining capacity, and providing real-time data for resource allocation.
[0051] The network status acquisition module is used to obtain the current network resources in real time.
[0052] The requirement parsing module performs requirement parsing on user equipment requirement data and obtains the parsing results. Here, the requirement parsing of user equipment requirement data includes: using a layered architecture to classify user equipment requirement data and assign data priorities.
[0053] Here, the requirements parsing module uses a layered architecture to analyze user requirements. The first layer performs initial data classification, and the second layer performs detailed analysis and requirements prioritization. Through this layered architecture design, the system can quickly respond to high-priority requirements while ensuring that low-priority tasks are handled appropriately.
[0054] Depending on the characteristics of different devices, requirements can be categorized into several types, such as:
[0055] (1) For mobile terminals, the requirements include high-definition video streaming, real-time navigation, emergency calls, etc.
[0056] (2) For IoT devices, the requirements include low-latency device status uploading and real-time environmental monitoring.
[0057] (3) For intelligent vehicles, the requirements include high-precision environmental perception and real-time communication.
[0058] Assuming intelligent connected vehicle users , No. individual users In the Each user requests to watch a high-definition movie from the network during a specific time slot. The user needs analysis submodule can then analyze the user's needs. The requirement is expressed as: ;in, Indicates the amount of data required by the requested service. This indicates the actual time cost required to complete the task. express The maximum acceptable latency when requesting this service.
[0059] The strategy calculation module is used to calculate network resource allocation strategies based on the parsing results and preprocessed environmental data. Given the high cost and long latency of satellite communication, the intervention of space-based networks is only sought when base stations and drones cannot meet user needs.
[0060] Here, the network resource allocation strategy is calculated based on the parsing results and the preprocessed environmental data, such as... Figure 2 As shown, it includes the following steps S1 to S9.
[0061] S1, determine the first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network respectively;
[0062] Here, the first data transmission rate is expressed as:
[0063] (1)
[0064] yes bandwidth, yes The transmission power, for Path loss between drones White noise power
[0065] The second data transmission rate is expressed as:
[0066] (2)
[0067] in, for Path loss between drones represents the Gaussian white noise power in the network.
[0068] S2, determine the third interference plus noise ratio of the signal between the terrestrial network and the user terminal, and calculate the third data transmission rate based on the third interference plus noise ratio; determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio; determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio;
[0069] Here, the interference plus noise ratio of the signal between the terrestrial network and the user terminal is expressed as:
[0070] (3)
[0071] in, This indicates the base station's transmit power. Indicates the distance between the base station and the user. Indicates base station and The path loss index of the link between them. This represents additive white Gaussian noise in the network.
[0072] Therefore, the third data transmission rate between the terrestrial network and user equipment can be derived:
[0073] (4)
[0074] in, This indicates the number of users accessing the base station in the same time slot. This indicates that the bandwidth obtained by users accessing the base station is related to the number of users accessing the station, and an average allocation strategy is implemented.
[0075] When the resources of terrestrial and airborne networks cannot meet the needs of ground users, these networks need to seek assistance from satellite networks. The fourth data transmission rate, representing the data transfer rate between space-based and terrestrial networks, is expressed as:
[0076] (5)
[0077] in, Indicates the satellite's communication bandwidth. Indicates the satellite's transmission power. h Indicates channel gain. This represents additive white Gaussian noise in the network.
[0078] Similarly, the interference-to-noise ratio and the fifth data transmission rate between the space-based network and the user terminal are expressed as:
[0079] (6)
[0080] in, Indicates path loss. This represents the power of white noise.
[0081] S3, calculate the actual time consumed by parsing the results based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate;
[0082] S4, determine the temporary reward function for each user device based on the actual time consumed, and determine the long-term reward function for each user device based on the temporary reward function;
[0083] S5. Based on the access volume constraint function and the storage volume constraint function of the terrestrial and airborne networks, determine the time slots. The initial network environment status at that time;
[0084] S6, according to time slot The long-term reward function is optimized based on the first network environment state at that time, resulting in the optimized reward function.
[0085] S7, according to time slot The first network environment state and time slot The second network environment state at that time determines the optimal policy for the current user equipment and the optimal policies for the other user equipment.
[0086] S8. Optimize the reward function based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. Solve the final reward function to obtain the network resource allocation policy for each user equipment.
[0087] For example, to model the uncertainty of a stochastic environment, the problem of vehicle users choosing resource providers is formulated as a stochastic game. This applies to users in a connected vehicle system. To meet the diverse business needs of users, the objective of this invention is to select a suitable service provider within an integrated air-space-ground network, and, under constraints of latency and economic cost, provide services to as many users as possible to achieve the highest overall system service rate. In the method provided by this invention, the overall system service rate is the most critical issue. (Users) In the time slot Based on the current environmental conditions and its own business needs and latency requirements, each user selects its own service provider. It then receives a reward to evaluate the performance of its chosen action. Therefore, the design of the reward function directly guides the learning process. In the method, the user is defined... The temporary reward function is expressed as:
[0088] (7)
[0089] in, This indicates the maximum acceptable latency when a user device requests a service. This indicates the actual time consumed by the task; Indicates cost weight; Indicates the overall service rate The weights; This indicates the cost of using a land-based network; This indicates the cost of using a space-based network; Indicates the overall service rate; This represents the probability that a space-based network will provide services. =0 or =1; Indicates the number of airborne network accesses; This indicates the probability that a land-based network will provide services. =0 or =1; Indicates the number of terrestrial network connections; Indicates the first One user device.
[0090] Overall system service rate Represented as:
[0091] (8)
[0092] in, This indicates the actual time consumed when a user requests the current service and the service provider provides the service.
[0093] user Any time slot The immediate rewards in the system depend primarily on the information observed.
[0094] 1) Observed information: The type of service currently requested (M), the action taken (A, the selected service provider), and the previous time slot of the base station / drone. Number of access points, storage capacity of base stations / drones, and current system service rate.
[0095] 2) Unobserved information: Other users on the network The actions taken and their benefits.
[0096] Next, consider maximizing long-term rewards by selecting appropriate actions in each time slot. Specifically, at a certain time slot in the process, the discounted reward is the sum of the rewards in the current time slot, plus the sum of future rewards multiplied by a constant factor. Therefore, the user's long-term reward function can be expressed as:
[0097] (9)
[0098] in, Denotes the first constant; The slot index represents the time step. Indicates time slot Temporary reward; Indicates the time step; Indicates time slot No. Long-term rewards for individual user devices.
[0099] Specifically, The value reflects the impact of future rewards on the optimal decision: if A value close to 0 indicates that the decision-making process emphasizes short-term gains; in contrast, if... Approaching 1 gives more weight to future rewards, a decision considered farsighted. Each user's action space. They are represented as follows: Abandon the service request, choose a terrestrial network for service, or choose an airborne network for service. In any time slot t, each user's goal is to take the optimal action. This maximizes the long-term reward in formula (9). Therefore, the end user , The optimal matching problem can be formulated as follows:
[0100] (10)
[0101] in, Indicates the first Actions corresponding to each user device; Indicates the first Long-term rewards for individual user devices; Indicates the time step; Indicates the first The best action for each user device; It represents a set of actions.
[0102] Note that the resource-on-demand matching problem under consideration consists of I subproblems, corresponding to I distinct users. Furthermore, each user has no information about other users, such as actions taken or rewards received. In the network under consideration, it is assumed that all users... They are all selfish and rational. Therefore, in any time slot t, all users observe the current environmental state. They choose their actions non-cooperatively. To maximize the long-term reward in (9).
[0103] Random games are a generalization of Markov decision processes in the multi-agent scenario, also known as Markov games, and are represented by tuples: . It refers to the number of users; It is a set of states that includes the state of each user; For action sets , User Action set; Let be the state transition probability function; Includes rewards for all users.
[0104] Due to users When connecting to a base station or drone, the service provider's resources are allocated equally. The resources each user receives are related to the number of users accessing the service and the service provider's own capacity limitations. Changes in environmental conditions will affect users. The choice, due to competing users There is no cooperation between them, therefore the environment state is based on each user. Defined by local observations. In time slots. ,user The observed environmental state is given by the following formula:
[0105] (11)
[0106] in express The service type requested in time slot t represents respectively Request data upload, computing, and storage services.
[0107] Because the resource provider adopts a resource equalization strategy, each user is allocated computing resources and bandwidth resources accordingly. It's related to the number of connections; too many connections... Choosing the same resource provider reduces the resources available, significantly increases computation and communication time, and consequently impacts overall efficiency. Therefore, it's necessary to set a maximum allowed number of base stations / drones to connect. Indicating base stations and drones Whether the access volume has reached its peak is indicated by:
[0108] (12)
[0109] in, These represent the maximum number of users allowed to access the base station and the drone, respectively. This indicates that neither the base station nor the drone's capacity has reached its peak access capacity. This indicates that the base station's access capacity is full. This indicates that the number of drones that can be connected has been reached.
[0110] Indicates the storage capacity limits for base stations and drones:
[0111] (13)
[0112] in, Indicates user The amount of data uploaded in time slot t , These represent the cache capacity of the base station and the drone, respectively.
[0113] It's a mapping from state to action. The state is given Take action The probability. In other words, for each state... of It has many different strategy options, which can be represented as: In random games, I The joint strategy of each player is defined as a strategy vector. Based on the above discussion, in a formalized random game, each user... The optimization objective is to maximize the expected return over time. Therefore, the optimization objective in equation (7) can be restated as:
[0114] (14)
[0115] in, This represents the expectation operation; Indicates the current environmental state; Indicates in Take action in state Optimization rewards; Denotes the first constant; The slot index represents the time step. Indicates the time step; Represents the state function; Indicates time slot A temporary reward.
[0116] From the environmental state To the new state The state transitions are determined by the joint policy of all users. Furthermore, in non-cooperative games, at each time slot t, each user is in state... Each independently selects its strategy to maximize and optimize rewards. Then, based on the joint strategy (set of actions) It receives its current personal reward. Therefore, it cannot be simply expected that the user's device will maximize the corresponding expected reward.
[0117] Here, each user's goal is to start from any state. Learning the optimal strategy The optimal strategy of other users is learned as Then the expected return can be restated as Equation (15). A solution to a stochastic game is described by Nash equilibrium, which is defined by Definition 1.
[0118] The final reward function is expressed as:
[0119] (15)
[0120] make , Formula (15) can be broken down and transformed into:
[0121] (16)
[0122] in, This represents the expectation operation; Denotes the first constant; Indicates the first Temporary rewards corresponding to each user device; The slot index represents the time step. Indicates the time step; Indicates time slot State function; Indicates the current policy of the user device; Indicates the policies of other user devices; Represents the state function.
[0123] Nash equilibrium is a set of I-optimal strategies. No user can gain a higher reward simply by changing their own strategy. In other words, for each user... In each state ,have:
[0124] (17)
[0125] in, User A set of strategies that can be adopted.
[0126] This means that in the formulaic on-demand matching game problem, there always exists a no-predictor (NE) strategy. In the network, each user obtains their own optimal strategy, and no end-user can obtain a better strategy by changing their own. Therefore, in this method, each user device... The goal is to be in any state Find an optimal strategy for NE (Non- ...
[0127] Each user observes their local information And receive your own reward. Due to future state Depends only on the current state and the actions taken Therefore, this dynamic multi-agent reinforcement learning process possesses Markov properties, and thus it is formulated as a stochastic game. Specifically, the stochastic game of a single user is modeled as a Markov decision process (MDP).
[0128] This paper applies Q-learning to solve the independent MDP problem and proposes a multi-agent q-learning (MA-Q) algorithm based on independent learning to address the resource on-demand matching problem among multiple users and services in the SAGIN network. In the proposed algorithm, each user runs an independent Q-learning algorithm and simultaneously learns a separate optimal policy for their MDP. Specifically, the choice of the optimal action depends on the Q function. According to the definition of Q-learning, the Q value is updated according to formula (18):
[0129] (18)
[0130] in It is the step size (learning rate). ; It is the next state observed after the current state executes the behavior policy.
[0131] To ensure the convergence of Q-learning, the learning rate is set according to the following equation (17):
[0132] (19)
[0133] in, , They are given initial and final values, It represents the maximum number of iterations for the learning algorithm.
[0134] The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation strategy; the network resources include: land-based network, air-based network and space-based network.
[0135] The network side also includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.
[0136] During system operation, the feedback module continuously optimizes resource allocation strategies through a feedback mechanism to ensure maximum task priority adjustment and resource utilization. If improper resource utilization or execution failures are detected, adjustments are made through the feedback mechanism. Based on real-time feedback information, the resource allocation strategy is dynamically adjusted. For example, if a resource is overloaded, some tasks are reassigned to other resources to optimize system performance.
[0137] The method provided by this invention realizes an on-demand service architecture that integrates air, space, and ground through multi-level demand analysis and resource allocation strategies. The system can not only analyze user needs in real time and quantify them into specific KPIs, but also intelligently allocate resources through advanced algorithms to ensure efficient utilization of network bandwidth, computing resources, and device endurance, thereby improving the overall traffic management system.
[0138] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0139] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present invention.
Claims
1. A system for on-demand service strategies in air-space-ground scenarios based on user demand analysis, characterized in that, It includes: a user terminal and a network terminal, wherein the user terminal and the network terminal are connected through a data processing module and a demand parsing module; The user terminal includes a data acquisition module, a data processing module, and a demand upload module connected in sequence; the data acquisition module is used to acquire environmental data collected in real time by sensors on each user device, and the data processing module is used to preprocess the environmental data to obtain preprocessed environmental data; The demand upload module is used to propose business demands based on the preprocessed environmental data and obtain user equipment demand data. The network side includes a network status acquisition module, a demand parsing module, a strategy calculation module, and a scheduling module connected in sequence; The network status acquisition module is used to acquire current network resources in real time; The demand parsing module is used to parse the user equipment demand data to obtain the parsing result; the strategy calculation module is used to calculate the network resource allocation strategy based on the parsing result and the preprocessed environmental data. The scheduling module is used to schedule air-space-ground network resources according to the network resource allocation strategy; wherein, the network resources include: land-based network, air-based network, and space-based network; The requirement parsing module performs requirement parsing on the user equipment requirement data, including: classifying the user equipment requirement data and assigning data priorities using a hierarchical architecture; The policy calculation module calculates a network resource allocation policy based on the parsing results and the preprocessed environmental data, including: The first data transmission rate between the user terminal and the air-based network and the second data transmission rate between the user terminal and the land-based network are determined respectively. The third interference plus noise ratio of the signal between the land-based network and the user terminal is determined, and the third data transmission rate is calculated based on the third interference plus noise ratio. Determine the fourth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fourth data transmission rate based on the fourth interference plus noise ratio; Determine the fifth interference plus noise ratio of the signal between the space-based network and the user terminal, and calculate the fifth data transmission rate based on the fifth interference plus noise ratio; The actual time consumed by the parsing result is calculated based on the first data transmission rate, the second data transmission rate, the third data transmission rate, the fourth data transmission rate, and the fifth data transmission rate. A temporary reward function for each user device is determined based on the actual time consumed, and a long-term reward function for each user device is determined based on the temporary reward function. Based on the access constraint function and storage constraint function of the land-based network and the air-based network, time slots are determined. The initial network environment status at that time; According to time slot The long-term reward function is optimized based on the first network environment state at that time to obtain the optimized reward function; According to time slot The first network environment state and time slot The second network environment state at that time determines the optimal policy for the current user equipment and the optimal policies for the other user equipment. The optimization reward function is optimized based on the optimal policy of the current user equipment and the optimal policies of the other user equipment to obtain the final reward function. The final reward function is then solved to obtain the network resource allocation policy for each user equipment.
2. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 1, characterized in that, The data acquisition module performs data preprocessing on the environmental data to obtain preprocessed environmental data, including: data filtering, time redundancy elimination, and spatial redundancy elimination on the environmental data to obtain preprocessed environmental data.
3. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 1, characterized in that, The temporary reward function is expressed as follows: ; in, This indicates the maximum acceptable latency when a user device requests a service. This indicates the actual time consumed by the task; Indicates cost weight; Indicates the overall service rate The weights; This indicates the cost of using a terrestrial network; This indicates the cost of using a space-based network; Indicates the overall service rate; This indicates the probability that a space-based network will provide services. Indicates the number of airborne network accesses; This indicates the probability that a terrestrial network will provide services. Indicates the number of terrestrial network connections; Indicates the first One user device.
4. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 3, characterized in that, The long-term reward function is expressed as follows: ; in, Denotes the first constant; The slot index represents the time step. Indicates time slot Temporary reward; Indicates the time step; Indicates time slot No. Long-term rewards for individual user devices.
5. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 1, characterized in that, The optimized reward function is expressed as follows: ; in, This represents the expectation operation; Indicates the current environmental state; Indicates in Take action in state Optimization rewards; Denotes the first constant; The slot index represents the time step. Indicates the time step; Represents the state function; Indicates time slot A temporary reward.
6. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 1, characterized in that, The final reward function is expressed as follows: ; in, This represents the expectation operation; Denotes the first constant; Indicates the first Temporary rewards corresponding to each user device; The slot index represents the time step. Indicates the time step; Indicates time slot State function; Indicates the current policy of the user device; Indicates the policies of other user devices; Represents the state function.
7. The air-space-ground scenario on-demand service strategy system based on user demand analysis according to claim 1, characterized in that, The network terminal also includes a feedback module, which feeds back the network resource allocation strategy to the strategy calculation module through a feedback mechanism.
Citation Information
Patent Citations
Unmanned aerial vehicle resource dynamic deployment method based on differentiated services
CN113242556A
Network service access and slice resource configuration method based on deep reinforcement learning
CN116095720A