Information processing method and device, network system and storage medium
By applying Q learning methods in NTN gateways and optimizing network resource allocation, the problem of difficulty in providing high performance and full coverage of existing networks is solved, and low latency and high real-time network communication is achieved.
Patent Information
- Application Number
- CN202311514939.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
Existing networks are difficult to provide high-performance requirements such as VR and AR and full coverage networks, resulting in large delays and difficult to meet users' real-time communication needs.
By applying the Q-learning method in the gateway of a non-terrestrial network (NTN), the expected reward table for network status and network resource allocation method is initialized, the expected reward table is updated based on the return information, and the network resource allocation method is optimized to reduce delay.
It realizes optimization of network resource allocation, reduces latency, improves the real-time and coverage capabilities of the network, and meets high-performance needs such as VR and AR.
Smart Images

Figure CN120017119A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of communications, and in particular to an information processing method, device, network system and storage medium. Background Art
[0002] With the widespread application of technologies such as VR (Virtual Reality) and AR (Augmented Reality), how to achieve high real-time data transmission, low latency, high transmission rate and other requirements is an important research direction in this field. The existing WIFI network may not be sufficient to provide the high performance requirements of VR and AR, and it is difficult to achieve full network coverage. It is difficult to meet the need to minimize latency under the communication needs of VR, AR and other users with a wider range of communication needs. Summary of the invention
[0003] In order to overcome the problems existing in the related art, the present disclosure provides an information processing method, device and storage medium to overcome the problem that the existing network is difficult to provide high performance requirements and full coverage network.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an information processing method, the method being applied to a gateway of a non-terrestrial network (NTN), the gateway being connected to the NTN via a satellite signal; the method comprising:
[0005] Initialize the expected reward table of network status and network resource allocation method;
[0006] Determine the current network status;
[0007] Allocating network resources based on one of the network resource allocation modes in the network state;
[0008] Obtaining reward information obtained from allocating the network resources;
[0009] According to the reward information, updating the expected reward table;
[0010] The network resource allocation method is determined according to the updated expected reward table.
[0011] In some embodiments, updating the expected reward table according to the reward information includes:
[0012] Determine, based on the reward information, an actual reward for executing the network resource allocation method in the current state;
[0013] Determine each of the expected rewards after update according to the current value of each expected reward, the actual reward obtained in the current state, and the maximum expected reward value of the expected reward table.
[0014] In some embodiments, the expected reward table includes:
[0015] The expected reward for each network resource allocation method corresponding to each of the network states.
[0016] In some embodiments, the network resource allocation includes:
[0017] Providing time slots to user terminals;
[0018] Allocate a base station for the user terminal; wherein the base station includes: a terahertz femtocell base station.
[0019] In some embodiments, the network status includes at least one of the following:
[0020] Channel information;
[0021] Network resources.
[0022] In some embodiments, each expected reward in the expected reward table is a parameter representing a delay-related parameter;
[0023] Among them, the shorter the delay is, the greater the value of the expected reward is.
[0024] In some embodiments, determining the network resource allocation method according to the updated expected reward table includes:
[0025] Generate random numbers;
[0026] When the random number is greater than or equal to a preset probability value, executing any network resource allocation method;
[0027] When the random number is less than the probability value, the network resource allocation method corresponding to the maximum reward value in the current environment state in the expected reward table is executed.
[0028] According to a third aspect of an embodiment of the present disclosure, an information processing device is provided, the device being applied to a gateway of an NTN, the gateway being connected to the NTN via a satellite signal; the device comprising:
[0029] An initialization module, configured to initialize the network state and the expected reward table of the network resource allocation method;
[0030] A first determination module, configured to determine a current network state;
[0031] An execution module, configured to perform network resource allocation based on one of the network resource allocation modes under the network state;
[0032] An acquisition module, configured to acquire reward information obtained by performing the network resource allocation;
[0033] An updating module, configured to update the expected reward table according to the reward information;
[0034] The second determination module is configured to determine a network resource allocation method according to the updated expected reward table.
[0035] In some embodiments, the update module includes:
[0036] A first determination submodule is configured to determine, based on the reward information, an actual reward for executing the network resource allocation method in the current state;
[0037] The second determination submodule is configured to determine each of the expected rewards after update according to the current value of each expected reward, the actual reward obtained in the current state, and the maximum expected reward value of the expected reward table.
[0038] In some embodiments, the expected reward table includes:
[0039] The expected reward for each network resource allocation method corresponding to each of the network states.
[0040] In some embodiments, the network resource allocation includes:
[0041] Provide time slots (TS) to user terminals;
[0042] Allocate a base station for the user terminal; wherein the base station includes: a terahertz femtocell base station.
[0043] In some embodiments, the network status includes at least one of the following:
[0044] Channel information;
[0045] Network resources.
[0046] In some embodiments, each expected reward in the expected reward table is a parameter representing a delay-related parameter;
[0047] Among them, the shorter the delay is, the greater the value of the expected reward is.
[0048] In some embodiments, the second determining module includes:
[0049] A generation submodule configured to generate random numbers;
[0050] A first execution submodule is configured to execute any network resource allocation method when the random number is greater than or equal to a preset probability value;
[0051] The second execution submodule is configured to execute the network resource allocation method corresponding to the maximum reward value in the current environment state in the expected reward table when the random number is less than the probability value.
[0052] According to a third aspect of an embodiment of the present disclosure, a network system is provided, including:
[0053] Non-terrestrial network NTN;
[0054] A gateway connected to the NTN via a satellite signal;
[0055] A plurality of base stations are connected to the gateway via an optical fiber network; wherein the base stations are used to provide terahertz network connections for a plurality of user terminals.
[0056] In some embodiments, the base station is: a terahertz femtocell base station;
[0057] The user terminal includes at least: a terminal supporting VR services.
[0058] In some embodiments, the NTN network is used to transmit VR data with the gateway via Ka-band transmission;
[0059] The gateway is used to transmit the VR data to the base station via an optical fiber link;
[0060] The base station is used to transmit the VR data to the user terminal via a terahertz frequency band.
[0061] In some embodiments, the NTN includes: Low Earth Orbit (LEO) satellites.
[0062] According to a fourth aspect of an embodiment of the present disclosure, there is provided an information processing device, including:
[0063] processor;
[0064] a memory for storing processor-executable instructions;
[0065] Wherein, the processor is configured to: implement the steps in any of the above-mentioned information processing methods when executed.
[0066] According to a fifth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, and when instructions in the storage medium are executed by a processor of an information processing device, the device is enabled to perform the steps in any one of the above-mentioned information processing methods.
[0067] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:
[0068] The disclosed embodiment provides an information processing method, which updates the expected reward table for network status and network resource allocation mode according to the reward information under network resource allocation, and further allocates network resources according to the updated expected reward table. In this way, the network resource allocation mode can be continuously optimized, network resource conflicts can be reduced, and transmission delays can be reduced.
[0069] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0071] Figure 1 The following is a flow chart of an information processing method according to an exemplary embodiment. Figure 1 ;
[0072] Figure 2 is a schematic diagram of a method for reducing transmission delay in a network system architecture according to an exemplary embodiment;
[0073] Figure 3 is a block diagram of an information processing device according to an exemplary embodiment;
[0074] Figure 4 A network system architecture is shown according to an exemplary embodiment. Figure 1 ;
[0075] Figure 5 A network system architecture is shown according to an exemplary embodiment. Figure 2 ;
[0076] Figure 6 A hardware structure frame of an information processing device according to an exemplary embodiment is shown Figure 1 ;
[0077] Figure 7 A hardware structure frame of an information processing device according to an exemplary embodiment is shown Figure 2 . DETAILED DESCRIPTION
[0078] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0079] The following is a detailed description:
[0080] Figure 1 The following is a flow chart of an information processing method according to an exemplary embodiment. Figure 1 ,This method is applied to the gateway of NTN, and the gateway is connected to NTN via satellite signals; Figure 1 As shown, the method includes:
[0081] In step S101, the expected reward table of network status and network resource allocation mode is initialized;
[0082] In step S102, the current network status is determined;
[0083] In step S103, network resource allocation is performed based on a network resource allocation method under the network state;
[0084] In step S104, obtaining the reward information obtained by allocating network resources;
[0085] In step S105, the expected reward table is updated according to the reward information;
[0086] In step S106, the network resource allocation method is determined according to the updated expected reward table.
[0087] NTN is one of the technical solutions for mobile terminals such as mobile phones to directly connect to satellites. It can use the integration of satellite communication networks and ground-based 5G networks to provide coverage without the restrictions of terrain and topography. 5G NTN technology can enable terminals to directly connect to cellular broadband networks via satellites, and can also break through the one-way transmission limitations of mobile phones to achieve two-way communication and network connectivity.
[0088] The NTN system may include a satellite communication network, a high altitude platform system (HAPS) and an air-to-ground network. The satellite communication network includes satellite-borne platforms including low earth orbiting (LEO) satellites, medium earth orbiting (MEO) satellites and geosynchronous earth orbiting (GEO) satellites. In the disclosed embodiment, a network system architecture is established by taking a low earth orbit satellite in a satellite communication network as an example, but it is not limited to low earth orbit satellites. Other forms of NTN systems may also be applicable to the technical solutions of the disclosed embodiments.
[0089] The above method is the application of Q-learning in the field of artificial intelligence. Q-learning is a reinforcement learning (RL) method that does not require prior information about the system model. This method can automatically learn the optimal strategy of the system through mutual learning with the environment, and make decisions in practical applications based on the optimal strategy.
[0090] In the disclosed embodiment, the Q learning method is used to perform reinforcement learning based on the network communication between NTN and the terahertz base station architecture in the NTN environment to obtain the optimal solution to the latency problem, thereby achieving the purpose of reducing transmission delay in the above network architecture, and providing VR users with a good performance experience.
[0091] In the disclosed embodiment, the above Q learning process can be performed by an agent or a "learning agent" to realize the allocation of network resources. Therefore, the agent can be set in the gateway, and the gateway can be used to realize intelligent network resource allocation. For the convenience of description, the "gateway" is directly used as the execution subject for description.
[0092] like Figure 2 As shown, the agent 210 and the network environment 220 learn from each other, and the agent detects the current environment s i , and perform action a based on the current environment i , and get the return information r i .
[0093] Specifically, by corresponding different network states to different network resource allocation methods, a corresponding expected reward table, i.e., a Q table, can be established. In some embodiments, the expected reward table includes: the expected reward for each network resource allocation method corresponding to each network state. That is, each item in the Q table represents the expected reward value for formulating the network resource allocation method under the corresponding network state, i.e., the Q value, as shown in Table 1 below. The network state is represented by s i Indicates that the network resource allocation method is composed of a i The Q value is represented by q(s i ,a i ) represents. Since the target problem to be solved by the above method in the embodiment of the present disclosure is to reduce network delay and seek the optimal solution with low network delay, the Q value here represents a parameter related to network delay. For example, the larger the network delay, the smaller the Q value. Therefore, the process of seeking the optimal solution is to continuously update the Q table and take the item with the largest Q value as the optimal combination. Therefore, when the network resource allocation is subsequently performed, for a network state s k , we can select the a corresponding to the network resource allocation method when its corresponding Q value is the largest, that is, max Q(sk,ai) k Performs allocation of network resources.
[0094] Q Table <![CDATA[a1]]> <![CDATA[a2]]> <![CDATA[s1]]> <![CDATA[q(s1,a1)]]> <![CDATA[q(s1,a2)]]> <![CDATA[s2]]> <![CDATA[q(s2,a1)]]> <![CDATA[q(s2,a2)]]> <![CDATA[s3]]> <![CDATA[q(s3,a1)]]> <![CDATA[q(s3,a2)]]>
[0095] Table 1
[0096] Exemplarily, the gateway may initialize the Q table by setting each directional Q value to 0 or setting it to the same fixed value.
[0097] After initialization, the gateway can determine the current network status by detecting the occupancy of each channel in the network, delay information, etc. Since it is the first allocation after initialization, the gateway allocates network resources based on the current network status and any network resource allocation method.
[0098] After executing a network resource allocation, the network state will change accordingly, that is, it will enter the next network state. At the same time, based on the changed network state, the system will return corresponding reward information. The reward information and Q value are the same type of parameters, but the Q value represents the expected reward value, while the reward information represents the actual reward value after the action is executed. Therefore, the gateway can update the corresponding Q value according to the reward information.
[0099] In the next network state, the gateway continues to perform network resource allocation based on a network resource allocation method. The action performed at this time can be based on the current optimal network resource allocation method in the Q table, or it can be a random network resource allocation method. Then the network state changes further, and the system returns the corresponding feedback information again. In this cycle, the gateway can update the corresponding Q value based on the received feedback information during the process of network resource allocation, thereby updating the Q table, and further determine the network resource allocation method based on the updated Q table and allocate network resources.
[0100] In some embodiments, updating the expected reward table according to the return information includes:
[0101] Based on the reward information, determine the actual reward for executing the network resource allocation method in the current state;
[0102] Determine each updated expected reward based on the current value of each expected reward, the actual reward obtained in the current state, and the maximum expected reward value of the expected reward table.
[0103] Specifically, the above-mentioned expected reward table, i.e., the Q table, can be updated according to the following formula (1).
[0104]
[0105] Among them, Q(s i ,a i ) means in state s iTake action in i The expected reward, r i represents the actual reward after the action is executed, α represents the learning rate, γ· represents the attenuation coefficient (or discount factor), is the maximum expected reward value in the future. According to the above formula, the expected reward table can be continuously updated to gradually approach the optimal expected reward.
[0106] In some embodiments, network resource allocation includes: providing time slots to user terminals; allocating base stations to user terminals; wherein the base stations include: terahertz femtocell base stations.
[0107] In some embodiments, the network status includes at least one of the following: channel information; network resources.
[0108] In some embodiments, each expected reward in the expected reward table is a parameter related to the delay; wherein, the shorter the delay, the greater the expected reward value. Exemplarily, the maximum expected reward value is set to 1 and the minimum is set to 0. If the delay satisfies the predetermined parameter, the expected reward value is 1; if it is greater than the predetermined parameter, the expected reward value is a value between 0 and 1.
[0109] In some embodiments, determining a network resource allocation method according to the updated expected reward table includes:
[0110] Generate random numbers;
[0111] When the random number is greater than or equal to the preset probability value, any network resource allocation method is executed;
[0112] When the random number is less than the probability value, the network resource allocation method corresponding to the maximum reward value in the current environment state in the expected reward table is executed.
[0113] In order to avoid the above method of updating the Q table and executing subsequent actions based on the Q table from falling into the local optimal solution, a part of the actions for executing network resource allocation can be specified as exploration actions, that is, the corresponding reward information is obtained after random execution. If a larger reward value is encountered, the maximum expected reward in the Q table can be further updated. The other part executes the current optimal action based on the Q table, thereby avoiding executing actions with lower Q values, and performing subsequent network resource allocation in a more optimal allocation method as much as possible, that is, the "greedy strategy" of the following formula (2).
[0114]
[0115] Among them, RandomAction represents a random action, ψ represents a random number, and ε represents the probability of action exploration.
[0116] Here, each time a network resource allocation action is executed, a random number can be generated, and then the random number can be compared with a preset probability value. The preset probability value is also called the probability of action exploration, which represents the probability of randomly executing the network resource allocation action for exploration. If the random number is greater than the probability value, any network resource allocation action is executed to achieve new exploration. If the random number is less than the probability value, the current optimal action is executed based on the Q table.
[0117] Figure 3 is a block diagram of an information processing device according to an exemplary embodiment, the device is applied to a gateway of an NTN, and the gateway NTN is connected via a satellite signal; Figure 3 The device 300 shown comprises:
[0118] Initialization module 301, configured to initialize the expected reward table of network status and network resource allocation mode;
[0119] A first determination module 302, configured to determine a current network state;
[0120] An execution module 303 is configured to perform network resource allocation based on a network resource allocation method under a network state;
[0121] An acquisition module 304 is configured to acquire reward information obtained by allocating network resources;
[0122] An updating module 305 is configured to update the expected reward table according to the reward information;
[0123] The second determination module 306 is configured to determine a network resource allocation method according to the updated expected reward table.
[0124] In some embodiments, the update module includes:
[0125] A first determination submodule is configured to determine an actual reward for executing a network resource allocation method in a current state according to the reward information;
[0126] The second determination submodule is configured to determine each updated expected reward according to the current value of each expected reward, the actual reward obtained in the current state, and the maximum expected reward value of the expected reward table.
[0127] In some embodiments, the expected reward table includes:
[0128] The expected rewards for each network resource allocation method corresponding to each network state.
[0129] In some embodiments, network resource allocation includes:
[0130] Providing time slots to user terminals;
[0131] Allocate a base station to the user terminal; wherein the base station includes: a terahertz femtocell base station.
[0132] In some embodiments, the network status includes at least one of the following:
[0133] Channel information;
[0134] Network resources.
[0135] In some embodiments, each expected reward in the expected reward table is a parameter representing a delay-related parameter;
[0136] Among them, the shorter the delay, the greater the expected reward value.
[0137] In some embodiments, the second determining module includes:
[0138] A generation submodule configured to generate random numbers;
[0139] A first execution submodule is configured to execute any network resource allocation method when the random number is greater than or equal to a preset probability value;
[0140] The second execution submodule is configured to execute the network resource allocation method corresponding to the maximum reward value in the current environment state in the expected reward table when the random number is less than the probability value.
[0141] Figure 4 is a diagram showing a network system architecture according to an exemplary embodiment. Figure 4 As shown, a network system 400 is provided, including:
[0142] Non-terrestrial network (NTN) 110;
[0143] The gateway 120 is connected to the NTN 110 via a satellite signal; the gateway 120 can be used as an intelligent agent in the system to execute the method in any of the above embodiments;
[0144] A plurality of base stations 130 are connected to the gateway 120 via an optical fiber network; wherein the base stations 130 are used to provide terahertz network connections for a plurality of user terminals 140 .
[0145] In other embodiments, the method in the above embodiments may also be executed by other devices, and the allocation of network resources may be achieved by communicating with the gateway 120 .
[0146] The above-mentioned NTN is one of the technical solutions for mobile terminals such as mobile phones to directly connect to satellites. It can use the integration of satellite communication networks and ground-based 5G networks to provide coverage without the restrictions of terrain and topography. 5G NTN technology can enable terminals to directly connect to cellular broadband networks via satellites, and can also break through the one-way transmission limitations of mobile phones to achieve two-way communication and network connectivity.
[0147] In the embodiments of the present disclosure, a network system architecture is established by taking a low-orbit satellite in a satellite communication network as an example, but it is not limited to low-orbit satellites. Other forms of NTN systems may also be applicable to the technical solutions of the embodiments of the present disclosure.
[0148] Here, Terahertz (THz) is one of the units of wave frequency, which is equal to 1,000,000,000,000 Hz. Terahertz waves are electromagnetic waves with a frequency of 0.1 to 10 Thz and a wavelength of 3000 to 30 um. They overlap with millimeter waves in the long wave band and with infrared light in the short wave band. The signal-to-noise ratio of the time domain spectrum of Terahertz is very high, and the instantaneous bandwidth is very wide, so it is very conducive to high-speed communication.
[0149] In the disclosed embodiment, the network system is constructed by combining the NTN system with the terahertz base station, so that the system can achieve full network coverage while meeting the high-speed and low-latency data transmission requirements of VR, AR, etc. This facilitates the expansion of the application areas of VR, AR and other services, for example, VR games that are globally networked can be realized.
[0150] In some embodiments, the above-mentioned base station is: a terahertz femtocell base station; the user terminal at least includes: a terminal supporting virtual reality VR services.
[0151] In some embodiments, the above-mentioned NTN network is used to transmit VR data with the gateway via Ka-band transmission; the gateway is used to transmit VR data to the base station via an optical fiber link; and the base station is used to transmit VR data to the user terminal via the terahertz frequency band.
[0152] In some embodiments, the above-mentioned NTN includes: a low-orbit satellite LEO.
[0153] like Figure 4 In the network architecture shown, NTN 110 may include one or more LEO satellites, and gateway 120 is a satellite gateway, which can provide network connections for multiple terahertz home base stations and can provide services for multiple virtual reality users (VRUser).
[0154] In the disclosed embodiment, wireless transmission may include two stages, namely, an outdoor stage and an indoor stage. In the outdoor transmission stage, the LEO satellite transmits VR content to the gateway via the Ka band, and the gateway transmits the received information to the terahertz home base station via an optical fiber link. In the indoor transmission stage, the flying cell can provide VR users with a high-speed link to transmit data via the terahertz frequency band.
[0155] Here, Ka-band is a part of the microwave band of the electromagnetic spectrum, and the frequency range of Ka-band is 26.5 to 40 GHz. Ka stands for Higher than K-band, and Ka-band is also called 30 / 20 GHz band, which can be used for satellite communications.
[0156] The present disclosure provides the following examples:
[0157] Virtual Reality (VR) services enable users to enjoy and interact with an immersive environment from a first-person perspective. VR technology has broad application scenarios in education, medical care, industrial engineering, and entertainment.
[0158] Although 5G networks have been widely deployed, they are still insufficient to support full coverage, low latency, and high-speed services due to their limited coverage and capacity.
[0159] Through the deep integration of satellite and terrestrial communication networks, the NTN network will become an indispensable part of future mobile communication systems, providing a path to achieve true global connectivity and bridging the coverage gap.
[0160] In addition, terahertz and millimeter wave communications can be applied to VR technology. Compared with wifi systems, terahertz communication can provide better reliability and latency performance for VR services.
[0161] Therefore, the embodiments of the present disclosure provide a system model that combines satellite networks with ground networks and uses terahertz and millimeter wave communications to achieve full coverage of VR applications. In addition, the embodiments of the present disclosure provide a time delay minimization solution under this model.
[0162] like Figure 5 As shown, the system model includes a low-orbit satellite 501 (LEO), a satellite gateway 502 (Gateway), a plurality of terahertz home base stations, namely, femtocells 503 (femtocells), and a plurality of VR users 504 (VR users).
[0163] In such Figure 5Under the system model shown, wireless transmission can include two stages, namely the outdoor stage and the indoor stage. In the outdoor transmission stage, the LEO satellite transmits VR content to the gateway via the Ka band, and the gateway forwards the received information to the terahertz home base station via an optical fiber link. In the indoor transmission stage, the flying cell can provide VR users with a high-speed link to transmit data through the terahertz frequency band.
[0164] The disclosed embodiment also provides a signal processing method based on the above system model, which is a network resource allocation method that minimizes time delay.
[0165] The above method is applied to a learning agent, which can also be called an intelligent agent of artificial intelligence, that is, a subject that executes the method. The learning agent can be set in the satellite gateway in the above system model, and the satellite gateway is used to allocate network resources.
[0166] As mentioned above Figure 2 As shown, the method represents the interaction between the learning agent and the satellite-assisted wireless network in the Q-learning environment. The method includes the following steps:
[0167] 1. Given a decision set, which includes various actions and states of the satellite gateway for network allocation, the decision set can be expressed as A[si,ai].
[0168] 2. Set up a Q table for the decision set, including multiple Q values and initialize them. It can be understood that the Q value in the initialized Q table represents the initial value of the expected reward when each state corresponds to each action, and the subsequent learning agent updates the Q table through perceptual learning.
[0169] 2. The learning agent perceives the state s from the network environment i , including channel status and network resource status.
[0170] 3. Based on the status of the current network environment i , the learning agent selects action a in the above decision set A[n] i This action is to allocate resources to VR users, including providing time slots (TS) and mobile base stations (femtocells).
[0171] 4. The network environment enters the next state s i+1 , the learning agent obtains the above action a i The reward r i , that is, feedback corresponding to the goal of reducing latency, for example, it can be the difference between the current latency and the target latency, or a score corresponding to the current latency, etc.
[0172] 5. Based on the current reward r i , determine the expected reward Q(s) corresponding to the current action i ,a i ). The determination formula is the above formula (1):
[0173]
[0174] Here, Q(s i ,a i ) means in state s i Take action in i The expected reward, r i It represents the actual reward after the action is executed, α represents the learning rate, and γ· represents the attenuation coefficient (or discount factor).
[0175] The Q table is continuously updated through the above formula.
[0176] 6. Based on the updated Q table, the satellite gateway performs corresponding network resource allocation according to the action corresponding to the optimal Q value in each state, thereby reducing network delay and improving the performance of the above system model.
[0177] In addition, in the embodiment of the present disclosure, in order to prevent the above process from falling into a local optimal solution, the satellite gateway can also use an ∈-greedy strategy (greedy strategy) to select the action to be executed, such as the above formula (2):
[0178]
[0179] Among them, RandomAction represents a random action, ψ represents a random number, and ε represents the probability of action exploration.
[0180] That is to say, when the proxy gateway selects an action to be executed, if the generated random number ψ is greater than or equal to the preset action exploration probability ε, a random action is selected for execution; if the generated random number ψ is less than the preset action exploration probability ε, the action with the largest Q value obtained based on the above formula (1) is selected.
[0181] In this way, we can avoid falling into the local optimal solution and choose the relatively optimal network allocation method as much as possible.
[0182] Figure 6 is a block diagram of an information processing device 900 according to an exemplary embodiment. For example, the device 900 may be a mobile phone, a mobile computer, etc.
[0183] Reference Figure 6, the device 900 may include one or more of the following components: a processing component 902 , a memory 904 , a power component 906 , a multimedia component 908 , an audio component 910 , an input / output (I / O) interface 912 , a sensor component 914 , and a communication component 916 .
[0184] The processing component 902 generally controls the overall operation of the device 900, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 902 may include one or more modules to facilitate interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate interaction between the multimedia component 908 and the processing component 902.
[0185] The memory 904 is configured to store various types of data to support operations on the device 900. Examples of such data include instructions for any application or method operating on the device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0186] The power supply component 906 provides power to the various components of the device 900. The power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 900.
[0187] The multimedia component 908 includes a screen that provides an output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0188] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), and when the device 900 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 904 or sent via the communication component 916. In some embodiments, the audio component 910 also includes a speaker for outputting audio signals.
[0189] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0190] The sensor assembly 914 includes one or more sensors for providing various aspects of status assessment for the device 900. For example, the sensor assembly 914 can detect the open / closed state of the device 900, the relative positioning of components, such as the display and keypad of the device 900, and the sensor assembly 914 can also detect the position change of the device 900 or a component of the device 900, the presence or absence of user contact with the device 900, the orientation or acceleration / deceleration of the device 900, and the temperature change of the device 900. The sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 914 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 914 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0191] The communication component 916 is configured to facilitate wired or wireless communication between the device 900 and other devices. The device 900 can access a wireless network based on a communication standard, such as Wi-Fi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0192] In an exemplary embodiment, the apparatus 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0193] In an exemplary embodiment, a terminal is also provided, which may include the modules or components in the above-mentioned information processing device, the terminal also includes a first security chip, and the terminal may also include a second security chip. Exemplarily, the terminal may be a smart phone, a tablet computer, a computer, a smart wearable device, and a vehicle-mounted device.
[0194] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the instructions can be executed by the processor 920 of the device 900 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0195] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an information processing device, enables the information processing device to perform an information processing method, the method is applied to a gateway of NTN, the gateway is connected to NTN via a satellite signal, the method comprising:
[0196] Initialize the expected reward table of network status and network resource allocation method;
[0197] Determine the current network status;
[0198] Allocating network resources based on one of the network resource allocation modes in the network state;
[0199] Obtaining reward information obtained from allocating the network resources;
[0200] According to the reward information, updating the expected reward table;
[0201] The network resource allocation method is determined according to the updated expected reward table.
[0202] Figure 7 1900 is a hardware structure block diagram of an information processing device according to an exemplary embodiment. For example, the device 1900 may be provided as a server. Figure 7The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform an information processing method, the method being applied to a gateway of an NTN, the gateway being connected to the NTN via a satellite signal, the method comprising:
[0203] Initialize the expected reward table of network status and network resource allocation method;
[0204] Determine the current network status;
[0205] Allocating network resources based on one of the network resource allocation modes in the network state;
[0206] Obtaining reward information obtained from allocating the network resources;
[0207] According to the reward information, updating the expected reward table;
[0208] The network resource allocation method is determined according to the updated expected reward table.
[0209] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output (I / O) interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or the like.
[0210] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present disclosure, the size of the serial number of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure are for description only and do not represent the advantages and disadvantages of the embodiments.
[0211] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0212] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0213] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0214] In addition, all functional units in the embodiments of the present disclosure may be integrated into one processing unit, or each unit may be separately configured as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0215] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0216] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An information processing method, characterized in that: The method is applied to a gateway of a non-terrestrial network NTN, wherein the gateway is connected to the NTN via a satellite signal; the method comprises: Initialize the expected reward table of network status and network resource allocation method; Determine the current network status; Allocating network resources based on one of the network resource allocation modes in the network state; Obtaining reward information obtained from allocating the network resources; According to the reward information, updating the expected reward table; The network resource allocation method is determined according to the updated expected reward table.
2. The method according to claim 1, characterized in that: The updating of the expected reward table according to the return information includes: Determine, based on the reward information, an actual reward for executing the network resource allocation method in the current state; Determine each of the expected rewards after update according to the current value of each expected reward, the actual reward obtained in the current state, and the maximum expected reward value of the expected reward table.
3. The method according to claim 1, characterized in that The expected reward table includes: The expected reward for each network resource allocation method corresponding to each of the network states.
4. The method according to claim 1, characterized in that: The network resource allocation includes: Providing time slots TS to user terminals; Allocate a base station for the user terminal; wherein the base station includes: a terahertz femtocell base station.
5. The method according to claim 1, characterized in that Each expected reward in the expected reward table is a parameter related to the delay; Among them, the shorter the delay is, the greater the value of the expected reward is.
6. The method according to any one of claims 1 to 5, characterized in that: Determining the network resource allocation method according to the updated expected reward table includes: Generate random numbers; When the random number is greater than or equal to a preset probability value, executing any network resource allocation method; When the random number is less than the probability value, the network resource allocation method corresponding to the maximum reward value in the current environment state in the expected reward table is executed.
7. An information processing device, characterized in that: The device is applied to a gateway of NTN, and the gateway is connected to the NTN via a satellite signal; the device comprises: An initialization module, configured to initialize the network state and the expected reward table of the network resource allocation method; A first determination module, configured to determine a current network state; An execution module, configured to perform network resource allocation based on one of the network resource allocation modes under the network state; An acquisition module, configured to acquire reward information obtained by performing the network resource allocation; An updating module, configured to update the expected reward table according to the reward information; The second determination module is configured to determine a network resource allocation method according to the updated expected reward table.
8. A network system, characterized in that: include: Non-terrestrial network NTN; A gateway connected to the NTN via a satellite signal; The gateway is used to execute the method according to any one of claims 1 to 8; A plurality of base stations are connected to the gateway via an optical fiber network; wherein the base stations are used to provide terahertz network connections for a plurality of user terminals.
9. An information processing device, characterized in that: include: processor; A memory for storing processor executable instructions; wherein the processor is configured to: implement the steps of any one of the information processing methods in claims 1 to 6 when executed.
10. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an information processing device, the device is enabled to perform the steps of any one of the information processing methods of claims 1 to 6.
Citation Information
Cited By
Video low-delay transmission method and device, storage medium, computer program product and audio and video transmission equipment
CN121151596A