An access network selection method and apparatus

By using a reinforcement learning model to select the best access network for terminal devices in an integrated air-space-ground network, the performance degradation caused by users selecting an unsuitable network is solved, and efficient adaptive network selection and service quality improvement are achieved.

CN115396915BActive Publication Date: 2026-05-29HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-05-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In an integrated air-space-ground network, users choosing an inappropriate access network can affect business data transmission, leading to a decline in network performance and service quality. Existing methods are difficult to update and adjust in a timely manner in dynamic network environments.

Method used

By employing a reinforcement learning model combined with environmental state information, and using action neural networks and evaluation neural networks, the optimal access network is selected for multiple terminal devices, thereby optimizing the total throughput and dynamically adapting to network changes.

Benefits of technology

It enables efficient and adaptive access network selection in an integrated air-space-ground network, improving overall network performance and service quality, reducing costs, and increasing learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115396915B_ABST
    Figure CN115396915B_ABST
Patent Text Reader

Abstract

The application discloses an access network selection method and device, relates to the technical field of communication, and is used for selecting an access network for a terminal device. The specific scheme is as follows: acquiring environment state information; wherein the environment state information comprises position information of a plurality of terminal devices, first access network model parameters, and second access network model parameters; the plurality of terminal devices are located in coverage areas of a first access network and a second access network, and the first access network and the second access network support different access technologies; obtaining first decision information according to the environment state information and a reinforcement learning model; wherein an optimization target of the reinforcement learning model is total throughput of the plurality of terminal devices, input parameters of the reinforcement learning model comprise the environment state information, and output parameters of the reinforcement learning model comprise access networks selected for the plurality of terminal devices; and wherein the first decision information is used for indicating the access networks selected for the plurality of terminal devices from the first access network and the second access network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method and apparatus for selecting an access network. Background Technology

[0002] With the rapid development of 5G technology, data traffic has exploded. Traditional terrestrial cellular networks are facing severe challenges such as capacity and coverage. Wireless networks based on different access technology standards have long coexisted, forming heterogeneous wireless networks that are mutually integrated. Among them, the air-space-ground integrated network is a typical result of the integration of heterogeneous wireless networks.

[0003] The integrated air-space-ground network includes terrestrial, airborne, and space-based networks. The terrestrial network mainly consists of terrestrial internet and cellular access networks, and can cover most densely populated areas on the ground. The space-based network mainly consists of high-orbit and low-orbit satellites, while the airborne network mainly consists of unmanned aerial vehicles (UAVs) and high-altitude communication platforms. The space-based and airborne networks can serve as extensions of the terrestrial network, resulting in a wider network coverage.

[0004] In an integrated air-space-ground network architecture, some densely populated areas may simultaneously have multiple access networks, such as satellite access networks and cellular access networks. These different access networks differ in network capacity, communication resources, access protocols, and application scenarios, thus corresponding to varying network performance and quality of service (QoS). If a user selects an access network from among these options that does not meet the network performance requirements for their service transmission, it will affect the user's data transmission. Therefore, selecting a suitable access network for a user becomes a crucial problem to solve. Summary of the Invention

[0005] This application provides an access network selection method and apparatus that can efficiently and adaptively select access networks for users in an integrated air-space-ground network, thereby effectively improving overall network performance and service quality.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, an access network selection device is provided, the method comprising: acquiring environmental state information; wherein the environmental state information includes location information of multiple terminal devices, parameters of a first access network model, and parameters of a second access network model; the multiple terminal devices are located within the coverage areas of the first access network and the second access network, and the first access network and the second access network support different access technologies; obtaining first decision information based on the environmental state information and a reinforcement learning model; wherein the optimization objective of the reinforcement learning model is the total throughput of the multiple terminal devices, the input parameters of the reinforcement learning model include the environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices; wherein the first decision information is used to indicate the access network selected for the multiple terminal devices from the first access network and the second access network.

[0008] Based on the method described in the first aspect, environmental state information can be obtained through real-time interaction with the environment. Based on the obtained environmental state information and reinforcement learning model, the access network selected for multiple terminal devices can be determined, thereby achieving efficient and adaptive network access selection and effectively improving the overall network performance and service quality of the access network.

[0009] In one possible design, the method further includes: configuring training parameters of a reinforcement learning model; wherein the reinforcement learning model includes an action neural network and an evaluation neural network, the action neural network being used to decide on the access network selected by multiple terminal devices, and the evaluation neural network being used to evaluate the access network selected by the multiple terminal devices; obtaining second decision information based on environmental state information and the action neural network in a first decision time slot, and executing the second decision information to obtain the total throughput of the multiple terminal devices; wherein the second decision information is used to indicate the access network initially selected for the multiple terminal devices from the first access network and the second access network; determining the reward value of the terminal device when executing the second decision information based on the total throughput of the multiple terminal devices; determining a time difference error based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network; wherein the time difference error is used to adjust the training parameters of the reinforcement learning model; determining the loss function of the neural network based on the time difference error, and updating the training parameters of the reinforcement learning model based on the loss function until the reward value converges.

[0010] Based on this possible design, training reinforcement learning models by building an integrated air-space-ground network simulation platform can effectively reduce costs and improve learning efficiency.

[0011] In one possible design, the method further includes: calculating a first rate based on first access network model parameters and location information of the terminal device, and calculating a second rate based on second access network model parameters and location information of the terminal device; wherein the first rate is the data rate when the terminal device accesses the first access network, and the second rate is the data rate when the terminal device accesses the second access network.

[0012] Based on this possible design, the data rate of multiple terminal devices after selecting to access the network under the current decision information can be obtained.

[0013] In one possible design, obtaining the total throughput of multiple terminal devices by executing the second decision information includes: determining the total throughput of multiple terminal devices based on a first rate, a second rate, a first decision variable, and a second decision variable, wherein the total throughput of multiple terminal devices is the sum of the data rates of the multiple terminal devices when they access a first access network or a second access network; wherein the first decision variable and the second decision variable can be obtained based on the second decision information; the first decision variable is used to represent the access status of the terminal devices to the first access network, and the second decision variable is used to represent the access status of the terminal devices to the second access network.

[0014] Based on this possible design, the total network throughput can be determined, and network performance and quality of service can be further optimized.

[0015] In one possible design, the second decision time slot is the next decision time slot after the first decision time slot; obtaining environmental state information includes: at the beginning of each decision time slot, obtaining environmental state information under the decision time slot.

[0016] Based on this possible design, reinforcement learning models can dynamically adapt to network changes.

[0017] Secondly, an access network selection device is provided, the device comprising: an acquisition unit for acquiring environmental state information; wherein the environmental state information includes location information of multiple terminal devices, parameters of a first access network model, and parameters of a second access network model; the multiple terminal devices are located within the coverage areas of the first access network and the second access network, and the first access network and the second access network support different access technologies; and a training unit for obtaining first decision information based on the environmental state information and a reinforcement learning model; wherein the optimization objective of the reinforcement learning model is the total throughput of the multiple terminal devices, the input parameters of the reinforcement learning model include the environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices; wherein the first decision information is used to indicate the access network selected for the multiple terminal devices from the first access network and the second access network.

[0018] In one possible design, the acquisition unit is further configured to configure the training parameters of the reinforcement learning model; wherein the reinforcement learning model includes an action neural network and an evaluation neural network, the action neural network being used to decide the access network selected by multiple terminal devices, and the evaluation neural network being used to evaluate the access network selected by multiple terminal devices; the training unit is further configured to obtain second decision information based on environmental state information and the action neural network in a first decision time slot, and execute the second decision information to obtain the total throughput of multiple terminal devices; wherein the second decision information is used to indicate the access network initially selected for multiple terminal devices from the first access network and the second access network; the training unit is further configured to determine the reward value of the terminal device when executing the second decision information based on the total throughput of multiple terminal devices; the training unit is further configured to determine the time difference error based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network; wherein the time difference error is used to adjust the training parameters of the reinforcement learning model; the training unit is further configured to determine the loss function of the reinforcement learning model based on the time difference error, and update the training parameters of the reinforcement learning model based on the loss function until the reward value converges.

[0019] In one possible design, the training unit is further configured to calculate a first rate based on the parameters of the first access network model and the location information of the terminal device, and to calculate a second rate based on the parameters of the second access network model and the location information of the terminal device; wherein the first rate is the data rate when the terminal device accesses the first access network, and the second rate is the data rate when the terminal device accesses the second access network.

[0020] In one possible design, the training unit is further configured to determine the total throughput of multiple terminal devices based on a first rate, a second rate, a first decision variable, and a second decision variable. The total throughput of the multiple terminal devices is the sum of the data rates of the multiple terminal devices when they access the first access network or the second access network. The first decision variable and the second decision variable can be obtained based on the second decision information. The first decision variable is used to represent the access status of the terminal device to the first access network, and the second decision variable is used to represent the access status of the terminal device to the second access network.

[0021] In one possible design, the second decision time slot is the next decision time slot after the first decision time slot; the acquisition unit is used to acquire environmental state information, including: at the beginning of each decision time slot, the acquisition unit acquires the location information of the terminal device under the decision time slot.

[0022] The technical effects of the second aspect or any possible design of the second aspect can be found in the first aspect or any possible design of the first aspect mentioned above, and will not be repeated here.

[0023] Thirdly, an access network selection device is provided, the access network selection device including one or more processors and a communication interface, the one or more processors and the communication interface being used to support the access network selection device in performing the access network selection method as described in the first aspect or any possible design of the first aspect.

[0024] Fourthly, a computer-readable storage medium is provided, which may be a readable non-volatile storage medium storing instructions that, when executed on a computer, cause the computer to perform the access network selection method described in the first aspect or any possible design of the first aspect.

[0025] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the access network selection method described in the first aspect or any possible design of the first aspect.

[0026] In a sixth aspect, a communication system is provided, the communication system including an access network selection device as described in the second aspect or any possible design of the second aspect, capable of performing an access network selection method as described in the first aspect or any possible design of the first aspect.

[0027] The technical effects of any of the design methods in aspects three through six can be found in the first aspect or any possible design of the first aspect, and will not be repeated here. Attached Figure Description

[0028] Figure 1 A schematic diagram of an integrated air-space-ground network architecture;

[0029] Figure 2 This is a schematic diagram of an access network selection system provided in an embodiment of this application;

[0030] Figure 3 A schematic diagram of an air-space-ground integrated vehicle network system provided in an embodiment of this application;

[0031] Figure 4 This is a schematic diagram of the structure of an access network selection device provided in an embodiment of this application;

[0032] Figure 5 A flowchart illustrating an access network selection method provided in this application embodiment;

[0033] Figure 6 A flowchart illustrating a reinforcement learning model training method provided in this application embodiment;

[0034] Figure 7An algorithm block diagram of a reinforcement learning model training method provided in an embodiment of this application;

[0035] Figure 8 This is a performance comparison diagram of an access network selection method and an optimal search method provided in an embodiment of this application;

[0036] Figure 9 This is a schematic diagram of an access network selection device provided in an embodiment of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will now be described in further detail with reference to the accompanying drawings.

[0038] The terms "first," "second," etc., used below are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0039] Furthermore, in this application, directional terms such as "upper" and "lower" are defined relative to the orientation of the components shown in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the orientation of the components in the accompanying drawings.

[0040] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0041] With the rapid development of communication technology, more and more wireless access technologies have emerged. However, currently, no single access network can fully meet users' requirements for wide coverage, high reliability, high bandwidth, and low latency. Different access networks have their own advantages in terms of network capacity, coverage, bandwidth, mobility support, and latency, and their application scenarios also differ. Therefore, it is difficult for various access networks to replace each other, and using only a single network cannot meet the diverse needs of broadband services. Therefore, the future trend of mobile communication development is the convergence and interconnection of existing and emerging wireless access networks, while fully leveraging the advantages of heterogeneous wireless access networks to effectively integrate and utilize various network resources, providing users with any form of ubiquitous service with QoS guarantees. For example, an integrated air-space-ground network can fully utilize access networks in different spatial dimensions to achieve more comprehensive network coverage for users within a region.

[0042] An integrated air-space-ground network is a type of network that breaks through the limitations of terrestrial networks and achieves seamless coverage of terrestrial, satellite, and airborne networks. Figure 1 This is a schematic diagram of an integrated air-space-ground network architecture, such as... Figure 1 As shown, the integrated air-space-ground network includes terrestrial, airborne, and space-based networks. The terrestrial network mainly consists of three parts: cellular wireless networks, satellite ground stations and mobile satellite terminals, and ground-based data and processing centers. Due to different application scenarios, it can also be divided into cellular access networks, the Internet of Things (IoT), terrestrial interconnection networks, maritime networks, emergency communication networks, and military operational networks. The airborne network mainly consists of various aircraft such as drones, civil airliners, and airships. It can include drone self-organizing networks, edge service drones, high-altitude communication platforms, and cruising drones. The space-based network mainly consists of various orbital satellites, such as high-orbit satellites and medium- and low-orbit satellites. In some special circumstances, such as when ground equipment is damaged and cannot be used normally due to natural disasters, space-based or airborne networks can be used to replace terrestrial networks to ensure people's daily communication.

[0043] In the integrated air-space-ground network architecture, the network resources and access protocols of the terrestrial, airborne, and space-based networks are heterogeneous. In some densely populated areas, there may be overlapping coverage of the terrestrial, airborne, and space-based networks. In this case, the choice of different access networks by users will have different impacts on the overall network performance and service quality of the integrated air-space-ground network. Therefore, the issue of user access network selection in the integrated air-space-ground network is of great research significance.

[0044] In one possible design, users select the access network within the integrated air-space-ground network based on the maximum received signal strength (max-RSS). For example, in densely populated areas with overlapping access network coverage, when users are forced to switch access networks due to location changes, they tend to choose high-power base stations for network access. This significantly reduces the number of users accessing the network from low-power base stations, leading to congestion of services at high-power base stations and wasted communication resources at low-power base stations, thus reducing network resource utilization and user experience.

[0045] In another possible design, mathematical models such as combinatorial optimization and game theory are used to find the optimal solution for access network selection in heterogeneous wireless networks. This optimal search method requires obtaining complete network environment information in advance, but it cannot update and adjust the network environment information in a timely manner when the network environment changes dynamically, and the signaling interaction cost is high.

[0046] To achieve efficient and adaptive access network selection for users in an integrated air-space-ground network, thereby effectively improving overall network performance and service quality, this application provides an access network selection method. The method includes: acquiring environmental state information; wherein the environmental state information includes location information of multiple terminal devices, parameters of a first access network model, and parameters of a second access network model; the multiple terminal devices are located within the coverage areas of the first and second access networks, and the first and second access networks support different access technologies; obtaining first decision information based on the environmental state information and a reinforcement learning model; wherein the optimization objective of the reinforcement learning model is the total throughput of the multiple terminal devices, the input parameters of the reinforcement learning model include the environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices; wherein the first decision information is used to indicate the access network selected for the multiple terminal devices from the first and second access networks.

[0047] For example Figure 2 This is a schematic diagram of an access network selection system provided in an embodiment of this application, such as... Figure 2 As shown, the reinforcement learning model is initially trained through the integrated air-space-ground network simulation platform. The trained reinforcement learning model is then applied to the actual integrated air-space-ground network. After obtaining environmental state information from the actual integrated air-space-ground network, the first decision information is obtained based on the reinforcement learning model and executed. Subsequently, the reinforcement learning model is updated, and the integrated air-space-ground network simulation platform is optimized. Finally, a closed-loop optimization system is formed to adapt to the ever-changing network environment and terminal device access requirements.

[0048] The access network selection method provided in the embodiments of this application will now be described with reference to the accompanying drawings.

[0049] The access network selection method provided in this application can be applied to various communication systems, such as: Long Term Evolution (LTE) systems, 5th Generation (5G) mobile communication systems, Wireless Fidelity (Wi-Fi) systems, future communication systems, or systems integrating multiple communication systems, etc., and this application does not limit the application. 5G can also be referred to as New Radio (NR).

[0050] The access network selection method provided in this application embodiment can be applied to various communication scenarios, such as one or more of the following communication scenarios: enhanced mobile broadband (eMBB), ultra-reliable low latency communication (URLLC), machine type communication (MTC), massive machine type communication (mMTC), device to device (D2D), vehicle to everything (V2X), vehicle to vehicle (V2V), and Internet of Things (IoT).

[0051] Figure 3 This application provides a schematic diagram of an air-space-ground integrated vehicle network system, including multiple terminal devices, network access sites, and a network server. Figure 3 The terminal equipment includes vehicle terminals, network access sites include cellular base stations and low-Earth orbit satellites, and network servers may include electronic map servers, which can be deployed in laboratories. For example... Figure 3 As shown, within this selected area, network coverage is provided by both a low-Earth orbit (LEO) satellite constellation and cellular base stations, offering access services to vehicle terminals within this area. The LEO satellite constellation access network sites include LEO satellites, and the cellular base station access network sites include: Cellular Base Station 1, Cellular Base Station 2, and Cellular Base Station 3. The terminal devices covered by the access network in this selected area include: Vehicle 1, Vehicle 2, Vehicle 3, Vehicle 4, Vehicle 5, and Vehicle 6. Vehicles 1 and 3 can choose Cellular Base Station 1 for access, Vehicles 2 and 5 can choose LEO satellites for access, Vehicle 4 can choose Cellular Base Station 2 for access, and Vehicle 6 can choose Cellular Base Station 3 for access.

[0052] In this application embodiment, the terminal device, except for Figure 3 Besides the vehicle terminal, the device can also be a user equipment (UE), a mobile station (MS), or a mobile terminal (MT). Specifically, the terminal can be a mobile phone, a tablet computer, or a computer with wireless transceiver capabilities. It can also be a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in autonomous driving, a wireless terminal in telemedicine, a wireless terminal in a smart grid, a wireless terminal in a smart city, a smart home, etc. In the embodiments of this application, the device used to implement the terminal device function can be a vehicle terminal or a device capable of supporting the terminal device to implement the function, such as a chip system. This device can be installed in the terminal device or used in conjunction with the terminal device.

[0053] In this application embodiment, the access network site is except Figure 3 In addition to the cellular base stations and low-Earth orbit satellites shown, the network may also include any of the following nodes: radio access point, transmission receive point (TRP), transmission point (TP), evolved Node B (gNB), transmission reception point (TRP), evolved Node B (eNB), radio network controller (RNC), Node B (NB), base station controller (BSC), base transceiver station (BTS), home base station (e.g., home evolved Node B, or home Node B, HNB), base band unit (BBU), or wireless fidelity (Wi-Fi) access point (AP), satellite ground station, mobile satellite terminal, and some other type of access node. In the embodiments of this application, the device used to implement the function of the access network site can be a cellular base station or a low-Earth orbit satellite; it can also be a device capable of supporting the access network site to implement this function, such as a chip system, which can be installed in the access network site or used in conjunction with the access network site.

[0054] It should be noted that the above Figure 3 The nodes and their connections in the air-space-ground integrated vehicle network system shown are merely an example. In a specific implementation, the nodes and their connections can take other forms, and this application does not impose any specific limitations on them. Furthermore, Figure 3 This is just an example model framework diagram. Figure 3 The number of nodes included and the access method of terminal devices are unrestricted. Except... Figure 3 In addition to the functional nodes shown, other nodes may be included without restriction.

[0055] In practical implementation, Figure 2 The integrated air-space-ground network simulation platform shown can be built in an access network selection device, which can adopt... Figure 4 The shown composition or includes Figure 4 The components shown. Figure 4 This is a schematic diagram of the access network selection device 400 provided in the embodiments of this application. When the device 400 has the function of the integrated air-space-ground network simulation platform described in the embodiments of this application, the access network selection device 400 can be a chip or system-on-a-chip in the integrated air-space-ground network simulation platform.

[0056] like Figure 4 As shown, the access network selection device 400 may include a processor 401, a communication line 402, and a communication interface 403. Furthermore, the access network selection device 400 may also include a memory 403. The processor 401, memory 402, and communication interface 403 can be connected via the communication line 402.

[0057] The processor 401 can be a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 401 can also be other devices with processing functions, such as circuits, devices, or software modules. The network access selection methods provided in the following embodiments of this application can be implemented by running the computer programs, software code, or instructions therein, or by calling the computer programs, software code, or instructions stored in memory 404.

[0058] Communication line 402 is used to transmit information between the components included in the access network selection device 400.

[0059] Communication interface 403 is used for communication with other devices or other communication networks. These other communication networks can be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Communication interface 403 can be a radio frequency module, transceiver, or any device capable of communication.

[0060] Memory 404 is used to store instructions. These instructions can be computer programs.

[0061] The memory 404 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disk storage, magnetic disk storage media or other magnetic storage devices. Optical disk storage includes compressed optical discs, laser discs, universal digital discs, Blu-ray discs, etc.

[0062] It should be noted that the memory 404 can exist independently of the processor 401 or can be integrated with the processor 401. The memory 404 can be used to store instructions, program code, or some data, etc. The memory 404 can be located inside or outside the access network selection device 400, without limitation. The processor 401 is used to execute the instructions stored in the memory 404 to implement the access network selection method provided in the following embodiments of this application.

[0063] In one example, processor 401 may include one or more CPUs, for example Figure 4 CPU0 and CPU1 in the example. As one exemplary implementation, the access network selection device 400 includes multiple processors, for example, besides... Figure 4 In addition to processor 401, it may also include processor 407.

[0064] As an example implementation, the network access selection device 400 may also include an output device 405 and an input device 406. The input device 406 may be a keyboard, mouse, microphone, or joystick, etc., and the output device 405 may be a display screen, speaker, etc.

[0065] It should be noted that the network access selection device 400 can be a desktop computer, laptop computer, network server, mobile phone, tablet computer, wireless terminal, embedded device, chip system, or other device. Figure 4 Equipment with a similar structure. Furthermore... Figure 4 The structural composition shown does not constitute a limitation on the access network selection device, except... Figure 4 In addition to the components shown, the network access selection device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0066] The following is combined with Figure 3 The illustrated air-space-ground integrated vehicle network system model describes the access network selection method provided in this application embodiment. The access network selection devices in the following embodiments may have… Figure 4 The components are shown. The actions, terminology, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples; other names may be used in specific implementations without limitation.

[0067] Figure 5 This is a flowchart illustrating an access network selection method provided in an embodiment of this application. The method can be... Figure 4 The network access selection device shown performs the operation, such as... Figure 5 As shown, the method may include:

[0068] Step 501: Obtain environmental status information.

[0069] The environmental status information includes the location information of multiple terminal devices, parameters of the first access network model, and parameters of the second access network model; the multiple terminal devices are located in the coverage areas of the first access network and the second access network, and the first access network and the second access network support different access technologies.

[0070] Optionally, environmental status information can be obtained directly from the network prior data set by the network access selection device, or it can be obtained by the network access selection device through real-time interaction with the actual integrated air-space-ground network environment. The network prior data set may include trajectory information of multiple vehicle terminals within the network coverage area, and the network access selection device can obtain this information from data deployed in the actual air-space-ground integrated network environment. Figure 3 Obtain the network prior dataset from the network server.

[0071] Optionally, the access network selection device can acquire environmental status information at the beginning of each decision time slot, and the location information of each terminal device is different in different decision time slots. For example, the environmental status information can be centrally acquired by the control center of the access network selection device at the beginning of each decision time slot. The length of the decision time slot can be pre-configured, and in each decision time slot, the access network selection device can select an access network for multiple terminal devices within the current network coverage area.

[0072] For example, the access network selection device can acquire environmental state information at the start time T1 of the first decision time slot, including the location information of multiple terminal devices, the parameters of the first access network model, and the parameters of the second access network model at time T1; and acquire environmental state information at the start time T2 of the second decision time slot, including the location information of multiple terminal devices, the parameters of the first access network model, and the parameters of the second access network model at time T2. In the first decision time slot, the access network selection device can select an access network for multiple terminal devices within the network coverage area; in the second decision time slot, the access network selection device can decide whether multiple terminal devices maintain or switch access networks.

[0073] The access network selection device can pre-number multiple decision time slots, such as time slot 1, time slot 2, ..., time slot N. The first decision time slot can be any one of the multiple decision time slots, and the second decision time slot is the next decision time slot after the first decision time slot.

[0074] Optionally, at the beginning of each decision time slot, the access network selection device can also acquire network topology information, which is used by the access network selection device to establish a network system model.

[0075] Before acquiring environmental status information, the access network selection device can establish an integrated air-space-ground network system model based on the acquired network topology information. This integrated air-space-ground network system model can include a set of terminal devices V = {1, 2, ..., V}, a first set of access network sites M = {1, 2, ..., M}, and a second set of access network sites S = {1, 2, ..., S}. For example... Figure 3 The air-space-ground integrated vehicle network system shown includes a vehicle set V = {vehicle 1, vehicle 2, vehicle 3, ..., vehicle 6}, a cellular base station set M = {cellular base station 1, cellular base station 2, cellular base station 3}, and a low-Earth orbit satellite set S = {low-Earth orbit satellites}.

[0076] The location information of the terminal devices can include latitude and longitude information of multiple terminal devices across multiple decision-making time slots, or the movement trajectory information of the terminal devices. The location information of the terminal devices can be obtained by a network server deployed in the laboratory based on the positioning of the terminal devices, or directly from an electronic map server. For example, in... Figure 3 Within the selected area shown, low-orbit satellites and cellular base stations provide access services for vehicles within the area, and network servers can obtain the location information of vehicles within the area.

[0077] Multiple terminal devices can be covered by multiple access networks. It should be noted that multiple access networks include two or more networks using different access technologies, for example, Figure 3 Multiple vehicle users are located in areas covered by a first access network and a second access network. The first access network is a cellular network, and the second access network is a satellite network.

[0078] The first and second access network model parameters can include one or more channel parameters of the heterogeneous wireless network. Taking a cellular network as the first access network and a satellite network as the second access network as an example, the first access network model parameters can include available bandwidth, transmission power, and path loss index of the cellular base station; the second access network model parameters can include available bandwidth, transmission power, and noise factor of the low-Earth orbit (LEO) satellite. Specific values ​​for the first and second access network model parameters are shown in Table 1. Optionally, the second access network model parameters can also include other LEO satellite orbit parameters, such as orbital altitude, orbital inclination, right ascension of the ascending node, and sensor conus angle. Values ​​for the LEO satellite orbit parameters are shown in Table 2.

[0079] Table 1 Access Network Model Parameters

[0080] Available bandwidth of cellular base stations 200MHz Available bandwidth for low-Earth orbit satellites 18MHz Cellular base station transmission power 100W Low Earth Orbit Satellite Transmission Power 100W Path loss exponent α 3 Noise factor β <![CDATA[10 -19.9 ]]>

[0081] Table 2 Satellite orbital parameters

[0082] orbital altitude 550km track inclination 53° Right ascension of ascending node 80° Sensor cone half angle 44.85°

[0083] Optionally, the access network selection device calculates a first rate based on the first access network model parameters and the location information of the terminal device, and calculates a second rate based on the second access network model parameters and the location information of the terminal device.

[0084] Wherein, the first rate is the data rate when the terminal device accesses the first access network, and the second rate is the data rate when the terminal device accesses the second access network. Figure 3Taking the air-space-ground integrated vehicle network system shown as an example, the first rate can be the data rate when the vehicle accesses the cellular base station, and the second rate can be the data rate when the vehicle accesses the low-orbit satellite. The first rate can be calculated by the following formula (1), and the second rate can be calculated by the following formula (2).

[0085] The data rate when vehicle v accesses cellular base station m is C v,m :

[0086]

[0087] The data rate when vehicle v accesses low-Earth orbit satellite s is C v,s :

[0088]

[0089] Among them, B v,m B represents the average bandwidth of the vehicle v connected to the cellular base station m. v,s This represents the average bandwidth of the connection between vehicle v and low-Earth orbit satellite s. (B) v,m The available bandwidth of the cellular base station in Table 1 can be obtained by dividing the number of vehicles accessing the cellular base station m. v,s This can be obtained by dividing the available bandwidth of the low-Earth orbit satellites in Table 1 by the number of vehicles accessing the low-Earth orbit satellite s. M P represents the transmission power of a cellular base station. S λ represents the transmission power of a low-Earth orbit satellite. M λ represents the wavelength of the cellular base station antenna. S d represents the wavelength of a low-Earth orbit satellite antenna. v,m d represents the distance between the cellular base station m and the vehicle v. v,s The distance between the low-orbit satellite s and the vehicle v is represented by α and β, respectively, which represent the path loss exponent and the noise factor.

[0090] Step 502: Obtain the first decision information based on the environmental state information and the reinforcement learning model.

[0091] The first decision information is used to indicate the access network selected for multiple terminal devices from the first access network and the second access network.

[0092] Optionally, the first decision information may include the network selection status of multiple terminal devices, and the first decision information may be represented in vector form. For example, Figure 3The areas shown include LEO satellite, cellular base station 1, cellular base station 2, and cellular base station 3, which provide access services. These access network sites, which can be selected by terminal devices within the network coverage area, are numbered 0, 1, 2, and 3 respectively. There are currently 6 vehicles within the network coverage area: vehicle 1, vehicle 2, vehicle 3, vehicle 4, vehicle 5, and vehicle 6. Based on the environmental status information obtained from the network server, the access network selection device can obtain first decision information a = (1, 0, 1, 2, 0, 3). This first decision information a indicates that the access network selection device selects cellular base station 1 for vehicle 1 and vehicle 3, LEO satellite for vehicle 2 and vehicle 5, cellular base station 2 for vehicle 4, and cellular base station 3 for vehicle 6.

[0093] The optimization objective of the reinforcement learning model is the total throughput of multiple terminal devices. The input parameters of the reinforcement learning model include environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices.

[0094] Specifically, reinforcement learning is a mathematical framework for autonomous policy learning through experience. It involves an agent learning through trial and error, receiving rewards from interactions with the environment to guide its behavior, with the goal of maximizing the agent's reward value. In this embodiment, the optimization objective of the reinforcement learning model is the total throughput of multiple terminal devices, achieved by... Figure 2 The reinforcement learning model trained on the integrated air-space-ground simulation platform is shown to interact with the actual integrated air-space-ground network environment to obtain reward values, which further guide the network access decisions of multiple terminal devices to obtain the maximum reward value. The reward value is directly proportional to the total throughput of the multiple terminal devices; when the reinforcement learning model obtains the maximum reward value, the total throughput of the multiple terminals also reaches its optimal level.

[0095] based on Figure 5 The method allows the access network selection device to decide the access network to be selected for the terminal device based on the terminal device's location information and reinforcement learning model. This enables efficient and adaptive access network selection for users in an integrated air-space-ground network, thereby effectively improving the overall network performance and service quality of the access network.

[0096] Figure 5 The reinforcement learning model involved can be pre-trained by the access network selection device, and the training process of the reinforcement learning model can be as follows: Figure 6 As shown:

[0097] Step 601: Configure the training parameters of the reinforcement learning model.

[0098] The reinforcement learning model includes an action neural network and an evaluation neural network. The action neural network is used to make decisions about the access network for terminal device selection, and the evaluation neural network is used to evaluate the access network for terminal device selection.

[0099] Optionally, the actor neural network and the critic neural network can be predefined. The actor neural network can decide the access network selected by each terminal device based on the current network state information, and the critic neural network can evaluate the access network selected by the terminal device, thereby influencing the actor neural network to make a better choice in future states.

[0100] Specifically, both actor neural networks and critic neural networks can include three fully connected layers.

[0101] The fully connected layers of the Actor neural network have a structure of 32×16×24. These three fully connected layers include two hidden layers and one output layer. The first hidden layer has 32 neurons, and the second has 16 neurons. The two hidden layers can use the rectified linear unit (ReLU) function as their activation function. The final output layer can be used to output the probability distribution π of each terminal device selecting each access network under a given state. θ (s,a), for example, Figure 3 Of the 6 cars in the model, each car has 4 access options, so the output layer can correspond to 24 neurons. The initial values ​​of the weight matrices of the three fully connected layers of the actor neural network can follow a normal distribution with a mean of 0 and a standard deviation of 0.1, and the initial bias value is 0.1.

[0102] The fully connected layers of the critic neural network have a 32×16×1 structure, consisting of three fully connected layers: two hidden layers and one output layer. The first hidden layer has 32 neurons, and the second has 16 neurons. The ReLU function can be used as the activation function for both hidden layers. The final output layer is used to output the estimated value V of the state-value function. ω (s), corresponding to 1 neuron. The initial values ​​of the weight matrices of the three fully connected layers of the critic neural network can follow a normal distribution with a mean of 0 and a standard deviation of 0.1, and the initial bias value is 0.1.

[0103] Optionally, the training parameters of the reinforcement learning model may include ω and θ, where ω can be used to update the critic neural network and θ can be used to update the actor neural network.

[0104] Optionally, the training parameters of the reinforcement learning model also include hyperparameters. Hyperparameters are parameters with fixed values ​​in the reinforcement learning model and can be pre-configured. In this embodiment, the hyperparameters may include the learning rate α of the actor neural network, the learning rate β of the critic neural network, and the decay coefficient γ, where α can be 10. -5 β can take the value 10. -4 .

[0105] Step 602: In the first decision time slot, obtain the second decision information based on the environmental state information and the action neural network, and execute the second decision information to obtain the total throughput of multiple terminal devices.

[0106] Optionally, in each decision time slot, the access network selection device can input the terminal device location information from the environmental state information into the actor neural network to obtain the second decision information.

[0107] The second decision information can be used to indicate the access network initially selected for the terminal device from the first access network and the second access network, and the second decision information can be a vector.

[0108] Optionally, the access network selection device can obtain second decision information based on the location information of multiple terminals and an actor neural network. For example, the access network is... Figure 3 The diagram shows the initial network selection for the six vehicles. The location information of the six vehicles can be represented as the state space s = [s 1,1 ,s 1,2 ,s 2,1 s 2,2 , ..., s 6,1 s 6,2 ], where s v,1 The longitude value of vehicle v is represented by s. v,2 Let represent the latitude value of vehicle v, where v takes values ​​from 1, 2, ..., 6. Each vehicle has 4 access options (3 cellular base stations and 1 low-Earth orbit satellite), therefore each vehicle user has 4 candidate actions, and the candidate action space for vehicle v is 'a'. v ={0, 1, 2, 3}, the joint action space of the 6 vehicles is a = {a1, a2, a3, a4, a5, a6}.

[0109] The access network selection device can calculate the attenuation exploration factor ε using the following formula (3). t :

[0110] ε t =0.95 - t * 1.32 * 10 -4 Formula (3)

[0111] Where t is the slot number of the first decision slot.

[0112] Furthermore, the access network selection device can randomly generate a floating-point number σ, where the value of σ ranges from [0,1). If σ < ε t The terminal device access selection decision is made according to an equal probability distribution, meaning that each access network site has an equal probability of being selected by each terminal device, thus obtaining the second decision information; if σ≥ε t According to the probability distribution π output by the actor neural network θ (s,a) makes access selection decisions for terminal devices and obtains the second decision information.

[0113] For example, Figure 3 The π in the air-space-ground integrated vehicle network shown θ The probability distribution of (s,a) can be shown in Table 3.

[0114] Table 3π θ probability distribution of (s,a)

[0115] Access probability low-orbit satellites Cellular base station 1 Cellular base station 2 Cellular base station 3 Vehicle 1 <![CDATA[7.0435405×10 -1 ]]> <![CDATA[2.882081×10 -2 ]]> <![CDATA[1.1874273×10 -1 ]]> <![CDATA[1.4808242×10 -1 ]]> Vehicle 2 <![CDATA[9.8443663×10 -1 ]]> <![CDATA[8.4275212×10 -7 ]]> <![CDATA[1.4014948×10 -2 ]]> <![CDATA[1.5475409×10 -3 ]]> Vehicle 3 <![CDATA[6.7944760×10 -3 ]]> <![CDATA[1.9188672×10 -6 ]]> <![CDATA[9.8826665×10 -1 ]]> <![CDATA[4.9369712×10 -3 ]]> Vehicle 4 <![CDATA[9.8854393×10 -1 ]]> <![CDATA[1.0245915×10 -2 ]]> <![CDATA[3.6379950×10 -7 ]]> <![CDATA[1.2097722×10 -3 ]]> Vehicle 5 <![CDATA[3.156804×10 -2 ]]> <![CDATA[4.3488425×10 -1 ]]> <![CDATA[5.198045×10 -1 ]]> <![CDATA[1.374327×10 -2 ]]> Vehicle 6 <![CDATA[1.4258887×10 -7 ]]> <![CDATA[2.0432055×10 -6 ]]> <![CDATA[9.9790949×10 -1 ]]> <![CDATA[2.0882946×10 -3 ]]>

[0116] If σ≥ε t The network access selection device can select the network based on the π value in Table 3. θ The probability distribution (s,a) represents the access selection decision made by the vehicle. The access network selection device can output the second decision information a = (1,0,1,2,0,3). This first decision information a can indicate that the access network selection device selects cellular base station 1 for vehicle 1 and vehicle 3, selects low-orbit satellite for vehicle 2 and vehicle 5, selects cellular base station 2 for vehicle 4, and selects cellular base station 3 for vehicle 6.

[0117] The second decision time slot is the next decision time slot after the first decision time slot. The location information of the terminal device is different in different decision time slots.

[0118] Furthermore, the network selection device can obtain the first decision variable and the second decision variable based on the second decision information. Specifically, the first decision variable represents the access status of the terminal device with the first access network, and the second decision variable represents the access status of the terminal device with the second access network. Both the first and second decision variables are single-bit binary numbers, indicating that each device can only choose either the first or the second access network for access.

[0119] The total throughput of multiple terminal devices can be determined by the access network selection device based on a first rate, a second rate, a first decision variable, and a second decision variable. Specifically, the first rate can be the data rate when a terminal device accesses a first access network site, and the second rate can be the data rate when a terminal device accesses a second access network site. The total throughput of multiple terminal devices can be the sum of the data rates of the multiple terminal devices accessing the first or second access network.

[0120] For example, the network access selection device can calculate based on the first rate, the second rate, the first decision variable, and the second decision variable. Figure 3 The maximum total throughput of the multiple vehicles shown is the optimization objective of the reinforcement learning model. The objective function can be expressed as follows: (4)

[0121] max∑ v∈V,m∈M,s∈S (x v,m C v,m +x v,s C v,s ) (Formula 4)

[0122] Among them, C v,m C represents the first speed. v,s Indicates the second speed, x v,m Let x represent the first decision variable. v,s Let x represent the second decision variable, when vehicle v connects to cellular base station m. v,m The value of x is 1, otherwise x v,m The value of x is 0; when vehicle v is connected to low-orbit satellite s, x v,s The value of x is 1, otherwise x v,s The value of x is 0. v,m and x v,s The constraint condition is: x v,m +x v,s =1,x v,m ,x v,s ∈{0,1}. Under this constraint, a vehicle can choose at most one network to access.

[0123] For example, Figure 3 The air-space-ground integrated vehicle network system shown includes a vehicle set V = {vehicle 1, vehicle 2, vehicle 3, ..., vehicle 6}, a cellular base station set M = {cellular base station 1, cellular base station 2, cellular base station 3}, and a low-Earth orbit satellite set S = {low-Earth orbit satellites}. Figure 3The areas shown in the diagram provide access services via low-Earth orbit satellites, cellular base station 1, cellular base station 2, and cellular base station 3. These access network sites, available for selection by end users within the network coverage area, are numbered sequentially as 0, 1, 2, and 3. Currently, there are 6 vehicles within the network coverage area: vehicle 1, vehicle 2, vehicle 3, vehicle 4, vehicle 5, and vehicle 6. Based on the environmental status information obtained from the network server, the access network selection device obtains second decision information a = (1, 0, 1, 2, 0, 3). This second decision information a indicates that the access network selection device selects cellular base station 1 for vehicles 1 and 3, low-Earth orbit satellites for vehicles 2 and 5, cellular base station 2 for vehicle 4, and cellular base station 3 for vehicle 6. Based on this second decision information, a first decision variable and a second decision variable can be obtained. For example, when vehicle 1 accesses cellular base station 1, the first decision variable x for vehicle 1's access to cellular base station 1 is... 1,1 =1, second decision variable x 1,s =0. When vehicle 5 connects to a low-Earth orbit satellite, the first decision variable for vehicle 5 is x. 5,m =0, second decision variable x 5,1 =1.

[0124] Step 603: Determine the reward value when the terminal device executes the second decision information based on the total throughput of multiple terminal devices.

[0125] In this model, the reward value is directly proportional to the total network throughput.

[0126] Specifically, the access network selection device can obtain the reward value r(s, a) obtained under the selection strategy of executing the second decision information through the following formula (5).

[0127] r(s, a) = μ∑ v∈V,m∈M,s∈S (x v,m C v,m +x v,s C v,s ) (Formula 5)

[0128] Where μ > 0, μ is the proportionality coefficient of the reward value, ∑ v∈V,m∈M,s∈S (x v,m C v,m +x v,s C v,s This represents the total throughput of multiple terminal devices.

[0129] Step 604: Determine the time difference error based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network.

[0130] Temporal difference error is used to adjust the training parameters of the reinforcement learning model.

[0131] Specifically, the critic neural network can obtain the value V in the first decision time slot based on the location information s_ of the terminal device in the first decision time slot and the location information s_ of the terminal device in the second decision time slot. ω (s) and the value V under the second decision time slot ω (s_), calculate the time difference error δ according to the following formula (6).

[0132] δ=(r(s,a)+γ*V ω (s_))-V ω (s) (Formula 6)

[0133] Where r(s, a) is the reward value, γ∈[0,1] is the decay coefficient, and V ω (s) represents the value in the first decision time slot, V ω (s_) represents the value under the second decision slot.

[0134] Step 605: Determine the loss function of the neural network based on the temporal difference error, and update the parameters of the reinforcement learning model based on the loss function until the reward value converges.

[0135] Specifically, the access network selection device can determine the loss function loss_c of the critic neural network based on the time difference error.

[0136] The loss function loss_c of the critic neural network can be obtained by calculating the following formula (7).

[0137] loss_c = δ 2 (Formula 7)

[0138] Furthermore, the access network selection device can update the training parameters ω of the critic neural network using loss_c. Specifically, the access network selection device can use an adaptive moment estimation (ADAM) algorithm as an optimizer to update the training parameters ω of the critic neural network, thereby reducing the value r(s,a)+γ*V obtained by executing the second decision information a at position s. ω (s_) and the value V of staying at position s ω The difference between (s).

[0139] Specifically, the access network selection device can determine the loss function loss_a of the actor neural network based on the time difference error.

[0140] The loss function loss_a of the actor neural network can be obtained by calculating the following formula (8).

[0141] loss_a=-log(π θ (s, a)*δ) (Formula 8)

[0142] Furthermore, the access network selection device can update the training parameters θ of the actor neural network with loss-a, thereby increasing the probability of the second decision information appearing when δ is positive and decreasing the probability of the second decision information appearing when δ is negative.

[0143] Furthermore, the access network selection device repeats steps 601-605 above until the reward value converges.

[0144] Optionally, the access network selection device can pre-configure a maximum number of trajectory steps. The maximum number of trajectory steps can be the number of times the environmental status information of multiple terminal devices is acquired. The access network selection device repeats the above steps 601-605 until the preset maximum number of trajectory steps is reached, at which point it can be determined that the reward value has converged and the reinforcement learning model has completed one iteration of training.

[0145] Optionally, the access network selection device can be pre-configured to have a maximum number of iterations of K for the reinforcement learning model. After the access network selection device completes K training iterations, it can obtain a trained reinforcement learning model. For example, K can be configured to be 100000.

[0146] The algorithm flowchart for the reinforcement learning model training process described in steps 601-605 can be as follows: Figure 7 As shown. Figure 7 As shown, the access network selection device can obtain the location information s of the terminal device from the integrated air-space-ground network simulation platform or the actual integrated air-space-ground network environment, and obtain the second decision information a through the actor neural network. The second decision information can be distributed to the terminal device in the integrated air-space-ground network for decision-making regarding the access network to be selected. By inputting the location information s of the terminal device in the first decision time slot and the location information s_ in the second decision time slot into the critic neural network, the value and reward value of the terminal device in different positions are obtained, and then the time difference error δ is calculated to reduce the loss function δ. 2 The critic neural network is further updated, and the probability of the second decision information a is increased when δ>0 and decreased when δ<0. Thus, after multiple training iterations of the reinforcement learning model, it can determine the optimal access network allocation method for the terminal devices under the integrated air-space-ground network coverage.

[0147] based on Figure 6The reinforcement learning training method shown can effectively reduce costs and improve learning efficiency by building an integrated air-space-ground network simulation platform to train reinforcement learning models.

[0148] For example, comparing the data rate performance of the reinforcement learning-based access network selection method and the optimal search method proposed in this application, the data rate performance comparison results of the two algorithms are as follows: Figure 8 As shown, with the increase of the number of iterations, the network performance based on the reinforcement learning algorithm in this embodiment gradually stabilizes. Although there is a certain performance gap between the reinforcement learning algorithm and the optimal search algorithm due to the complexity of the environment, the computational cost of forward propagation in the neural network of reinforcement learning is extremely small. Compared with the optimal search algorithm, which requires a large amount of computation, the reinforcement learning algorithm has a significant advantage in running time. Table 4 shows the comparison of the average running time of the two algorithms in one decision time slot for different numbers of vehicles. As can be seen from Table 4, when there are 10 vehicles in the network, the running time of the optimal search algorithm is about 100,000 times that of the reinforcement learning algorithm. Therefore, the efficiency of the access network selection method provided in this embodiment is far superior to the traditional optimal search algorithm for selecting access networks.

[0149] Table 4 Comparison of Algorithm Running Time (Unit: ms)

[0150] Number of vehicles 5 6 7 8 9 10 Reinforcement learning methods 3.25 3.42 3.63 3.78 4.24 4.26 Optimal search method 292.34 1436.87 3773.59 13720.94 51829.38 518910.16

[0151] It is understood that each node, such as a terminal device or access network site, includes corresponding hardware structures and / or software modules to perform the aforementioned functions. Those skilled in the art should readily recognize that the methods of this application's embodiments can be implemented in hardware, software, or a combination of hardware and computer software, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein. Whether a function is executed in hardware or computer software-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application's embodiments.

[0152] This application embodiment can divide the access network selection device into functional modules according to the above method example. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0153] Figure 9A structural diagram of an access network selection device is shown, which can be used to perform the functions of the access network selection device involved in the above embodiments. As one possible implementation, Figure 9 The network access selection device shown includes: an acquisition unit 901 and a training unit 902.

[0154] The acquisition unit 901 is used to acquire environmental status information; wherein, the environmental status information includes the location information of multiple terminal devices, parameters of the first access network model, and parameters of the second access network model; the multiple terminal devices are located in the coverage area of ​​the first access network and the second access network, and the access technologies supported by the first access network and the second access network are different.

[0155] The acquisition unit 901 is also used to configure the training parameters of the reinforcement learning model; wherein the reinforcement learning model includes an action neural network and an evaluation neural network, the action neural network is used to decide the access network selected by multiple terminal devices, and the evaluation neural network is used to evaluate the access network selected by multiple terminal devices.

[0156] For example, the acquisition unit 901 can support Figure 9 The network access selection device shown performs step 501 or step 601.

[0157] Training unit 902 is used to obtain first decision information based on environmental state information and reinforcement learning model; wherein, the optimization objective of reinforcement learning model is the total throughput of multiple terminal devices, the input parameters of reinforcement learning model include environmental state information, and the output parameters of reinforcement learning model include the access network selected for multiple terminal devices; wherein, the first decision information is used to indicate the access network selected for multiple terminal devices from the first access network and the second access network.

[0158] The training unit 902 is also used to obtain second decision information based on environmental state information and action neural network in the first decision time slot, and execute the second decision information to obtain the total throughput of multiple terminal devices; wherein, the second decision information is used to indicate the access network initially selected for multiple terminal devices from the first access network and the second access network.

[0159] Training unit 902 is also used to determine the reward value when a terminal device executes the second decision information based on the total throughput of multiple terminal devices.

[0160] The training unit 902 is also used to determine the temporal difference error based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network; wherein, the temporal difference error is used to adjust the training parameters of the reinforcement learning model.

[0161] The training unit 902 is also used to determine the loss function of the reinforcement learning model based on the temporal difference error, and update the training parameters of the reinforcement learning model according to the loss function until the reward value converges.

[0162] For example, training unit 902 can support Figure 4 The access network selection device shown performs step 502 or steps 602-605.

[0163] The training unit 902 may include a processor or a controller. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination of functions implementing computation, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0164] Specifically, the above Figures 5-6 All relevant content regarding each step in the illustrated method embodiment can be referenced from the functional description of the corresponding functional module unit, and will not be repeated here. The network access selection device is used to perform... Figures 5-6 The method shown in the diagram has the same function as the access network selection method described above, and therefore can achieve the same effect.

[0165] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a terminal of any of the foregoing embodiments, such as an internal storage unit including a data sending end and / or a data receiving end, like a hard disk or memory of the terminal. The computer-readable storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal. Further, the computer-readable storage medium can include both the internal storage unit and the external storage device of the terminal. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0166] This application also provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the network access selection method described in any embodiment of this application.

[0167] It should also be understood that the first, second, third, fourth and various numerical designations used herein are merely for descriptive convenience and are not intended to limit the scope of this application.

[0168] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0169] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0170] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0171] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0172] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0175] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0176] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0177] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0178] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for selecting an access network, characterized in that, The method includes: Obtain environmental status information; The environmental status information includes the location information of multiple terminal devices, parameters of a first access network model, and parameters of a second access network model; the multiple terminal devices are located within the coverage areas of the first access network and the second access network, and the first access network and the second access network support different access technologies; Based on the environmental state information and the reinforcement learning model, first decision information is obtained; the reinforcement learning model includes an action neural network and an evaluation neural network, the action neural network is used to decide the access network selected by the multiple terminal devices, and the evaluation neural network is used to evaluate the access network selected by the multiple terminal devices. The optimization objective of the reinforcement learning model is the total throughput of the multiple terminal devices. The input parameters of the reinforcement learning model include environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices. The first decision information is used to indicate the access network selected for the plurality of terminal devices from the first access network and the second access network.

2. The method according to claim 1, characterized in that, The method further includes: Configure the training parameters of the reinforcement learning model; In the first decision time slot, second decision information is obtained based on the environmental state information and the action neural network, and the second decision information is executed to obtain the total throughput of the plurality of terminal devices; wherein, the second decision information is used to indicate the access network initially selected for the plurality of terminal devices from the first access network and the second access network; The reward value for the terminal device when executing the second decision information is determined based on the total throughput of the plurality of terminal devices; The temporal difference error is determined based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network; wherein, the temporal difference error is used to adjust the training parameters of the reinforcement learning model; The loss function of the neural network is determined based on the time difference error, and the training parameters of the reinforcement learning model are updated based on the loss function until the reward value converges.

3. The method according to claim 1 or 2, characterized in that, The method further includes: A first rate is calculated based on the parameters of the first access network model and the location information of the terminal device, and a second rate is calculated based on the parameters of the second access network model and the location information of the terminal device; wherein, the first rate is the data rate when the terminal device accesses the first access network, and the second rate is the data rate when the terminal device accesses the second access network.

4. The method according to claim 2, characterized in that, The step of obtaining the total throughput of the multiple terminal devices by executing the second decision information includes: The total throughput of the plurality of terminal devices is determined based on the first rate, the second rate, the first decision variable, and the second decision variable. The total throughput of the plurality of terminal devices is the sum of the data rates of the plurality of terminal devices when they access the first access network or the second access network. The first decision variable and the second decision variable can be obtained based on the second decision information; the first decision variable is used to represent the access status of the terminal device to the first access network, and the second decision variable is used to represent the access status of the terminal device to the second access network.

5. The method according to claim 2 or 4, characterized in that, The second decision time slot is the next decision time slot after the first decision time slot; The acquisition of environmental state information includes: acquiring environmental state information at the beginning of each decision time slot.

6. A network access selection device, characterized in that, The device includes: Acquisition unit, used to acquire environmental status information; The environmental status information includes the location information of multiple terminal devices, parameters of a first access network model, and parameters of a second access network model; the multiple terminal devices are located within the coverage areas of the first access network and the second access network, and the first access network and the second access network support different access technologies; The training unit is used to obtain first decision information based on the environmental state information and the reinforcement learning model; the reinforcement learning model includes an action neural network and an evaluation neural network, the action neural network is used to decide the access network selected by the multiple terminal devices, and the evaluation neural network is used to evaluate the access network selected by the multiple terminal devices. The optimization objective of the reinforcement learning model is the total throughput of the multiple terminal devices. The input parameters of the reinforcement learning model include environmental state information, and the output parameters of the reinforcement learning model include the access network selected for the multiple terminal devices. The first decision information is used to indicate the access network selected for the plurality of terminal devices from the first access network and the second access network.

7. The apparatus according to claim 6, characterized in that, The acquisition unit is also used to configure the training parameters of the reinforcement learning model; The training unit is further configured to obtain second decision information based on the environmental state information and the action neural network in the first decision time slot, and execute the second decision information to obtain the total throughput of the plurality of terminal devices; wherein, the second decision information is used to indicate the access network initially selected for the plurality of terminal devices from the first access network and the second access network; The training unit is also configured to determine the reward value when the terminal device executes the second decision information based on the total throughput of the plurality of terminal devices; The training unit is further configured to determine the temporal difference error based on the location information of the terminal device in the first decision time slot, the location information of the terminal device in the second decision time slot, the reward value, and the evaluation neural network; wherein the temporal difference error is used to adjust the training parameters of the reinforcement learning model; The training unit is also used to determine the loss function of the reinforcement learning model based on the temporal difference error, and update the training parameters of the reinforcement learning model based on the loss function until the reward value converges.

8. The apparatus according to claim 6 or 7, characterized in that, The training unit is further configured to calculate a first rate based on the first access network model parameters and the location information of the terminal device, and to calculate a second rate based on the second access network model parameters and the location information of the terminal device; wherein, the first rate is the data rate when the terminal device accesses the first access network, and the second rate is the data rate when the terminal device accesses the second access network.

9. The apparatus according to claim 7, characterized in that, The training unit is also used to determine the total throughput of the plurality of terminal devices based on the first rate, the second rate, the first decision variable, and the second decision variable. The total throughput of the plurality of terminal devices is the sum of the data rates when the plurality of terminal devices access the first access network or the second access network. The first decision variable and the second decision variable can be obtained based on the second decision information; the first decision variable is used to represent the access status of the terminal device to the first access network, and the second decision variable is used to represent the access status of the terminal device to the second access network.

10. The apparatus according to claim 7 or 9, characterized in that, The second decision time slot is the next decision time slot after the first decision time slot; The acquisition unit is used to acquire environmental status information, including: at the beginning of each decision time slot, the acquisition unit acquires the location information of the terminal device in the decision time slot.

11. A network access selection device, characterized in that, The access network selection device includes one or more processors and a communication interface, wherein the one or more processors and the communication interface are used to support the communication device in performing the access network selection method as described in any one of claims 1-5.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on a computer, cause the computer to perform the access network selection method as described in any one of claims 1-5.

13. A computer program product containing instructions, characterized in that, When the instructions are executed on a computer, the computer performs the network access selection method as described in any one of claims 1-5.

14. A communication system, characterized in that, The communication system includes an access network selection device as described in any one of claims 6-10, and is capable of executing the access network selection method as described in any one of claims 1-5.