Remote area remote education heterogeneous network vertical switching method
By combining fuzzy logic systems with Q-learning algorithms, the problem of handover strategies in heterogeneous networks in remote areas being unable to adapt to dynamic changes was solved, enabling differentiated service quality requirements for multiple services and improving network handover efficiency and user experience.
Patent Information
- Application Number
- CN202511691573.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing handover strategies are ill-suited to the dynamic changes in heterogeneous networks in remote areas and cannot meet the differentiated quality of service requirements of multiple services such as video, audio, and data transmission, resulting in high handover latency and poor service continuity.
By combining fuzzy logic systems with Q-learning algorithms, the teaching service type and network type of user terminals are obtained. The membership function of the fuzzy logic system is used to evaluate network performance indicators, and the reward function of the Q-learning algorithm is combined to optimize the switching strategy, thereby achieving seamless switching that adapts to dynamic network changes.
It effectively reduces unnecessary network switching frequency, improves user experience, enhances switching efficiency and service continuity between heterogeneous networks, and meets the QoS requirements of multiple services.
Smart Images

Figure CN121510184A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distance education technology, and in particular to a method, system, device and medium for vertical switching of heterogeneous networks for distance education in remote areas. Background Technology
[0003] In recent years, with the development of information and communication technologies, distance education has been regarded as an effective way to bridge the education gap. Through live video streaming, online courses, and other means, students in remote areas can access high-quality educational resources from cities. Addressing the challenge of network coverage in remote areas, a single network technology is insufficient to meet diverse needs. However, a heterogeneous network architecture encompassing satellite communication, terrestrial base stations, and Delay Tolerant Networks (DTN) teaching vehicles demonstrates significant advantages, effectively solving the problem of network coverage dead zones in remote areas.
[0004] Remote education applications typically need to support multiple service types simultaneously, such as live video streaming, online interaction, and courseware downloads. These services have significantly different requirements for Quality of Service (QoS) metrics such as bandwidth, latency, and packet loss rate. Current handover strategies based on fixed thresholds often fail to make optimal decisions under frequently fluctuating network conditions, leading to problems such as excessive handover latency and frequent service interruptions, severely impacting the continuity and stability of the remote teaching experience. Therefore, how to achieve handover decisions that adapt to dynamic network changes and meet the differentiated QoS requirements of multiple services in heterogeneous network environments has become a critical issue that urgently needs to be addressed to ensure the real-time performance, reliability, and continuity of remote education.
[0005] Traditional handover algorithms assume that mobile users can accurately obtain network attributes, primarily relying on static threshold decisions based on Received Signal Strength (RSS) or overall Quality of Service (QoS) parameters. While simple to implement, this approach struggles to adapt to dynamic network changes. Machine learning techniques have been introduced into handover decision-making, significantly improving decision-making capabilities in dynamic environments. However, in remote education scenarios in remote areas, existing research still faces key challenges: on the one hand, the differentiated bandwidth, latency, and reliability requirements of multiple service flows such as video, voice, and data necessitate QoS trade-offs in handover strategies; on the other hand, the large fluctuations in networks in remote areas may prevent the provision of stable connections, resulting in limited contextual information.
[0006] Therefore, existing handover strategies are difficult to adapt to dynamic network changes and cannot meet the differentiated Quality of Service (QoS) requirements of multiple services such as video, audio, and data transmission, resulting in problems such as high handover latency and poor service continuity. Summary of the Invention
[0007] The purpose of this invention is to provide a method, system, device and medium for vertical handover of heterogeneous networks for distance education in remote areas. This invention can solve the problems of high handover latency and poor service continuity caused by existing handover strategies being unable to adapt to dynamic network changes and failing to meet the differentiated quality of service (QoS) requirements of multiple services such as video, audio and data transmission.
[0008] To address the aforementioned technical problems, embodiments of the present invention provide a method for vertical handover of heterogeneous networks for distance education in remote areas, comprising the following steps: Obtain the current teaching service type and network type of the user terminal; among them, there are multiple different teaching service types and multiple different network types for the user terminal; Based on the membership function of the fuzzy logic system, the membership degree of each network performance indicator of the current teaching business type under each network type is determined by the current teaching business type and the multiple different network performance indicators of the current teaching business type under each network type. The membership degree of each network performance indicator under each network type is weighted and summed to obtain the service quality score of the current teaching business type under each network type. Based on the Q-learning algorithm, the switching strategy for the current network type is determined according to the service quality score of the current teaching service type under each network type; the reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
[0009] Furthermore, for each network performance indicator of the current teaching service type, the membership degree corresponds to a fuzzy level under each network type, and there are three fuzzy levels: low, medium, and high. The weighted summation of the membership degrees of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type includes: Based on the preset weights of each network performance indicator corresponding to the current teaching service type and the fuzzy levels corresponding to the membership degree of each network performance indicator of the current teaching service type under each network type, the service quality score of the current teaching service type under each network type is determined.
[0010] Furthermore, the Q-learning algorithm-based determination of the current network type switching strategy according to the service quality score of the current teaching service type under each network type includes: Normalize the quality of service scores for the current teaching service type under each network type; Based on the preset mapping relationship between the service quality score and the five service quality levels of very low, low, medium, high and very high, the normalized service quality score is mapped to the corresponding service quality level. Based on the Q-learning algorithm, the switching strategy for the current network type is determined according to the quality of service level of the current teaching service type under each network type.
[0011] Furthermore, the state space of the Q-learning algorithm consists of the current network type, candidate network types other than the current network type, and the current teaching business type; The action space includes all network types. At each time step, the agent chooses whether to maintain the current network type or switch to another candidate network type.
[0012] Furthermore, the teaching service types include the following: video teaching, homework upload, resource download, course on demand, audio learning, quizzes and exams, and data synchronization; The network types include the following: terrestrial access networks, low-Earth orbit satellite networks, and latency-tolerant networks. Terrestrial access networks include 5G, WiFi, and LTE. The network performance metrics include the following: bandwidth, latency, signal strength, jitter, packet loss rate, and cost.
[0013] Embodiments of the present invention also provide a vertical switching system for heterogeneous networks in remote areas for distance education, comprising: The data acquisition module is used to acquire the current teaching service type and network type of the user terminal; among them, there are multiple different teaching service types and multiple different network types for the user terminal; The fuzzy processing module is used to determine the membership degree of each network performance indicator of the current teaching business type under each network type based on the membership function of the fuzzy logic system, through the current teaching business type and multiple different network performance indicators of the current teaching business type under each network type. The weighted scoring module is used to perform a weighted summation of the membership degree of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type. The switching decision module is used to determine the switching strategy for the current network type based on the Q-learning algorithm and the service quality score of the current teaching service type under each network type. The reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
[0014] Embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described vertical switching method for heterogeneous networks in remote education.
[0015] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for vertical switching of heterogeneous networks for remote education in remote areas.
[0016] The vertical handover method for heterogeneous networks in remote education provided by this invention has at least the following beneficial effects: First, a fuzzy logic system is used to evaluate the service quality of the current teaching service type under different network types. This service quality comprehensively considers multiple different network performance indicators for each network type, characterizing the dynamic characteristics under different network access scenarios. To simplify fuzzy inference while maintaining adaptability to different teaching service types, the fuzzy logic system of this invention differs from traditional rule-based fuzzy systems. It does not rely on a detailed rule base, but instead, after obtaining membership degrees using a membership function, it performs a weighted sum of the membership degrees of each network performance indicator under each network type to obtain the service quality score for the corresponding teaching service type under each network type. Then, considering... In multi-network heterogeneous distance education scenarios, network states change frequently and dynamically. Relying solely on fuzzy logic evaluation makes it difficult to capture the continuous evolution of network topology in real time, which can easily lead to frequent switching and potentially cause connection instability, increased system latency, and a degraded user experience. This invention is based on the Q-learning algorithm, which determines the switching strategy for the current network type based on the service quality score of the current teaching service type under each network type. The reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action. It can suppress unnecessary switching behavior, balance switching frequency and service quality, and achieve selection and seamless switching between heterogeneous networks. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0018] Figure 1 A flowchart illustrating a method for vertical switching of heterogeneous networks in remote education provided by the present invention; Figure 2 This invention provides a schematic diagram of a heterogeneous network handover system architecture. Figure 3 A topology diagram of a heterogeneous network for distance education in remote areas provided by the present invention; Figure 4 A flowchart illustrating a fuzzy weighted scoring mechanism provided by the present invention; Figure 5 A schematic diagram of an input variable membership function provided by the present invention; Figure 6 A schematic diagram of an output variable membership function provided by the present invention; Figure 7 A comparative diagram of the average number of switching operations for different algorithms provided by this invention; Figure 8 This is a schematic diagram comparing the average network performance of different algorithms provided by the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0020] Distance education is crucial for promoting educational equity in underdeveloped and remote areas, but it is often constrained by limited network infrastructure. Building an efficient and stable heterogeneous network coverage system is key to improving the quality of education in remote areas. Currently, a heterogeneous network composed of satellites, terrestrial base stations, and latency-tolerant DTN teaching vehicles can provide coverage support, but traditional vertical handover strategies are difficult to adapt to dynamic network changes and cannot meet the differentiated Quality of Service (QoS) requirements of multiple services such as video, audio, and data transmission, resulting in problems such as high handover latency and poor service continuity.
[0021] This invention provides a method, system, device, and medium for vertical switching of heterogeneous networks in remote education, which can adapt to different network conditions and user needs, optimize the user experience in remote education, and learn to avoid frequent network switching when unnecessary, thereby reducing latency and bandwidth consumption caused by switching and improving the user experience.
[0022] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] One embodiment of the present invention relates to a vertical handover method for heterogeneous networks in remote distance education. The specific process of the vertical handover method for heterogeneous networks in remote distance education in this embodiment can be as follows: Figure 1As shown, it includes: Step 101: Obtain the current teaching service type and network type of the user terminal; whereby the user terminal may have multiple different teaching service types and multiple different network types.
[0024] Step 102: Based on the membership function of the fuzzy logic system, determine the membership degree of each network performance indicator of the current teaching business type under each network type by using the current teaching business type and the multiple different network performance indicators of the current teaching business type under each network type.
[0025] Step 103: Perform a weighted summation of the membership degree of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type.
[0026] Step 104: Based on the Q-learning algorithm, determine the switching strategy for the current network type according to the service quality score of the current teaching service type under each network type; wherein, the reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
[0027] The following is a detailed description of the implementation details of the vertical switching method for heterogeneous networks in remote education in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.
[0028] The vertical switching method for heterogeneous networks in remote education in this embodiment is based on... Figure 2 The system architecture shown is a three-layer logical architecture, namely the business layer, the network layer and the algorithm layer.
[0029] (1) Business Layer: Primarily targeting various teaching services in distance education application scenarios, covering typical distance education applications such as video teaching (VT), assignment uploading (AU), resource downloading (DL), course on demand (RC), audio learning (AL), quizzes and exams (EX), and data synchronization (DS). Different business types have significantly different sensitivities to system resources. Table 1 provides a service level requirement hierarchy categorized by business process type. The system needs to adaptively adjust its access and switching strategies according to the business type to meet differentiated quality of service (QoS) requirements, as follows:
[0030] Table 1 (2) Network Layer: To improve network availability and overall service capabilities in remote areas, the system incorporates various heterogeneous access resources in its network layer design. The system integrates existing 5G, 4G, and WiFi micro base station networking solutions, flexibly deploying them in suitable areas such as education centers and teaching stations to meet the needs of some end users for high-throughput, low-latency network resources. A low-Earth orbit (LEO) satellite network is introduced, utilizing its wide-area coverage and dynamic orbit characteristics to provide basic communication guarantees for large-scale remote areas, compensating for service blind spots caused by insufficient ground infrastructure. To further enhance the system's buffering and relay capabilities, delay-tolerant network (DTN) equipment is also introduced, enabling cached forwarding of data through deployment in teaching vehicles and cache relay stations, effectively supporting the needs of high-capacity, non-real-time data synchronization and remote education content distribution. The combined coverage of multiple heterogeneous networks provides strong support for stable services across various business types in remote education.
[0031] (3) Algorithm Layer: Addressing the complex handover decision-making problem in distance education scenarios involving multiple parallel services and frequent network dynamics, this paper introduces a joint optimization mechanism of fuzzy weighted scoring and Q-learning in the algorithm layer design, forming an intelligent handover decision-making module. This method uses a fuzzy inference system for real-time multi-dimensional service quality assessment (bandwidth / latency / RSS), while simultaneously utilizing the Q-learning algorithm to dynamically optimize the handover strategy. This dual mechanism minimizes handover overhead while ensuring service quality, achieving closed-loop optimization of network state awareness and handover decision-making.
[0032] To achieve comprehensive network coverage in remote areas, this embodiment designs a heterogeneous network architecture encompassing satellite, DTN equipment, and various terrestrial access networks (5G, WiFi, LTE), enhancing the overall communication system's coverage, access flexibility, and service continuity. The network topology is shown below. Figure 3 As shown, a low-Earth orbit (LEO) satellite network enables dynamic wide-area coverage, ensuring overall regional network availability. This paper assumes the deployment of one LEO satellite, covering the entire area. A 5G macro base station is deployed in the ground core area to provide high-bandwidth, low-latency communication for teaching buildings and large teaching centers; two LTE base stations and two WiFi hotspots are deployed at surrounding teaching sites. The various base stations form a partially overlapping cell layout, facilitating subsequent handover control and load balancing optimization.
[0033] Furthermore, considering the latency sensitivity and long-term intermittent connectivity issues at some extremely remote teaching sites, the teaching vehicle is equipped with a Delay-Tolerant Network (DTN) caching module. The teaching vehicle cruises along a pre-set periodic trajectory within the core and edge coverage areas, periodically synchronizing the core data cache to achieve data pre-distribution, offline transmission, and spatiotemporal joint optimization of the caching strategy. Its trajectory is modeled using a closed elliptical orbit, expressed by the following formula:
[0034] ; In the formula, and These are the major and minor semi-axes of the elliptical trajectory, respectively. This indicates the cruise cycle rate control parameter, ensuring that the vehicle completes full coverage cruise within one cycle.
[0035] Various networks differ significantly in terms of coverage, bandwidth capacity, latency characteristics, and service costs. In practical educational applications, end users need to dynamically select and switch between different network resources to ensure service continuity and Quality of Service (QoS) requirements.
[0036] Furthermore, in the heterogeneous network architecture for remote education in remote areas proposed in this embodiment, terminals need to dynamically switch between different available networks during actual business operations. To support subsequent fuzzy logic decision-making and Q-learning training, it is necessary to reasonably model the core performance indicators of the system to characterize the dynamic characteristics under different network access scenarios. This embodiment mainly models from six dimensions: bandwidth, latency, signal strength (RSS), jitter, packet loss rate, and cost.
[0037] (1) RSS: Received Signal Strength (RSS) is a key indicator for evaluating communication quality, reflecting the signal strength from the user terminal to the access network (such as satellite or terrestrial base stations). To simplify the model, this paper uses the path loss model to calculate RSS, as shown in the following formula:
[0038] ; In the formula, R0 represents the reference signal strength, d is the distance to the base station, and α is the path loss coefficient.
[0039] (2) Bandwidth: Bandwidth modeling takes into account the impact of dynamic network load and assumes the maximum bandwidth of each network node. Below, based on the current load factor The actual available bandwidth is dynamically adjusted, and the modeling formula is as follows: ; In the formula, The values are generated through random perturbations to simulate the dynamic impact of real-time network congestion fluctuations on resource availability, and β is the distance decay factor.
[0040] (3) Delay: For delay modeling, this embodiment considers the superposition effect of basic propagation and load-related queuing delay, and establishes the following model: ; In the formula, This indicates the delay between basic propagation and forwarding; This is a small-amplitude disturbance term.
[0041] (4) Packet loss rate: Packet loss rate is a crucial indicator affecting service continuity and learning experience. This embodiment combines network load and channel quality to comprehensively model the packet loss rate, and its calculation formula is as follows:
[0042] ; In the formula, This represents the base packet loss rate, which varies depending on the network type. λ represents the current network load ratio. This indicates the load's sensitivity to packet loss rate. This represents the signal strength influence coefficient.
[0043] (5) Shaking: Jitter describes the severity of latency fluctuations in a network, and jitter control is particularly important in real-time audio and video teaching scenarios. This embodiment uses the following model for description:
[0044] ; In the formula, Based on the jitter.
[0045] (6) Cost: Different network access methods result in different operation and maintenance costs. Considering the unit service cost of different network types in remote education deployments, the following simplified model is defined: ; In the formula, The unit cost constant is preset for different network types, satisfying: ; This is to reflect the differences in the complexity and resource consumption of different network deployments.
[0046] Based on the above architecture, a vertical handover strategy integrating a fuzzy weighted scoring mechanism and Q-learning (FIMQL) is proposed. This strategy utilizes a fuzzy logic system to obtain QoS values and introduces a service-weighted scoring mechanism to adaptively match various services with differentiated QoS requirements in distance education. Finally, reinforcement learning is used to determine whether to switch to the network with the highest QoS value. This approach can handle uncertain information in the scenario and learn from historical experience to make optimal decisions in these changing environments.
[0047] Specifically, based on the acquired current teaching service type and network type of the user terminal, firstly, according to the membership function of the fuzzy logic system, the membership degree of each network performance indicator for the current teaching service type under each network type is determined through the current teaching service type and multiple different network performance indicators for each network type. Each membership degree of the current teaching service type under each network type corresponds to a fuzzy level, with three levels: low, medium, and high. When performing a weighted summation of the membership degrees of each network performance indicator under each network type, the service quality score of the current teaching service type under each network type is determined based on the preset weights of each network performance indicator corresponding to the current teaching service type and the fuzzy levels corresponding to the membership degrees of each network performance indicator under each network type.
[0048] In one example, after determining the service quality score of the current teaching service type under each network type, the service quality score of the current teaching service type under each network type is first normalized. Then, according to the preset mapping relationship between the service quality score and five service quality levels (very low, low, medium, high, and very high), the normalized service quality score is mapped to the corresponding service quality level. This allows the Q-learning algorithm to determine the switching strategy for the current network type based on the service quality level of the current teaching service type under each network type.
[0049] In its implementation, this embodiment proposes a fuzzy weighted scoring mechanism to simplify fuzzy reasoning while maintaining adaptability to different service requirements. This mechanism combines membership function modeling with a parameterized scoring function. Unlike traditional rule-based fuzzy systems (which include fuzzification, reasoning, and defuzzification stages), this method simplifies the reasoning process by replacing manually defined rules with a structured scoring strategy, such as... Figure 4 As shown. This enables dynamic and interpretable quality of service assessments in different network environments.
[0050] First, latency, bandwidth, RSS, and service type are selected as input variables, and fuzzification is performed using a trapezoidal-triangle composite membership function. Each membership function is categorized into three linguistic levels: low, medium, and high, such as... Figure 5 and Figure 6 As shown, this system does not rely on an exhaustive rule base, but instead uses a parameterized scoring function to compute a continuous Qos score. This score is then mapped to five semantic categories—very low, low, medium, high, and very high—to facilitate the seamless integration of interpretable Qos evaluation and reinforcement learning.
[0051] In distance education scenarios, different teaching service types exhibit significantly different focuses on network performance metrics. To more reasonably reflect the diversity of services and the heterogeneity of requirements, this embodiment designs corresponding indicator weight matrices for three QoS metrics—bandwidth, latency, and signal strength (RSS)—based on the service characteristics and service sensitivity of each service type, as shown in the following formula. Each row corresponds to a different service type, and each column represents the weight of bandwidth, latency, and RSS, respectively. Specifically, the RSS weight is uniformly set to 0.2 to reflect the universal sensitivity of different services to differences in signal strength in remote scenarios.
[0052] ; During rule generation, each input variable is assigned a corresponding level number (1, 2, 3) based on its fuzziness level (Low, Medium, High). Then, the comprehensive score for each rule is calculated using the following weighted scoring function:
[0053] ; In the formula, B, D, and R represent the ranks of the input variables in their respective membership degrees, and W... b W d W r The weight corresponding to the business type.
[0054] To ensure a smooth mapping, the score can be normalized: ; To ensure a consistent mapping between the weighted scoring results and the QoS output levels of the fuzzy logic system, this embodiment divides the scoring interval into five discrete levels and directly maps them to the QoS membership function output. This ensures a consistent mapping between the scoring function and the fuzzy membership output in the fuzzy inference system, thereby improving the system's stability, scalability, and interpretability. The specific mapping relationships are shown in Table 2.
[0055] Table 2 This QoS assessment model enables the model to adapt to services, capture the nonlinear interaction between complex changes in heterogeneous networks and service characteristics, and provide a robust decision-making basis for intelligent handover.
[0056] After obtaining the Quality of Service (QoS) score of the current teaching service type under each network type, the switching strategy for the current network type is determined based on the Q-learning algorithm according to the QoS score of the current teaching service type under each network type.
[0057] Specifically, in multi-network heterogeneous distance education scenarios, network states change frequently and dynamically. Relying solely on fuzzy logic QoS assessment is insufficient to capture the continuous evolution of network topology in real time, easily leading to frequent handovers. This can cause connection instability, increased system latency, and a degraded user experience, particularly affecting video and audio quality, potentially severely impacting the smoothness and stability of distance education. Therefore, this paper introduces a reinforcement learning (RL) handover decision module based on fuzzy logic QoS assessment to achieve intelligent, dynamic, and adaptive vertical handover control. Q-Learning is a classic reinforcement learning algorithm used to solve Markov Decision Process (MDP) problems. The characteristics of an MDP can be described as a tuple. ,in The state space representing the intelligent agent; The action space representing the intelligent agent; Represents the state transition probability of an MDP; It is a reward function; It is a discount factor.
[0058] (1) State space: As shown in the following equation, the state space It consists of the QoS and service type of the current access network, all candidate networks (i.e., the current network type, candidate network types other than the current network type, and the current teaching service type), which enables the agent to consider the network environment and service requirements in order to make informed decisions: ; (2) Action space: The action space encompasses all available network types. At each time step, the agent chooses whether to maintain the current connection or switch to another network. In this paper, heterogeneous networks include five network types, therefore the action space is as follows:
[0059] ; In the formula, .
[0060] (3) Reward function: The following formula describes the service quality improvement resulting from state transitions in the system. A high reward is given when the service quality improvement is significant and the switching cost is low. Conversely, if the service quality change is small but the switching cost is high, a negative reward is applied to discourage unnecessary switching behavior.
[0061] ; (4) Q-value function: Q-learning evaluates a set of state-action pairs using a Q-value function to maximize the expected value of the discounted cumulative reward. Evaluation function Representing state and actions The reward value for the correct state, and it is defined as the reward value obtained from the state. Begin and take action And the maximum cumulative discount reward you receive. In other words, This represents the sum of immediate rewards and cumulative discount rewards for following the optimal strategy.
[0062] The Q function can be represented as: ; In the formula, Indicates the current state Take action The immediate reward value afterward; Indicates the discount factor; Indicates the state The maximum cumulative discount reward value that can be obtained is, i.e. Indicates the state The cumulative reward value follows the optimal strategy. In this embodiment, "state" refers to the network where the current user node can receive signals, and "action" refers to choosing which network to switch to.
[0063] state The optimal strategy under the given conditions is the action strategy that maximizes the immediate benefit and the cumulative benefit of subsequent states, that is: ; This results in the following expression: ; Therefore, the recursively defined Q-value evaluation function can be expressed as: ; In the formula, The learning rate is used to determine the impact of new information on existing information and to control the convergence speed.
[0064] During Q-learning, the agent only needs to compare each state-action pair. The function can determine the optimal policy action.
[0065] Therefore, the Q-learning-based vertical handover algorithm in this embodiment can be as shown in Algorithm 1, which implements the learning process of finding the optimal network access decision strategy. The inputs to Algorithm 1 include the learning rate α, parameter ϵ, number of learning rounds n, and the Quality of Service (QoS) for all users. The output is the optimal network access decision strategy. For each QoS, learning can be performed according to the following steps: ① Initialize the Q-table to 0. ② Perform multiple rounds of training. ③ Start the Q-learning process from LEO as the initial state. ④ The agent calculates the reward value of the action according to the above formula, updates the Q-value, and then moves to the new state until the end of the training round. ⑤ After training, find the accessible network with the maximum Q-value for the corresponding service in the Q-table.
[0066] To verify the effectiveness of the method proposed in this embodiment, an experiment based on MATLAB R2022a is conducted to simulate a heterogeneous network environment for remote education in a remote area. The scenario is set as a 15*15*15 grid containing one LEO satellite. It is assumed that the LEO network in this experiment has complete coverage. The ground is centered on a 5G base station, with one WiFi access point and one LTE base station node on each side. A DTN teaching vehicle moves periodically along an elliptical trajectory around the perimeter of the scenario, simulating the function of a real teaching vehicle as a buffer forwarding node. The teaching vehicle periodically passes through different user areas, providing intermittent supplementary services to users. Remote education users are distributed in various areas within the scenario, and users select the best access signal in real time based on their service needs and network status during movement. Network attributes are shown in Table 3.
[0067] Table 3 The baseline method comparison method is as follows: FL-VHO: A fuzzy logic-based algorithm is proposed to adjust the handover margin (HO) in HetNets by utilizing cell load and radio channel quality. Furthermore, a two-step target cell determination model considering cell load level and reference signal received power (RSRP) is presented.
[0068] MA-VHO proposes an intelligent access network selection algorithm based on multi-attribute decision-making. First, considering the characteristics of cognitive wireless networks, a cross-layer network framework based on cognitive cycles is proposed. Second, network selection decisions are made using cross-layer multi-attributes based on network conditions and user QoE.
[0069] AHP-EEM-GRA proposes a vertical switching decision scheme for cognitive heterogeneous networks that combines subjective and objective weighting with grey relational analysis. The algorithm ranks candidate networks and is integrated into the cognitive loop of the agent.
[0070] FRL-VHO: To address the problem that users cannot obtain accurate network attributes, a vertical switching algorithm for heterogeneous networks that integrates fuzzy logic systems and reinforcement learning is proposed.
[0071] The simulation includes seven types of remote education services, each corresponding to different QoS preferences. The service type is determined during user initialization using a weighted random sampling method and remains unchanged during user movement. The sampling probability distribution of the service types is shown in Table 4.
[0072] Table 4 The relevant parameters for Q learning are shown in Table 5: Table 5 This invention compares the proposed method with four other methods in terms of average handover times and network performance, and obtains relevant data through 100 experiments. The average value is as follows: Figure 7 and Figure 8 As shown. Figure 7 It shows the average total number of times a mobile terminal switches between different available networks, which serves as a qualitative criterion for evaluating the quality of vertical handover algorithms in heterogeneous networks.
[0073] As shown in the figure, the proposed method FIMQL-VHO has the lowest number of handovers, reducing it by 26.9% compared to FL-VHO, 13.6% compared to FRL-VHO, 20.8% compared to MA-VHO, and 38.7% compared to AHP-EEM-GRA. This indicates that in the scenario described in this invention, it can keenly perceive the network environment in which the user equipment is located and determine the optimal decision on whether to perform a network handover through continuous iteration and optimization. The red bar FRL-VHO shows the second-best performance. Both methods share the commonality of using reinforcement learning for handover decisions, demonstrating their ability to consider multiple factors, such as network bandwidth, signal strength, and handover latency, from a global and long-term perspective, thus making more informed decisions. In contrast, traditional algorithms, lacking a deep understanding of the network environment and dynamic adaptability, often lead to frequent and unnecessary handovers. For example, when the network signal experiences a brief fluctuation, some algorithms may immediately trigger a handover operation, ignoring the costs and risks associated with the handover itself.
[0074] Specific network performance metrics, such as Figure 8 As shown. From Figure 8 As can be seen from the above, the FIMQL-VHO method proposed in this invention outperforms other algorithms in comparison in several network performance indicators that are crucial for remote education in remote areas, such as bandwidth, latency, packet loss rate, jitter, and cost.
[0075] in, Figure 8 (a) shows a comparison of the five algorithms in terms of average bandwidth. The graph shows that FIMQL-VHO has the highest average bandwidth, exceeding FRL-VHO by 2.4%. This indicates that it can accurately and flexibly allocate limited bandwidth resources according to the actual needs of distance education. In contrast, the average bandwidth of the other algorithms is reduced to varying degrees, which may not be sufficient for large file resources such as videos. Figure 8 (b) shows a comparison of the five algorithms in terms of average latency. Latency is a key indicator of network response speed and is crucial for distance education services with high real-time requirements. Figure 8 As shown in (b), the FIMQL-VHO method and the FRL-VHO method perform similarly, with a difference of only about 1ms. Other algorithms, however, have latency increases of 26%, 13%, and 6.5% respectively compared to this method. Figure 8 (c) shows a comparison of the average cost of the five algorithms. For remote areas, many areas without terrestrial base station coverage require satellite communication, but satellite usage is expensive. For some less urgent services, DTN equipment can be used for communication. Figure 8 (c) It can be seen that the FIMQL-VHO method maintains the lowest cost while ensuring network performance. Jitter and packet loss rate are important indicators for measuring network reliability. Figure 8 (d) and Figure 8 (e) shows the average jitter and packet loss rate of the five algorithms. As can be seen from the figure, FIMQL outperforms the other algorithms in both of these metrics.
[0076] In summary, this invention aims to address the challenges of diverse QoS requirements across multiple services and frequent dynamic changes in network status faced by heterogeneous networks in remote education. By dynamically evaluating network status and optimizing decisions, this strategy can meet the differentiated QoS requirements of multiple service scenarios and improve network performance in remote education. Specifically, a fuzzy weighted scoring mechanism is used to comprehensively evaluate uncertain indicators such as network bandwidth, latency, and packet loss rate to generate a comprehensive network score. Based on this, Q-learning is used to dynamically optimize handover decisions, and a handover penalty mechanism is introduced to balance handover frequency and service quality, ensuring efficient and seamless handover between heterogeneous networks. Simulation results show that the proposed method outperforms other traditional methods in terms of network performance indicators such as average number of handovers, bandwidth, latency, jitter, packet loss rate, and cost. This strategy not only improves the efficiency of network handover but also enhances the stability and continuity of remote education services, which is of great significance for the popularization and quality improvement of education in remote areas.
[0077] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.
[0078] Another embodiment of the present invention relates to a vertical handover system for heterogeneous networks in remote areas for distance education. The implementation details of this embodiment's vertical handover system are described below. The following details are provided for ease of understanding and are not essential for implementing this solution. This embodiment's vertical handover system for heterogeneous networks in remote areas for distance education includes: The data acquisition module is used to acquire the current teaching service type and network type of the user terminal; among them, there are multiple different teaching service types and multiple different network types for the user terminal; The fuzzy processing module is used to determine the membership degree of each network performance indicator of the current teaching business type under each network type based on the membership function of the fuzzy logic system, through the current teaching business type and multiple different network performance indicators of the current teaching business type under each network type. The weighted scoring module is used to perform a weighted summation of the membership degree of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type. The switching decision module is used to determine the switching strategy for the current network type based on the Q-learning algorithm and the service quality score of the current teaching service type under each network type. The reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
[0079] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0080] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0081] Another embodiment of the present invention relates to a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the vertical switching method for heterogeneous networks for remote education in the above embodiments.
[0082] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0083] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0084] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0085] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0086] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A method for vertical switching of heterogeneous networks for distance education in remote areas, characterized in that, include: Obtain the current teaching service type and network type of the user terminal; among them, there are multiple different teaching service types and multiple different network types for the user terminal; Based on the membership function of the fuzzy logic system, the membership degree of each network performance indicator of the current teaching business type under each network type is determined by the current teaching business type and the multiple different network performance indicators of the current teaching business type under each network type. The membership degree of each network performance indicator under each network type is weighted and summed to obtain the service quality score of the current teaching service type under each network type. Based on the Q-learning algorithm, the switching strategy for the current network type is determined according to the service quality score of the current teaching service type under each network type; the reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
2. The vertical switching method for heterogeneous networks in remote education areas according to claim 1, characterized in that, The membership degree of each network performance indicator under each network type for the current teaching business type corresponds to a fuzzy level, and there are three fuzzy levels: low, medium, and high. The weighted summation of the membership degrees of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type includes: Based on the preset weights of each network performance indicator corresponding to the current teaching service type and the fuzzy levels corresponding to the membership degree of each network performance indicator of the current teaching service type under each network type, the service quality score of the current teaching service type under each network type is determined.
3. The vertical switching method for heterogeneous networks in remote education areas according to claim 2, characterized in that, The Q-learning algorithm determines the current network type switching strategy based on the quality of service score of the current teaching service type under each network type, including: Normalize the quality of service scores for the current teaching service type under each network type; Based on the preset mapping relationship between the service quality score and the five service quality levels of very low, low, medium, high and very high, the normalized service quality score is mapped to the corresponding service quality level. Based on the Q-learning algorithm, the switching strategy for the current network type is determined according to the quality of service level of the current teaching service type under each network type.
4. The vertical switching method for heterogeneous networks in remote education areas according to claim 1, characterized in that, The state space of the Q-learning algorithm consists of the current network type, candidate network types other than the current network type, and the current teaching business type. The action space includes all network types. At each time step, the agent chooses whether to maintain the current network type or switch to another candidate network type.
5. The vertical handover method for heterogeneous networks in remote education according to any one of claims 1 to 4, characterized in that, The teaching services include the following types: video teaching, homework upload, resource download, course on demand, audio learning, quizzes and exams, and data synchronization; The network types include the following: terrestrial access networks, low-Earth orbit satellite networks, and latency-tolerant networks. Terrestrial access networks include 5G, WiFi, and LTE. The network performance metrics include the following: bandwidth, latency, signal strength, jitter, packet loss rate, and cost.
6. A vertical switching system for heterogeneous networks in remote education areas, characterized in that, The system includes: The data acquisition module is used to acquire the current teaching service type and network type of the user terminal; among them, there are multiple different teaching service types and multiple different network types for the user terminal; The fuzzy processing module is used to determine the membership degree of each network performance indicator of the current teaching business type under each network type based on the membership function of the fuzzy logic system, through the current teaching business type and multiple different network performance indicators of the current teaching business type under each network type. The weighted scoring module is used to perform a weighted summation of the membership degree of each network performance indicator under each network type to obtain the service quality score of the current teaching service type under each network type. The switching decision module is used to determine the switching strategy for the current network type based on the Q-learning algorithm and the service quality score of the current teaching service type under each network type. The reward function of the Q-learning algorithm is constructed based on the change in service quality score after the action is executed and the cost of executing the action.
7. A computer device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the vertical switching method for heterogeneous networks for remote education as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for vertical switching of heterogeneous networks for remote education in remote areas as described in any one of claims 1 to 5.