Intelligent management method of portable WIFI based on reinforcement learning and portable WIFI

By monitoring users' data usage records and signal strength, and using reinforcement learning models to optimize the data transmission rate and data plan of portable Wi-Fi, the problems of insufficient speed and resource waste in existing technologies are solved, thereby improving user experience and operator management efficiency.

CN119676756BActive Publication Date: 2025-11-04SHENZHEN KUYI ELECTRONIC TECHNOLOGY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411926977.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-11-04
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Existing portable Wi-Fi management systems fail to adequately consider users' real-time needs and changes in the network environment, resulting in insufficient data transmission rates or wasted resources. Furthermore, data plan recommendations are not personalized enough, impacting user experience and the efficiency of operators' data management.

Method used

By monitoring users' data usage records, connection duration, signal strength, and other status data, reinforcement learning models are used to analyze data transmission rates and intelligently adjust data plans and recommendations to optimize data transmission rates and data usage status.

Benefits of technology

It enables dynamic adjustment of data transmission rates based on user needs and network environment, reducing insufficient or wasted data rates, improving user experience and operator traffic management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119676756B_ABST
    Figure CN119676756B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of reinforcement learning, WIFI, internet access and related services, and provides an intelligent management method and a personal WIFI based on reinforcement learning, wherein the state representation data of the personal WIFI of a current user is monitored, the state representation data of the personal WIFI includes the traffic usage record of the personal WIFI of the user, the continuous use duration of the personal WIFI of the user each time, and the signal strength of the personal WIFI of the user equipment, the state representation data of the personal WIFI is input into a trained reinforcement learning model, the data transmission rate of the personal WIFI adapted to the current user is obtained through reinforcement learning model analysis, the traffic usage state of the traffic package of the personal WIFI of the current user is intelligently adjusted according to the data transmission rate, and the traffic package of the personal WIFI adapted to the current user is intelligently recommended according to the traffic usage state, so that the data transmission rate and the traffic recommendation of the personal WIFI are optimized, and the product use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical fields of reinforcement learning, WIFI, internet access and related services, and particularly relates to an intelligent management method and system for portable WIFI based on reinforcement learning. BACKGROUND

[0002] As an internet access device, portable WIFI can connect devices connected to the portable WIFI signal to the internet through the communication network of a communication operator to realize data transmission. Before obtaining the communication permission to access the internet, the user of the device connected to the portable WIFI signal needs to log in to the portable WIFI management system of the portable WIFI operator to purchase a traffic package. Different traffic packages correspond to different traffic and prices. After the traffic package takes effect, in order to stimulate traffic use, the portable WIFI management system configures a data transmission rate according to the traffic balance. A larger traffic balance corresponds to a faster data transmission rate. However, the portable WIFI management system simply configures the data transmission rate according to the traffic balance, and does not fully consider the real-time demand of the user and the change of the network environment, resulting in insufficient data transmission rate or resource waste in some cases. In addition, the recommendation mechanism of the traffic package is simple, and cannot make personalized recommendations according to the user's usage habits and needs, affecting the user experience and the traffic management efficiency of the operator.

[0003] In summary, the existing management technology of portable WIFI has the technical problems of insufficient data transmission rate or resource waste, and cannot make personalized recommendations according to the user's usage habits and needs. SUMMARY

[0004] In view of the above-mentioned problems in the prior art, the present application provides an intelligent management method and system for portable WIFI based on reinforcement learning to optimize the data transmission rate and traffic recommendation of portable WIFI and improve the product use experience of users.

[0005] In a first aspect, the present application provides an intelligent management method for portable WIFI based on reinforcement learning, comprising:

[0006] monitoring state representation data of the current user's portable WIFI, wherein the state representation data of the portable WIFI includes the user's portable WIFI traffic usage record, the user's continuous use duration for each connection to the portable WIFI, and the signal strength of the portable WIFI of the user's device;

[0007] inputting the state representation data of the portable WIFI into a trained reinforcement learning model, and obtaining a data transmission rate adapted to the current user's portable WIFI through analysis of the reinforcement learning model;

[0008] According to the data transmission rate, the traffic usage state of the traffic package of the current user's portable WIFI is intelligently adjusted, and the traffic package suitable for the current user's portable WIFI is intelligently recommended according to the traffic usage state.

[0009] In a second aspect, the application provides a portable WIFI, which is managed by using the intelligent management method of the portable WIFI based on reinforcement learning.

[0010] Compared with the prior art, the application has the following beneficial effects:

[0011] The application provides an intelligent management method of portable WIFI based on reinforcement learning and a portable WIFI. The state representation data of the current user's portable WIFI is monitored, the state representation data of the portable WIFI includes the traffic usage record of the user's portable WIFI, the continuous use duration of the user's portable WIFI each time the user connects the portable WIFI, and the signal strength of the portable WIFI of the user's device, the state representation data of the portable WIFI is input into a trained reinforcement learning model, the data transmission rate suitable for the current user's portable WIFI is obtained by analyzing the reinforcement learning model, according to the data transmission rate, the traffic usage state of the traffic package of the current user's portable WIFI is intelligently adjusted, and the traffic package suitable for the current user's portable WIFI is intelligently recommended according to the traffic usage state, so that the data transmission rate and the traffic recommendation of the portable WIFI are optimized, and the product use experience of the user is improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are included to provide a further understanding of the application and constitute a part of this application, illustrate certain illustrative embodiments of the application and together with the description serve to explain the application. In the drawings, the same components have the same reference numerals. It should be understood that the drawings are not necessarily to scale, and that, in some instances, various aspects of the application can be shown exaggerated or enlarged to facilitate an understanding of the application. The application will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:

[0013] Figure 1 is a flowchart of the intelligent management method of the portable WIFI based on reinforcement learning according to an embodiment of the application;

[0014] Figure 2 is a system architecture diagram when the continuous use duration of the user's portable WIFI each time the user connects the portable WIFI is counted according to an embodiment of the application;

[0015] Figure 3 is a system architecture diagram when the signal strength of the portable WIFI of the user's device is counted according to an embodiment of the application;

[0016] Figure 4Fig. 1 shows a schematic diagram of an architecture of a server according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should belong to the protection scope of the present application.

[0018] Embodiment one

[0019] Referring to Figures 1-4 The embodiment provides an intelligent management method of a personal WIFI based on reinforcement learning, which comprises the following steps: S101, S102 and S103. The steps S101, S102 and S103 can be run on a server to implement the intelligent management method of the personal WIFI based on reinforcement learning. The state representation data of the personal WIFI of a current user is monitored, the state representation data of the personal WIFI comprises the traffic usage record of the personal WIFI of the user, the continuous use duration of the personal WIFI of the user each time, and the signal strength of the personal WIFI of the user equipment, the state representation data of the personal WIFI is input into a trained reinforcement learning model, the data transmission rate suitable for the personal WIFI of the current user is obtained by analyzing the reinforcement learning model, the traffic usage state of the traffic package of the personal WIFI of the current user is intelligently adjusted according to the data transmission rate, and the traffic package suitable for the personal WIFI of the current user is intelligently recommended according to the traffic usage state, so as to optimize the data transmission rate and the traffic recommendation of the personal WIFI and improve the product use experience of the user.

[0020] Referring to Figure 4 , Figure 4 Fig. 1 shows a schematic diagram of an architecture of a server according to an embodiment of the present application. It should be noted that the server is the execution subject of all or part of the steps in the intelligent management method of the personal WIFI based on reinforcement learning, and in addition to the steps S101, S102 and S103 in the embodiment, it can also run part or all of the steps of the methods involved below.

[0021] From Figure 4It can be known that the server can be an information processing device including a memory, a processor and a network interface connected to each other through a system bus. Those skilled in the art can understand that the server herein is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and the hardware thereof includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc. The server can interact with the user through a touchpad or a voice control device, etc. The memory includes at least one type of readable storage medium, which includes a flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the server, such as a hard disk or a memory of the server. In other embodiments, the memory can also be an external storage device of the server, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the server. Of course, the memory can include both the internal storage unit and the external storage device of the server.

[0022] Referring to Figures 1-4 The embodiment provides a smart management method of a portable WIFI based on reinforcement learning, which comprises the following steps:

[0023] S101, monitoring state representation data of a portable WIFI of a current user, wherein the state representation data of the portable WIFI of the user comprises a traffic usage record of a portable WIFI of the user, a continuous use time length of the user connecting the portable WIFI each time, and a signal strength of the portable WIFI of the user equipment;

[0024] S102, inputting the state representation data of the portable WIFI into a trained reinforcement learning model, and obtaining a data transmission rate suitable for the portable WIFI of the current user through analysis of the reinforcement learning model;

[0025] S103, according to the data transmission rate, intelligently adjust the traffic usage state of the current user's portable WIFI traffic package, and intelligently recommend an adaptive traffic package for the current user's portable WIFI according to the traffic usage state.

[0026] It should be noted that in step S101, monitoring the state of the current user's portable WIFI indicates data, including the user's portable WIFI traffic usage record, the user's each time connecting the portable WIFI duration of use and the user equipment's portable WIFI signal strength, so as to obtain the user's use behavior and current network environment information, reflecting the user's real-time demand and the change of network conditions. In step S102, the reinforcement learning model is used to analyze the monitored state of the current user's portable WIFI, predict and determine the optimal data transmission rate. Reinforcement learning model can learn user's use habits and dynamic changes of network environment through continuous training and feedback, and adjust the data transmission rate in real time, so as to avoid simple rate configuration based on traffic balance, improve resource utilization efficiency, and reduce the situation of insufficient or waste rate. In step S103, according to the data transmission rate, the usage state of the traffic package is intelligently adjusted, and the adaptive traffic package is intelligently recommended, to ensure that the user obtains the optimal traffic package, which not only meets the data transmission demand, but also avoids the waste of resources. At the same time, the personalized recommendation mechanism can improve the user experience and improve the traffic management efficiency of the operator.

[0027] In some preferred embodiments, when training the reinforcement learning model, the following steps are included: obtaining statistical state representation data of the personal WIFI, the state representation data including the traffic usage record of the personal WIFI of the user, the duration of continuous use of the personal WIFI of the user each time the user connects to the personal WIFI, and the signal strength of the personal WIFI of the user equipment; integrating the obtained state representation data of the personal WIFI into a state vector, and defining an adjustment action of the data transmission rate of the personal WIFI corresponding to the state vector; defining a reward function, the reward function evaluating the adjustment of the data transmission rate based on user satisfaction and resource utilization, the reward function being used to guide the reinforcement learning model to optimize the adjustment strategy of the data transmission rate in the training process; selecting a proximal policy optimization algorithm as a reinforcement learning algorithm, and training the reinforcement learning algorithm through the state vector, the adjustment action of the data transmission rate, and the reward function, to obtain a reinforcement learning model capable of analyzing the adaptive data transmission rate of the personal WIFI of the current user. It should be noted that the traffic usage record and the usage duration directly reflect the user's online demand and habits. These data can help the model understand the demand changes of the user in different time periods and different use scenarios. The signal strength affects the data transmission rate and stability. Understanding the signal strength can help the model adjust the strategy when the network condition is not good, avoiding excessive rate reduction or resource waste. These state data can provide sufficient information basis for the model to make intelligent decisions based on actual use, rather than relying on static or simple rules. In addition, in the present embodiment, the multi-dimensional data (traffic usage, usage duration, signal strength) is integrated into a unified state vector, so that the model can consider multiple factors for decision-making. Through the form of the state vector, the model can comprehensively consider user behavior and network environment, avoiding decision-making bias caused by a single factor. In addition, the adjustment action of the data transmission rate of the personal WIFI corresponding to the state vector is defined, so that the model can flexibly adjust the rate in different states to adapt to the real-time demand of the user and the network change, improve the overall management efficiency, and make the decision space of the model clear, facilitating the training and optimization of the strategy. In addition, in the present embodiment, a reward function is defined, which evaluates the adjustment of the data transmission rate based on user satisfaction and resource utilization. The reward function consists of user satisfaction and resource utilization. User satisfaction can reflect user experience, such as whether the data transmission rate meets the user's demand, avoiding dissatisfaction caused by insufficient rate. Resource utilization can measure the use efficiency of operator resources, avoiding unnecessary bandwidth occupation caused by excessive rate configuration. The reward function considers both user satisfaction and resource utilization, ensuring that the model improves user experience while not wasting operator resources. Through the reward function, the model can clearly optimize the target and learn how to balance user demand and resource constraints in different states, so as to develop a reasonable strategy.Moreover, the reward function enables the adjustment of the strategy based on real-time feedback, allowing the model to continuously optimize in a changing environment. Additionally, in this embodiment, the Proximal Policy Optimization (PPO) algorithm is chosen as the reinforcement learning algorithm. PPO is a policy gradient method that offers good sample efficiency and stability, enabling effective policy optimization in complex high-dimensional spaces. It is also worth noting that through the state vector, the adjustment action of the data transmission rate, and the training of the reward function, the model can understand and predict the user's usage patterns and changes in the network environment. This understanding ability allows the model to flexibly adjust the data transmission rate in different situations, meeting user needs while optimizing resource utilization.

[0028] In further embodiments, when defining the adjustment action of the data transmission rate of the portable WIFI corresponding to the state vector, the method comprises: mapping the state vector to a selected adjustment level from a set of predefined data transmission rate adjustment levels according to a pre-set strategy. It is worth noting that by mapping the state vector to a predefined adjustment level, the number of possible actions can be reduced, making the decision-making process of the reinforcement learning model more concise, which helps to speed up the training process and improve the learning efficiency of the model. Moreover, the predefined adjustment levels can be adjusted according to different traffic packages and network resource allocation strategies to ensure efficient resource utilization and avoid resource waste or deficiency due to improper rate adjustment. For example, assume that at a certain time, the network load is high, and multiple users are connected to the portable WIFI. User A is in a high-traffic usage state (e.g., streaming high-definition video), and his state vector is mapped to a high-speed level, maintaining a high data transmission rate to meet his needs. User B is in a medium-traffic usage state, and his state vector is mapped to a medium-speed level, ensuring a reasonable data transmission rate while saving network resources. User C is in a low-traffic usage state, and his state vector is mapped to a low-speed level, reducing resource occupation to ensure more users can have stable connections. In this way, the predefined adjustment levels not only simplify the decision-making process of the model, but also flexibly and efficiently allocate data transmission rates according to the needs of different users and the actual situation of network resources, avoiding resource waste or insufficient rates, and improving the efficiency of overall network management and user experience.

[0029] In further embodiments, in defining the reward function, the user satisfaction is determined by measuring the extent to which the on-body WIFI satisfies the user's data transmission demand at the current data transmission rate; the resource utilization is determined by evaluating the resource usage efficiency of the communication network resources at the current data transmission rate; a user experience score is calculated according to the user satisfaction; a resource efficiency score is calculated according to the resource utilization; the user experience score and the resource efficiency score are weighted and aggregated to obtain a comprehensive reward value; and the comprehensive reward value is used to guide the reinforcement learning model to optimize the adjustment strategy of the data transmission rate in the training process to maximize the comprehensive reward value. It should be noted that by evaluating whether the on-body WIFI satisfies the user's data transmission demand at the current data transmission rate, it is ensured that the user experience is not affected. High user satisfaction means that the user's demand for network service is fully met, avoiding user dissatisfaction due to insufficient rate. By evaluating the usage efficiency of the communication network resources at the current data transmission rate, it is ensured that the network resources are used efficiently, avoiding resource waste due to excessively high rate configuration, or resource shortage due to excessively low rate configuration. Based on the quantification result of the user satisfaction, the quality of the user experience is evaluated separately, which can clearly indicate the user's satisfaction degree at the current rate. Based on the quantification result of the resource utilization, the efficiency of resource usage is evaluated separately, which can help the operator to understand the impact of the current configuration on resource management. By weighting and aggregating the user experience score and the resource efficiency score, the balance between user satisfaction and resource utilization efficiency can be achieved in the optimization process, allowing the weight to be adjusted according to the actual demand to adapt to different operation strategies and business objectives. The comprehensive reward value as the feedback signal of the reinforcement learning model can guide the model to optimize the user experience and resource management simultaneously in the training process, achieving the balance of dual objectives. In this embodiment, by taking the comprehensive reward value as the optimization target, the reinforcement learning model can learn to select the optimal rate adjustment action under different states to maximize the comprehensive evaluation of user satisfaction and resource utilization efficiency.

[0030] In some preferred embodiments, when the traffic usage record of the user's pocket WIFI is being counted, the traffic usage record of the user's pocket WIFI is counted by integrating a traffic monitoring module in the firmware of the pocket WIFI; or, when the traffic usage record of the user's pocket WIFI is being counted, the traffic usage record of the user's pocket WIFI is counted by the network interface module of the pocket WIFI. It should be noted that by integrating a traffic monitoring module at the firmware level, the underlying data transmission information of the device can be directly accessed, reducing data loss or errors. The monitoring module at the firmware level can capture and record the traffic usage in real time, ensuring the timeliness and accuracy of the data, which helps to adjust the management strategy in a timely manner. In addition, through the network interface module, the data packet information passing through the network interface can be directly obtained, realizing efficient traffic counting. The network interface module has efficient data processing capability and can complete the traffic counting task with low delay, improving the response speed of the overall system.

[0031] Referring to Figures 1-4 In some preferred embodiments, when the duration of each connection of the user's pocket WIFI is being counted, the following steps are included: when the user's device is connected or disconnected from the pocket WIFI, a connection event or a disconnection event is triggered; the connection event and the disconnection event are listened to, and when the connection event and the disconnection event are captured, the current accurate time is obtained, and the timestamp corresponding to the connection event and the timestamp corresponding to the disconnection event are generated; the difference between the connection time and the disconnection time of each connection of the user's device to the pocket WIFI is calculated according to the connection event, the timestamp corresponding to the connection event, the disconnection event and the timestamp corresponding to the disconnection event, and the duration of each connection of the user's pocket WIFI is obtained. Specifically, when the user's device is connected or disconnected from the pocket WIFI, the connection event or the disconnection event is triggered by the network interface module of the pocket WIFI. Specifically, the connection event and the disconnection event triggered by the network interface module are listened to by the event listener integrated in the firmware of the pocket WIFI, and when the connection event and the disconnection event are captured by the event listener, the current accurate time is obtained by calling the system time function, and the timestamp corresponding to the connection event and the timestamp corresponding to the disconnection event are generated. Specifically, the connection event and its corresponding timestamp delivered by the event listener are received by the session management module of the pocket WIFI, and the disconnection event and its corresponding timestamp delivered by the event listener are received, so as to calculate the difference between the connection time and the disconnection time of each connection of the user's device to the pocket WIFI, and obtain the duration of each connection of the user's pocket WIFI.

[0032] It should be noted that in this embodiment, when the user equipment is connected or disconnected with the pocket WIFI, the network interface module automatically triggers the corresponding event. Through the event listener integrated in the firmware, these events are captured in real time. When the event occurs, the system time function is called to generate the corresponding timestamp. By triggering the event immediately when connecting or disconnecting and recording the timestamp, the calculation of the continuous use duration has high real-time and accuracy, which helps to accurately reflect the actual use behavior of the user and avoid inaccurate data due to delay or error. In addition, when the event occurs, the event listener immediately records the event type (connection or disconnection) and its accurate time. The captured events and timestamps are passed to the session management module for subsequent time difference calculation. By separating event listening and session management, the modularity and maintainability of the pocket WIFI system are enhanced. The event listener focuses on capturing and recording events, while the session management module is responsible for data processing and calculation. Accurate continuous use duration data can help operators deeply understand the user's usage habits and behavior patterns, supporting subsequent traffic management and package recommendation. By understanding the actual use duration of the user, the operator can more effectively allocate network resources, optimize bandwidth usage, and improve overall network efficiency. The statistics of continuous use duration help to monitor the quality of service and timely discover and solve problems that may affect user experience, such as frequent disconnection or unstable network conditions.

[0033] Referring to Figures 1-4 In some preferred embodiments, when the signal strength of the pocket WIFI of the user equipment is counted, it includes: collecting the wireless signal strength data provided by the network operator's cellular base station in real time through the wireless receiving module of the pocket WIFI device to obtain the signal strength of the pocket WIFI of the user equipment. It should be noted that the wireless signal strength is an important indicator to measure the quality of network connection. By collecting signal strength data in real time, the current network environment changes such as signal weakening or strengthening, network congestion, etc. can be accurately understood. The network environment is dynamically changing, and real-time monitoring of signal strength can help the model learn and respond to environmental changes in a timely manner, adjust data transmission strategies, and ensure that users always have the best network experience.

[0034] In some preferred embodiments, the method further comprises: receiving the data transmission rate adapted to the current user's pocket WIFI according to the analysis of the reinforcement learning model; adjusting the data transmission parameters of the current user's pocket WIFI based on the received data transmission rate, wherein the data transmission parameters include upload rate and download rate; monitoring and recording the traffic usage of the current user's pocket WIFI in real time, wherein the traffic usage includes used traffic and remaining traffic; and dynamically updating the usage state of the traffic package according to the monitored and recorded traffic usage and the adjusted data transmission rate. It should be noted that the network environment and user demand are dynamically changing, for example, the user's usage mode is different at different time periods, and the network load also fluctuates accordingly. In this embodiment, based on the adapted data transmission rate output by the reinforcement learning model, the upload rate and the download rate are dynamically adjusted, which can optimize the data transmission in real time and ensure that the best network performance and user experience can be provided under various environments. In addition, adjusting the upload rate and the download rate to match the current data transmission demand of the user and the network resource use efficiency can not only improve the user experience, but also avoid resource waste, achieving a balance between the two. In addition, by monitoring the used traffic and the remaining traffic of the user in real time, the traffic consumption of the user can be accurately understood. According to the monitored traffic usage and the adjusted data transmission rate, the usage state of the traffic package is dynamically updated, which can ensure the reasonable allocation of resources.

[0035] In further embodiments, when intelligently recommending a traffic package for the current user of the portable WIFI based on the traffic usage state, the method comprises: analyzing the traffic usage state, the traffic usage state comprising used traffic, remaining traffic, and a usage pattern of the user, the usage pattern of the user comprising a peak period usage frequency and an application type (e.g. video streaming, online gaming, file downloading, etc.); and intelligently recommending a traffic package for the current user of the portable WIFI based on the used traffic, the remaining traffic, and the usage pattern of the user. It should be noted that by analyzing the used traffic and the remaining traffic, the traffic consumption of the user can be accurately grasped, and it can be identified whether the user is close to the upper limit of the traffic, and whether the user needs to upgrade the package or adjust the usage strategy. In addition, the peak period usage frequency and the application type (e.g. video streaming, online gaming, file downloading, etc.) reflect the specific needs and usage habits of the user. Different usage patterns have different demands for traffic packages. For example, video streaming requires high download speed and large traffic package, while online gaming focuses more on low latency and stability. In this embodiment, based on the specific usage and habits of the user, a suitable traffic package is recommended, which can ensure that the user obtains a good network experience in different usage scenarios, avoids interrupting the use of the user due to insufficient traffic, or paying unnecessary fees due to excessive traffic, and improves the overall user satisfaction. In addition, understanding the traffic usage and usage pattern of the user helps the operator to optimize the allocation of network resources. For example, during the peak period, the transmission rate and traffic allocation can be dynamically adjusted according to the actual needs of the user, to ensure that the network remains efficient under high load. By monitoring the remaining traffic and the usage habits of the user, unreasonable traffic usage can be identified and adjusted, and resource waste can be avoided.

[0036] In further embodiments, when intelligently recommending a traffic package that adapts to the current user's portable WIFI, according to the used traffic, the remaining traffic, and the user's usage pattern, it includes: predicting the user's future traffic demand according to the used traffic, the remaining traffic, and the user's usage pattern; matching the predicted user's future traffic demand with the preset multiple traffic package options to obtain a traffic package of portable WIFI recommended to the current user in descending order of matching degree. It should be noted that by analyzing the user's used traffic, remaining traffic, and usage pattern (such as peak period usage frequency and application type), the user's traffic demand in the future can be accurately predicted. This prediction helps to identify the user's possible traffic shortage or excess in advance, ensuring that the user can obtain sufficient traffic support when needed, while avoiding resource waste. Matching the predicted future traffic demand with the preset multiple traffic package options, the user can obtain the most suitable traffic package for himself, avoiding frequent upgrade of the package due to insufficient traffic or paying unnecessary fees due to excess traffic, optimizing the user's cost-effectiveness. Sorting the traffic package options in descending order of matching degree and generating multiple recommended schemes can provide users with flexible choices, not only improving the accuracy of recommendations, but also giving users more decision-making power to meet the individual needs of different users.

[0037] Embodiment two

[0038] Reference Figures 1-4 The embodiment provides a portable WIFI, which is managed by using the intelligent management method of the portable WIFI based on reinforcement learning in any of the above embodiments. The state representation data of the portable WIFI of the current user is monitored, the state representation data of the portable WIFI including the traffic usage record of the portable WIFI of the user, the continuous use duration of the user connecting the portable WIFI each time, and the signal strength of the portable WIFI of the user equipment. The state representation data of the portable WIFI is input into the trained reinforcement learning model, the data transmission rate that adapts to the portable WIFI of the current user is obtained by analyzing the reinforcement learning model, the traffic usage state of the traffic package of the portable WIFI of the current user is intelligently adjusted according to the data transmission rate, and the traffic package of the portable WIFI that adapts to the current user is intelligently recommended according to the traffic usage state, so as to optimize the data transmission rate and the traffic recommendation of the portable WIFI and improve the product use experience of the user.

[0039] It should be noted that the above embodiments are only preferred specific embodiments of the present application, and the protection scope of the present application is not limited thereto. Any changes or replacements within the technical range disclosed by the present application can be easily thought of by those skilled in the art, which should be covered within the protection scope of the present application. The protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A smart management method for portable Wi-Fi based on reinforcement learning, characterized in that, include: The status data of the user's portable Wi-Fi is monitored. The status data of the portable Wi-Fi includes the user's portable Wi-Fi data usage records, the duration of each connection to the portable Wi-Fi, and the signal strength of the user's portable Wi-Fi. The state representation data of the portable Wi-Fi is input into the trained reinforcement learning model, and the data transmission rate of the portable Wi-Fi adapted to the current user is obtained through the analysis of the reinforcement learning model. Based on the data transmission rate, the system intelligently adjusts the data usage status of the current user's portable Wi-Fi data plan. Specifically, this includes: receiving the data transmission rate of the portable Wi-Fi adapted to the current user's data transmission rate obtained from the reinforcement learning model, and adjusting the data transmission parameters of the current user's portable Wi-Fi based on the received data transmission rate, including upload and download speeds; monitoring and recording the data usage of the current user's portable Wi-Fi data plan in real time, including used and remaining data; dynamically updating the data plan usage status based on the real-time monitored and recorded data usage and the adjusted data transmission rate; and intelligently recommending a data plan adapted to the current user's portable Wi-Fi based on the data usage status. When recommending a data plan for a user's portable Wi-Fi, the process includes: analyzing the data usage status, which includes used data, remaining data, and the user's usage patterns, including peak-hour usage frequency and application types; intelligently recommending a data plan suitable for the current user's portable Wi-Fi based on the used data, remaining data, and the user's usage patterns; and further, predicting the user's future data needs based on the used data, remaining data, and the user's usage patterns; and matching the predicted future data needs with a range of preset data plan options to recommend portable Wi-Fi data plans to the current user based on the highest to lowest matching degree. Training the reinforcement learning model includes: The system acquires statistical data representing the status of the portable Wi-Fi, including the user's portable Wi-Fi data usage records, the duration of each connection, and the signal strength of the user's portable Wi-Fi. The process of calculating the duration of each connection includes: triggering a connection event or disconnection event when the user's device connects to or disconnects from the portable Wi-Fi; monitoring the connection and disconnection events; upon detection, obtaining the current precise time and generating timestamps for the connection and disconnection events; and calculating the difference between the connection and disconnection times for each connection and disconnection event to obtain the duration of each connection. The acquired state representation data of the portable WIFI is integrated into a state vector, and the adjustment action of the data transmission rate of the portable WIFI corresponding to the state vector is defined; when defining the adjustment action of the data transmission rate of the portable WIFI corresponding to the state vector, it includes: according to a pre-set strategy, mapping the state vector to a set of predefined data transmission rate adjustment levels and selecting an adjustment level. A reward function is defined, which evaluates the adjustment of the data transmission rate based on user satisfaction and resource utilization. This reward function guides the reinforcement learning model to optimize the data transmission rate adjustment strategy during training. When defining the reward function, user satisfaction is determined by measuring the degree to which the portable Wi-Fi meets the user's data transmission needs at the current data transmission rate. Resource utilization is determined by evaluating the resource usage efficiency of the communication network resources at the current data transmission rate. A user experience score is calculated based on the user satisfaction score. A resource efficiency score is calculated based on the resource utilization rate. The user experience score and the resource efficiency score are weighted and summed to obtain a comprehensive reward value. This comprehensive reward value is used to guide the reinforcement learning model to optimize the data transmission rate adjustment strategy during training to maximize the comprehensive reward value. The near-end policy optimization algorithm is selected as the reinforcement learning algorithm. The reinforcement learning algorithm is trained through the state vector, the data transmission rate adjustment action, and the reward function to obtain a reinforcement learning model that can analyze the current user's portable WIFI and adapt to the data transmission rate.

2. The intelligent management method for portable Wi-Fi based on reinforcement learning as described in claim 1, characterized in that, When tracking a user's portable Wi-Fi data usage, a data usage monitoring module can be integrated into the portable Wi-Fi firmware to track the user's data usage; alternatively, the user's portable Wi-Fi data usage can be tracked through the portable Wi-Fi's network interface module.

3. The intelligent management method for portable Wi-Fi based on reinforcement learning as described in claim 1, characterized in that, When a user device connects to or disconnects from the portable Wi-Fi, a connection event or disconnection event is triggered through the network interface module of the portable Wi-Fi.

4. The intelligent management method for portable Wi-Fi based on reinforcement learning as described in claim 3, characterized in that, The event listener integrated into the portable Wi-Fi firmware listens for connection and disconnection events triggered by the network interface module. When the event listener captures a connection or disconnection event, it calls the system time function to obtain the current precise time and generates a timestamp corresponding to the connection event and the timestamp corresponding to the disconnection event.

5. The intelligent management method for portable Wi-Fi based on reinforcement learning as described in claim 4, characterized in that, The session management module of the portable Wi-Fi receives connection events and their corresponding timestamps transmitted by the event listener, and receives disconnection events and their corresponding timestamps transmitted by the event listener, in order to calculate the difference between the connection time and disconnection time between the user device and the portable Wi-Fi each time, and obtain the continuous usage time of the user's connection to the portable Wi-Fi each time.

6. The intelligent management method for portable Wi-Fi based on reinforcement learning as described in claim 1, characterized in that, When calculating the signal strength of a user's portable Wi-Fi, the following steps are taken: real-time collection of wireless signal strength data provided by the network operator's cellular base station through the wireless receiving module of the portable Wi-Fi device to obtain the signal strength of the user's portable Wi-Fi.

7. A portable WiFi device, characterized in that, The portable Wi-Fi is managed using the intelligent management method for portable Wi-Fi based on reinforcement learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Power management method and device for portable 3G wireless internet device

    CN102026348A

  • Model training method, information recommendation method and device, medium and electronic equipment

    CN117196721A