Compressible Access Control Method Based on Q-Learning in Large-Scale Internet of Things

Through the compressible access control method based on Q learning, a two-dimensional Q value table is generated and time slots and measurement vector selection is optimized, which solves the problem of low throughput under high load of the satellite Internet of Things, and improves the success rate of node access and resource utilization.

CN115884428BActive Publication Date: 2025-07-08ANHUI TELECOMM PLANNING & DESIGNING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211506154.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-07-08
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing satellite IoT access protocol has low throughput under high load conditions and poor node access success rate, which cannot effectively maintain the communication needs of large-scale terminal nodes.

Method used

A compressible access control method based on Q learning is adopted to generate a two-dimensional Q value table by receiving satellite broadcast information, divide time slots according to the length value of the data frame, select odd or even frame access, and use the target measurement vector for compression measurement, update the Q function evaluation value to determine the final measurement vector and the transmission time slot, and optimize resource utilization.

Benefits of technology

It improves the data throughput and node access success rate of satellite IoT under high loads, reduces the low resource utilization rate and collision probability, and is suitable for IoT nodes with limited energy and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115884428B_ABST
    Figure CN115884428B_ABST
Patent Text Reader

Abstract

The present application relates to a compressible access control method based on Q-learning in a large-scale Internet of Things. Users maintain a two-dimensional Q-table and select time slots and measurement vectors through Q-learning, effectively reducing the problems caused by the randomness of users' selection of time slots and measurement vectors and improving resource utilization. The odd-even frame access scheme is used to ensure that the reception of feedback information matches the large time delay problem of satellite networks. The adaptive frame length adjustment control combining two-dimensional Q-learning and compressive sensing is used to solve the problems of low system resource utilization and poor flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things communication technologies, and particularly to a compressible access control method based on Q-learning in large-scale Internet of Things. Background Art

[0002] Satellite communication technology is a communication between two or more terminals that uses an artificial earth satellite as a relay station to forward radio waves. The satellite Internet of Things implemented based on satellite communication technology can provide communication access services for Internet of Things terminal nodes in remote areas without relying on ground-based communication facilities. The satellite Internet of Things is widely applied in fields such as natural resource surveillance and management, large-scale energy infrastructure monitoring, natural disaster monitoring and early warning, and military.

[0003] In application scenarios such as natural resource surveillance and management, large-scale energy infrastructure monitoring, and natural disaster monitoring and early warning, the services of the satellite Internet of Things have characteristics such as short packets, burstiness, and randomness. Each Internet of Things terminal node accesses a shared channel in an independent and competitive manner. Existing access schemes mainly include the slotted ALOHA (SA) protocol, the contention resolution diversity SA (CRDSA) protocol improved based on the SA protocol, and the irregular repetition SA (IRSA) protocol, etc. Among them, the peak throughput of the SA protocol is only about 0.36, and the peak throughput of the CRDSA and IRSA protocols is between 0.55 and 0.8. Existing access protocols randomly select time slots to send data packets, which may result in multiple data packets to be sent in some time slots and no data packets to be sent in some time slots. Therefore, in the case of high communication load, the data throughput of existing access protocols is low.

[0004] Moreover, the coverage area of the satellite Internet of Things is large, and the number of Internet of Things terminal nodes covered by an artificial earth satellite is often in the tens of thousands, with an extremely large communication load. Therefore, when facing the access requirements of a large number of terminal nodes under the coverage of the satellite Internet of Things, existing access protocols cannot maintain a high data throughput. Summary of the Invention

[0005] In view of the above-mentioned disadvantages of the prior art, this application provides a communication method and system to solve the technical problems of low throughput of existing access schemes and poor node access success rate.

[0006] To achieve the above object, this application provides a compressible access control method based on Q-learning in large-scale Internet of Things, including:

[0007] Receive the broadcast information broadcast by the satellite terminal, where the broadcast information includes the length value and serial number of the data frame, the Q-learning window length, and the measurement matrix, and the measurement matrix includes a plurality of measurement vectors;

[0008] Divide a plurality of transmission time slots according to the length value of the data frame, and generate a two-dimensional Q-value table, where the two-dimensional Q-value table includes the Q-function evaluation values between each measurement vector and each transmission time slot, and the initial value of each Q-function evaluation value is a preset value;

[0009] Divide the data frames into odd frames and even frames according to the serial number, and select to access either odd frames or even frames;

[0010] Determine the target measurement vector and the target transmission time slot in the data frame to be transmitted according to the multiple Q-function evaluation values included in the two-dimensional Q-value table;

[0011] Use the target measurement vector to perform compressive measurement on the first data packet to be transmitted to obtain a second data packet to be transmitted, and send the second data packet to be transmitted to the satellite terminal at the target time slot in the data frame to be transmitted;

[0012] Receive the feedback information for the second data packet to be transmitted broadcast by the satellite terminal, and update the Q-function evaluation value corresponding to the target measurement vector and the target transmission time slot according to the feedback information to obtain an updated two-dimensional Q-value table;

[0013] Determine the final measurement vector and the final transmission time slot according to the updated two-dimensional Q-value table;

[0014] Access the satellite Internet of Things using the final measurement vector and the final transmission time slot.

[0015] In an optional embodiment of the present application, determining the final measurement vector and the final transmission time slot according to the updated two-dimensional Q-value table specifically includes:

[0016] Judge whether the total length of the data frames that have been sent is less than the Q-learning window length;

[0017] If the judgment result is no, determine the final measurement vector and the final transmission time slot, where the final measurement vector and the final transmission time slot are the measurement vector and the transmission time slot corresponding to the maximum Q-function evaluation value in the updated two-dimensional Q-value table;

[0018] If the judgment result is yes, use the next-next data frame of the data frame to be transmitted as the data frame to be transmitted, and return to execute the step of determining the target measurement vector and the target transmission time slot in the data frame to be transmitted according to the multiple Q-function evaluation values included in the two-dimensional Q-value table until the total length of the data frames that have been sent is not less than the Q-learning window length.

[0019] In an alternative embodiment of the present application, determining the target measurement vector and the target transmission time slot in the data frame to be transmitted according to the multiple Q function evaluation values included in the two-dimensional Q value table includes:

[0020] Selecting a target Q function evaluation value from the multiple Q function evaluation values included in the two-dimensional Q value table, where the target Q function evaluation value is the maximum value among the Q function evaluation values in the two-dimensional Q value;

[0021] Taking the measurement vector and the transmission time slot corresponding to the target Q function evaluation value as the target measurement vector and the target transmission time slot in the data frame to be transmitted.

[0022] In an alternative embodiment of the present application, the feedback information is NACK or ACK, where NACK indicates that the decoding of the data frame to be transmitted fails, and ACK indicates that the decoding of the data frame to be transmitted is successful; updating the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot according to the feedback information to obtain an updated two-dimensional Q value table includes:

[0023] If the feedback information is NACK, then reducing the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot to obtain an updated two-dimensional Q value;

[0024] If the feedback information is ACK, then increasing the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot to obtain an updated two-dimensional Q value.

[0025] In an alternative embodiment of the present application, the following formula is used to increase or decrease the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot:

[0026]

[0027] where, Q i (j) (t, m) represents the Q function evaluation value corresponding to the updated target measurement vector and the target transmission time slot, represents the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot before update, α represents the learning rate, and r represents the reward and punishment factor.

[0028] In an alternative embodiment of the present application, using the target measurement vector to perform compressive measurement on the first data packet to be transmitted to obtain a second data packet to be transmitted, which is implemented by the following formula:

[0029]

[0030] where, is the second data packet to be transmitted, is the first data packet to be transmitted, is the target measurement vector.

[0031] In an alternative embodiment of the present application, after determining the final measurement vector and the final transmission time slot, the method further includes:

[0032] Obtain the length value of the updated data frame broadcast by the satellite side, and return to execute the step of dividing multiple transmission time slots according to the length value of the data frame and generating a two-dimensional Q-value table until the load of the entire system is stable.

[0033] In an alternative embodiment of the present application, the length value of the data frame is updated in the following manner:

[0034] Take the total number of elements of all time slot support sets in the last data frame within the Q-learning window as the estimated number of active nodes, and determine whether the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes;

[0035] If the estimated number of active nodes is greater than the upper limit value of the number of accessible nodes, increase the data frame length;

[0036] If the estimated number of active nodes is less than the lower limit value of the number of accessible nodes, decrease the data frame length;

[0037] If the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes, the load is stable.

[0038] In an alternative embodiment of the present application, the upper limit value of the accessible nodes is calculated by the following formula:

[0039] S up = G * · N · T;

[0040] The lower limit value of the accessible nodes is calculated by the following formula:

[0041] S down = (G * - Δ s ) · (N - 1) · T;

[0042] Wherein, G * is the load threshold, N is the number of rows of the measurement matrix, T is the length value of the data frame, and Δ s is the adjustment buffer interval.

[0043] In an alternative embodiment of the present application, the length value of the data frame is increased by the following formula:

[0044] T = 2T;

[0045] The length value of the data frame is decreased by the following formula:

[0046]

[0047] Among them, round() is a floor function, T is the length value of the data frame, S is the estimated number of active nodes, and N is the number of rows of the measurement matrix.

[0048] Advantages of this application:

[0049] This application provides a compressible access control method based on Q-learning in a large-scale Internet of Things. The main application scenario is a satellite Internet of Things with large-scale node access. Nodes receive broadcast information broadcast by the satellite side, and the broadcast information includes the length value and sequence number of the data frame, the Q-learning window length, and the measurement matrix. The measurement matrix includes multiple measurement vectors; multiple transmission time slots are divided according to the length value of the data frame, and a two-dimensional Q-value table is generated. The two-dimensional Q-value table includes the Q-function evaluation values between each measurement vector and each transmission time slot, and the initial value of each Q-function evaluation value is a preset value; the data frame is divided into odd frames and even frames according to the sequence number, and an odd frame or an even frame is selected for access; according to the multiple Q-function evaluation values included in the two-dimensional Q-value table, the target measurement vector and the target transmission time slot in the data frame to be transmitted are determined; the first data packet to be transmitted is compressed and measured by using the target measurement vector to obtain a second data packet to be transmitted, and the second data packet to be transmitted is sent to the satellite side in the target time slot in the data frame to be transmitted; the feedback information for the second data packet to be transmitted broadcast by the satellite side is received, and the Q-function evaluation value corresponding to the target measurement vector and the target transmission time slot is updated according to the feedback information to obtain an updated two-dimensional Q-value table; the final measurement vector and the final transmission time slot are determined according to the updated two-dimensional Q-value table; the satellite Internet of Things is accessed by using the final measurement vector and the final transmission time slot. It combines two-dimensional Q-learning and compressive sensing for node access control. Through two-dimensional Q-learning technology, the time slots and measurement vectors of the access nodes are learned, the best occupied time slots and measurement vectors are selected, the resource utilization rate is optimized, and the throughput is improved.

[0050] For the compressible access control method based on Q-learning in the large-scale Internet of Things of this application, the time slots and measurement vectors occupied during node access are determined by the way of initial randomness and then autonomous learning, without frequent interaction between nodes and resource allocation. The access method is simple, avoiding the collision retransmission of traditional random access protocols, and is suitable for Internet of Things nodes with limited energy and computing resources.

[0051] This application also obtains the number of elements in the support set as an estimation index of network load by means of support set estimation in the compressive sensing reconstruction process, which is used to adjust the length of a frame of signal, improve the resource utilization rate and system flexibility, thereby improving the node access success rate and maintaining a high node access success rate under high load. Description of the Drawings

[0052] Figure 1 This is the flowchart of the compressible access control method based on Q - learning in the large - scale Internet of Things of this application.

[0053] Figure 2 This is an exemplary schematic diagram of the two - dimensional Q - value table adopted by this application.

[0054] Figure 3 This is the schematic diagram of satellite delay adopted by this application.

[0055] Figure 4 This is the schematic diagram of the update of the Q - evaluation value of this application.

[0056] Figure 5 This is the comparison chart of load - throughput performance between the solution of this application and other traditional solutions.

[0057] Figure 6 This is the comparison curve graph of the estimated number of users and the actual number of users of this application.

[0058] Figure 7 This is the overall flowchart of the compressible access control method based on Q - learning in the large - scale Internet of Things of a specific example of this application.

[0059] Figure 8 This is the comparison chart of load - throughput performance under different numbers of measurement vectors of this application.

[0060] Figure 9 This is the graph of frame - length change under different numbers of active nodes of this application.

[0061] Figure 10 This is the performance graph of measurement - vector utilization under different numbers of active nodes of this application.

[0062] Figure 11 This is the performance graph of access success rate under different numbers of active nodes of this application. Detailed implementation manners

[0063] The following uses specific specific examples to illustrate the implementation manners of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0064] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0065] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.

[0066] In the description of the embodiments of the present disclosure, terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0067] Unless otherwise specified, the term "plurality" means two or more.

[0068] In the embodiments of the present disclosure, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.

[0069] The term "and / or" is an associative relationship describing objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.

[0070] In order to increase the data throughput in the satellite Internet of Things under high load conditions, the embodiments of the present application provide a compressible access control method based on Q learning in a large-scale Internet of Things, as Figure 1 shown, the method includes:

[0071] S101. Receive the broadcast information broadcast by the satellite side.

[0072] In the embodiments of the present application, the satellite side can broadcast to the active Internet of Things terminal nodes within its coverage area during the over-the-top time. Among them, the broadcast information includes the length value and serial number of the data frame, the Q learning window length, and the measurement matrix, and the measurement matrix includes a plurality of measurement vectors.

[0073] The Q learning window length of the satellite side and the length value of the data frame can be preset according to the actual application scenario, and the Q learning window length is greater than the length value of the data frame.

[0074] The measurement matrix is randomly generated. The dimension of the measurement matrix is N*M, and the elements in the measurement matrix follow a Gaussian distribution. Among them, the number of rows N can be set according to the number of OFDM subcarriers, and the number of columns M is the number of measurement vectors included in the measurement matrix. The value of M can be set according to the storage resources of the IoT terminal node.

[0075] For example, the measurement matrix can be expressed as Ψ=[ψ1,ψ2,...,ψ M , where each column vector is a measurement vector, and the measurement vector can be expressed as ψ m ∈C N×1 , m = 1, 2,..., M, and the length is N.

[0076] S102. Divide multiple transmission time slots according to the length value of the data frame, and generate a two-dimensional Q value table.

[0077] Among them, the two-dimensional Q value table includes the Q function evaluation values between each measurement vector and each transmission time slot. The initial value of each Q function evaluation value is a preset value, and the preset value can be set according to the actual application scenario. For example, the preset value can be set to 0.

[0078] Assume that the length value of the data frame is set to T time slots, then T transmission time slots can be divided.

[0079] As Figure 2 shown, Figure 2 is an exemplary schematic diagram of a two-dimensional Q value table provided by an embodiment of the present application. Figure 2 The initial value of each Q function evaluation value in the two-dimensional Q value table in Figure 2 is the preset value 0. This two-dimensional Q value table can be regarded as a matrix, where the number of rows corresponds to the measurement vector serial number, and the number of columns corresponds to the serial number of the transmission time slot.

[0080] S103. Divide the data frame into odd frames and even frames according to the serial number, and select to access the odd frames or the even frames.

[0081] Among them, the odd frame is the data frame with an odd serial number, and the even frame is the data frame with an odd serial number.

[0082] S104. Determine the target measurement vector and the target transmission time slot in the data frame to be sent according to the multiple Q function evaluation values included in the two-dimensional Q value table.

[0083] The specific method for determining the target measurement vector and the target transmission time slot of the data frame to be sent will be introduced below.

[0084] S105. Use the target measurement vector to perform compressive measurement on the first data packet to be sent, so as to obtain the second data packet to be sent, and send the second data packet to the satellite terminal in the target time slot within the data frame to be sent.

[0085] In the embodiment of the present application, the Internet of Things terminal node needs to perform channel coding and digital modulation processing on the transmission data to be sent first. The processed data symbol can be expressed as That is, the first data packet to be sent, where i is the node number, t is the sequence number of the target transmission time slot within the data frame to be sent, L is the length of the data symbol, and the value of L can be determined according to the transmission rate of the satellite Internet of Things system.

[0086] Assume that the target measurement vector is It represents the m-th column measurement vector selected by any Internet of Things terminal node i from the known measurement matrix Ψ = [ψ1, ψ2,..., ψ M , where m = 1, 2,..., M, and M is the number of columns of the measurement matrix Ψ. Then the compressive measurement process can be expressed as: That is, after any Internet of Things terminal node i performs channel coding and digital modulation processing on the data packet to be sent, it multiplies with the target measurement vector, where is the second data packet to be sent, is the first data packet to be sent, is the target measurement vector.

[0087] The data packet to be sent after compressive measurement accesses the channel according to the target transmission time slot t, and sends the second data packet to be sent after compressive measurement through N OFDM subcarriers. It is forwarded by the satellite terminal to the ground receiving end. For a frame signal containing T time slots, the received signal Y at any t-th time slot in a frame t is expressed as:

[0088]

[0089] Among them, represents the data symbol of the first data packet to be sent after being modulated by node i at the transmission time slot t. For the nodes in the non-active state in the data frame to be sent, its corresponds to 0, and H i,t ∈C N×N is the channel coefficient matrix of node i at the target transmission time slot t, and n i, is the additive white Gaussian noise during the transmission of node i at the target transmission time slot t.

[0090] At this time, the sparsity K in a single transmission time slot is equal to the number of measurement vectors used within a single transmission time slot, and is less than or equal to the estimated number of active nodes in the current transmission time slot. According to the theory of compressive sensing technology, N << I, that is, N is much smaller than the total number I of IoT terminal nodes, and multiple IoT terminal nodes select the same target transmission time slot t for data transmission. At the receiving end, using the compressive sensing reconstruction algorithm (abbreviation: CS reconstruction algorithm), the identities of the IoT terminal nodes can be detected and the second data packet to be transmitted can be recovered.

[0091] S106. Receive the feedback information for the second data packet to be transmitted broadcast by the satellite side, and update the Q function evaluation value corresponding to the target measurement vector and the target transmission time slot according to the feedback information, so as to obtain an updated two-dimensional Q value table.

[0092] Among them, the feedback information is used to indicate whether the receiving end successfully decodes the data packet to be transmitted.

[0093] In the above S103, the data frame is divided into odd frames and even frames, and the odd-even frame strategy is to adapt to the large delay of the satellite.

[0094] Taking the odd frame as an example, as Figure 3 shown, Figure 3 frames 1 and 3 are both odd frames. If the propagation delay of the uplink from the IoT terminal node to the satellite side is T f , and the propagation delay from the satellite side to the ground receiving end is T b , then when the IoT terminal node sends out the data packet to be transmitted, the minimum delay until receiving the feedback information for the data packet to be transmitted sent by the ground receiving end is 2(T f + T b ). Assuming that the data frame length is T F , in order for the feedback information of frame 1 to be fed back to the IoT terminal node before frame 3 starts to be transmitted, then T F > 2(T f + T b ).

[0095] S107. Determine the final measurement vector and the final transmission time slot according to the updated two-dimensional Q value table.

[0096] Specifically, determine whether the total length of the transmitted data frames is less than the Q learning window length.

[0097] If the judgment result is negative, determine the final measurement vector and the final transmission time slot, where the final measurement vector and the final transmission time slot are the measurement vector and the transmission time slot corresponding to the maximum Q function evaluation value in the updated two-dimensional Q value table.

[0098] If the judgment result is yes, use the next-next data frame of the to-be-sent data frame as the to-be-sent data frame, and return to execute the step of determining the target measurement vector and the target transmission time slot in the to-be-sent data frame according to the two-dimensional Q value table including multiple Q function evaluation values, until the total length of the already-sent data frames is not less than the Q learning window length, so as to obtain the final measurement vector and the final transmission time slot.

[0099] S108. Access the satellite Internet of Things by using the final measurement vector and the final transmission time slot.

[0100] By adopting the embodiment of the present application, by receiving the broadcast information broadcast by the satellite end, the Internet of Things terminal node generates a two-dimensional Q value table through the broadcast information, determines the target measurement vector and the target transmission time slot according to the two-dimensional Q value table, and then compresses and measures the to-be-sent data packet by using the two-dimensional Q value table, sends the to-be-sent data packet within the target transmission time slot, updates the Q function evaluation value between the target measurement vector and the target transmission time slot according to the feedback information, and repeatedly updates the two-dimensional Q value table when the length of the already-sent data frame is less than the Q learning window length until the Q learning converges. In this way, the final measurement vector and the final transmission time slot determined by the Internet of Things terminal node can be the optimal measurement vector and the optimal transmission time slot, which can reduce the low resource utilization rate caused by the random selection of the measurement vector and the transmission time slot by the Internet of Things terminal node, effectively reduce the probability of collision of multiple data packets in the same transmission time slot, and improve the data throughput of the satellite Internet of Things under high load conditions.

[0101] In another embodiment of the present application, the above S104. Determining the target measurement vector and the target transmission time slot in the to-be-sent data frame according to the multiple Q function evaluation values included in the two-dimensional Q value table can be specifically implemented as follows:

[0102] Select a target Q function evaluation value from the multiple Q function evaluation values included in the two-dimensional Q value table, where the target Q function evaluation value is the maximum value among the Q function evaluation values in the two-dimensional Q value;

[0103] Use the measurement vector and the transmission time slot corresponding to the target Q function evaluation value as the target measurement vector and the target transmission time slot in the to-be-sent data frame.

[0104] In the embodiment of the present application, there may be multiple Q function evaluation values that are the maximum value. At this time, it is necessary to randomly select one Q function evaluation value from the multiple Q function evaluation values that are the maximum as the target Q function evaluation value.

[0105] In addition, when initially selecting the target Q function evaluation value, since the Q function evaluation values in the two-dimensional Q value table are all initial values, at this time, a Q function evaluation value can be randomly selected as the target Q function evaluation value.

[0106] Through the above operations, the Q - function evaluation value selected as the maximum value is used as the target Q - function evaluation value, and the measurement vector and the transmission time slot corresponding to the target Q - function evaluation value are used as the target measurement vector and the target transmission time limit, which can increase the transmission success rate of the data packet to be sent.

[0107] In another embodiment of the present application, the above feedback information is Negative Acknowledgement (NACK) or Acknowledgement (ACK). Among them, NACK indicates that the decoding of the data frame to be sent fails, and ACK indicates that the decoding of the data frame to be sent fails. The specific implementation of updating the Q - function evaluation value between the target measurement vector and the target transmission time slot according to the feedback information can be as follows:

[0108] If the feedback information is NACK, then reduce the Q - function evaluation value between the target measurement vector and the target transmission time slot; if the feedback information is ACK, then increase the Q - function evaluation value between the target measurement vector and the target transmission time slot.

[0109] In the embodiment of the present application, after the receiving end receives the data packet to be sent forwarded by the satellite end, it will, according to the time - slot sequence in the data frame, use the compressive sensing reconstruction algorithm to reconstruct and decode the received data packet to be sent, and recover the information of all Internet of Things terminal nodes slot by slot. In the case of successful decoding, the feedback information sent to the corresponding Internet of Things terminal node is ACK, and in the case of decoding failure, the feedback information sent to the corresponding Internet of Things terminal node is NACK.

[0110] If the feedback information is NACK, it means that the target measurement vector and the target transmission time slot used by the Internet of Things terminal node will cause the data packet to be sent to fail. At this time, it is necessary to reduce the Q - function value evaluation function between the target measurement vector and the target transmission time slot, so as to obtain the updated two - dimensional Q value, so that the Internet of Things terminal can re - select the appropriate target measurement vector and target transmission time slot in the next - next data frame.

[0111] If the feedback information is ACK, it means that the target measurement vector and the target transmission time slot used by the Internet of Things terminal node can successfully send the data packet to be sent. At this time, it is necessary to increase the Q - function value evaluation function between the target measurement vector and the target transmission time slot to obtain the updated two - dimensional Q value, so that the Internet of Things terminal can continue to select the target measurement vector and the target transmission time slot in the next - next data frame.

[0112] Specifically, the following formula can be used to update the Q - function evaluation value between the target measurement vector and the target transmission time slot:

[0113]

[0114] Among them, represents the Q - function evaluation value between the target measurement vector and the target transmission time slot before the update, α represents the learning rate, r represents the reward - punishment factor. When the feedback information is ACK, r takes the value of 1; when the feedback information is NACK, r takes the value of - 1.

[0115] In this way, the IoT terminal node can continuously increase the Q - function evaluation value between the measurement vector that can successfully send data packets and the transmission time slot, and continuously decrease the Q - function evaluation value between the measurement vector that may cause data packet transmission failure and the transmission time slot, so that the IoT terminal node can select the best measurement vector and transmission time slot according to the size of the Q - function evaluation value, improving the data throughput of the satellite IoT.

[0116] The following combines Figure 4 to illustrate the embodiments of the present application. Figure 4 Taking nodes 1 and 2 selecting the target measurement vector and the target time slot to send data packets as an example.

[0117] In odd - numbered frames: Node 1 selects the measurement vector 1 and the transmission time slot 2 to send data packets in the first frame. If the transmission is successful, the Q - function evaluation value corresponding to the measurement vector 1 and the transmission time slot 2 increases by 0.1. In the third frame, 0.1 is the maximum value, and node 1 will continue to select the measurement vector and the transmission time slot corresponding to 0.1 to send data packets in the third frame. Node 2 selects the measurement vector 2N and the time slot 2 to send data packets in the first frame. The transmission is successful, and the Q - function evaluation value corresponding to the measurement vector 2N and the transmission time slot 2 increases by 0.1. In the third frame, 0.1 is the maximum value, and node 2 will continue to select the measurement vector and the transmission time slot corresponding to 0.1 to send data packets.

[0118] In even - numbered frames: Node 1 selects the measurement vector 1 and the transmission time slot 2 to send data packets in the second frame. If the transmission fails, the Q - function evaluation value corresponding to the measurement vector 1 and the transmission time slot 2 decreases by 0.1. In the third frame, - 0.1 is not the maximum value, and node 1 will not continue to select the measurement vector and the time slot corresponding to - 0.1 to send data packets in the fourth frame. Node 2 selects the measurement vector 2N and the transmission time slot 1 to send data packets in the second frame. The transmission is successful, and the Q - function evaluation value corresponding to the measurement vector 2N and the transmission time slot 1 increases by 0.1. In the fourth frame, 0.1 is the maximum value, and node 2 will continue to select the measurement vector and the transmission time slot corresponding to 0.1 to send data packets.

[0119] After determining the final measurement vector and the final transmission time slot in S107 above, if the current entire low-orbit satellite IoT access system is in an underloaded or overloaded state, the satellite side will re-broadcast the updated data frame length value to the IoT terminal nodes covered by it. At this time, the IoT terminal nodes will re-obtain the updated data frame length value broadcast by the satellite side, and return to execute the steps of dividing multiple transmission time slots according to the data frame length value and generating a two-dimensional Q-value table until the system load stabilizes.

[0120] In the above embodiments, the data frame length value is updated in the following manner:

[0121] Take the total number of elements in the support sets of all time slots in the last data frame within the Q-learning window as the estimated number of active nodes, and the ground station determines whether the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes.

[0122] If the estimated number of active nodes is greater than the upper limit value of the number of accessible nodes, the system is overloaded, and the data frame length is increased.

[0123] If the estimated number of active nodes is less than the lower limit value of the number of accessible nodes, the system is underloaded, and the data frame length is decreased.

[0124] If the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes, it proves that the estimated number of active nodes is between the upper and lower limits at this time. While ensuring a high access success rate, the transmission resources are also ideally utilized, which can be used as a termination condition for this frame length adjustment process. When this condition is reached, the frame length adjustment process stops, and it is determined that the system load is stable.

[0125] In the embodiments of the present application, when the receiving end reconstructs the data packet, it will obtain the total sum of the element data of the support sets of all time slots in the data frame to be transmitted. According to the compressive sensing theory, the number of elements in the support set is equal to the sparsity K, and the sum of the number of support sets of all time slots is the estimated number of active nodes of the data frame to be transmitted. The estimated number of active users adopted in the embodiments of the present application is the estimated number of active nodes of the last data frame within the Q-learning window.

[0126] With the continuous access of IoT terminal nodes, it will cause the normalized load of the system to increase. When the normalized load is greater than the preset load threshold, the sparsity of each transmission time slot will exceed the tolerance of the system, resulting in system overload. At this time, the existing transmission resources cannot meet the access requirements of all IoT terminal nodes, and it is impossible to achieve the unique allocation of resources, resulting in the Q-learning never reaching the convergence state.

[0127] Figure 5 Shows the throughput performance of the technical solution (QSA-RABCS) of the present application as Figure 4As shown, the hexagram is the QSA-RABCS scheme of this application with 100 measurement vectors, the rhombus is the traditional SA scheme, the pentagram is the traditional QSA scheme, and the circle is the throughput effect diagram of the random access scheme (RABCS) based on compressive sensing CS and slotted ALOHA. It can be seen that the throughput of the QSA-RABCS scheme of this application is higher than that of the traditional SA, slotted ALOHA based on Q-learning (QSA), and RABCS. From Figure 5 It can be seen that the normalized load when the QSA-RABCS system of this application maintains a high throughput is G * = 1.1. And there is Figure 6 It can be known that when the actual number of nodes does not exceed 1100 (the normalized load G is 1.1), the number of nodes estimated by QSARABCS of this application is the same as the actual number of nodes.

[0128] The specific normalized load is calculated by the following formula:

[0129]

[0130] Among them, λ represents the average packet arrival rate in a data frame, with the unit of packet / slot, r is the coding rate, M1 is the modulation order, the normalized load G represents the average number of packets transmitted on a single subcarrier per time slot, with the unit of bits / symbol / carrier, and the normalized throughput T represents the average number of successfully decoded packets on a single subcarrier per time slot, with the same unit as G.

[0131] In order to ensure that each data frame can be in the best convergence state, a preset load threshold can be set. According to the preset load threshold, the upper limit value of the number of accessible nodes can be calculated. At the same time, in order to ensure the utilization rate of the measurement vector, the lower limit value of the number of accessible nodes can also be calculated according to the preset load threshold.

[0132] Specifically, the upper limit value of the above-mentioned number of accessible nodes is calculated by the following formula:

[0133] S up = G * ·N·T;

[0134] The lower limit value of the number of accessible nodes is calculated by the following formula:

[0135] S down =(G * -Δ s )·(N - 1)·T;

[0136] Among them, G *is the load threshold, N is the number of rows of the measurement matrix, determined according to the number of OFDM subcarriers, T is the length value of the data frame, that is, the number of time slots, Δ s is the adjustment buffer interval to ensure that the system operates near the load threshold.

[0137] Specifically, the length value of the data frame is increased by the following formula:

[0138] T = 2T.

[0139] The length value of the data frame is decreased by the following formula:

[0140]

[0141] where T is the length value of the data frame, S is the estimated number of active nodes, and N is the number of rows of the measurement matrix.

[0142] When the estimated number of active nodes S is between the upper limit and the lower limit, that is, S down ≤ S ≤ S up , it indicates that the access success rate of the system is relatively high at this time, and the utilization rate of the transmission resources is high.

[0143] As Figure 7 shown, Figure 7 is a flowchart for adjusting the length of the data frame provided by an embodiment of the present application.

[0144] In the prior art, the length value of the data frame in the RABCS scheme based on two-dimensional Q learning is fixed, the system lacks flexibility, and the resource utilization rate is low. However, the compressible access control method based on Q learning provided by the present application can flexibly adjust the length value of the data frame according to the system load, increasing the flexibility of the system and improving the utilization rate of the transmission resources.

[0145] Taking the satellite side covering 100 to 2000 Internet of Things terminal nodes, the initial length value of the data frame is T = 20, the number of subcarriers N = 50 (the normalized load is 0 to 2), and the learning window length L of Q learning Q = 10, and the number of measurement vectors is M = 100 as an example.

[0146] As Figure 8 shown, in the case of M = 100, the throughput improvement of RABCS is the largest. Due to the relatively small total number of measurement vectors, the problem of transmission resource collision is more serious compared to other M values. However, by using the Q learning scheme to continuously select and learn the optimal transmission strategy, the relatively serious problem of transmission resource collision caused by random selection can be significantly improved.

[0147] Figure 9The compressible access control method based on Q - learning proposed in this application, for the optimal frame lengths corresponding to different numbers of access nodes (100 - 1000 corresponding to the normalized load from 0.1 to 1), from Figure 9 it can be seen that the frame length changes dynamically with the load and is always less than the initially set frame length of 20. An appropriate frame length can improve the utilization rate of the measurement vector, and at the same time, the reduction of the frame length also reduces the processing delay at the receiving end to a certain extent.

[0148] The compressible access control scheme based on Q - learning proposed in this application is Figure 10 and Figure 11 marked as ADFL in Figure 10 and Figure 11 respectively compares the performance of the utilization rate of the measurement vector and the access success rate with the QSA - RABCS scheme that only uses Q - learning without frame length control. It can be seen from the figure that the ADFL scheme of this application has a higher utilization rate of the measurement vector and can maintain a high access success rate when a large number of nodes access.

[0149] The above - mentioned embodiments only illustratively explain the principles and effects of this application, rather than limiting this application. Any person familiar with this technology can modify or change the above - mentioned embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.

[0150] In the description herein, many specific details are provided, such as examples of components and / or methods, to provide a complete understanding of the embodiments of this application. However, those skilled in the art will recognize that the embodiments of this application can be practiced without one or more of the specific details or by other devices, systems, components, methods, parts, materials, parts, etc. In other cases, well - known structures, materials, or operations are not specifically shown or described in detail to avoid obscuring aspects of the embodiments of this application.

[0151] Throughout the specification, the mention of "an embodiment", "embodiment" or "specific embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of this application and not necessarily in all embodiments. Thus, the appearances of the phrases "in an embodiment", "in an embodiment" or "in a specific embodiment" in different places throughout the specification are not necessarily referring to the same embodiment. In addition, the specific features, structures, or characteristics of any specific embodiment of this application can be combined with one or more other embodiments in any suitable manner. It should be understood that other variations and modifications of the invention embodiments described and shown herein may be according to the teachings herein and will be regarded as part of the spirit and scope of this application.

[0152] It should also be understood that one or more of the elements shown in the figures can also be implemented in a more separated or more integrated manner, or even removed because they cannot be operated in some cases or provided because they can be useful according to a particular application.

[0153] In addition, unless otherwise expressly specified, any reference signs in the figures should be regarded only as exemplary and not restrictive. Furthermore, unless otherwise indicated, the term "or" as used herein generally intends to mean "and / or". In cases where the term is foreseen to be unclear due to the ability to provide separation or combination, the combination of components or steps will also be regarded as having been specified.

[0154] As used in the description herein and throughout the claims below, unless otherwise indicated, the singular forms "a", "an", and "the" include plural referents. Also, as used in the description herein and throughout the claims below, unless otherwise indicated, the meaning of "in" includes "in" and "on".

[0155] The foregoing description of the embodiments shown in this application (including what is described in the abstract) is not intended to be exhaustive or to limit the application to the precise forms disclosed herein. Although specific embodiments of the application and examples of the application have been described herein for illustrative purposes only, various equivalent modifications will be within the spirit and scope of the application as will be recognized and understood by those skilled in the art. As noted, these modifications can be made to the application in accordance with the foregoing description of the embodiments of the application, and these modifications will be within the spirit and scope of the application.

[0156] The systems and methods have been generally described herein to assist in understanding the details of the application. In addition, various specific details have been given to provide an overall understanding of the embodiments of the application. However, those skilled in the relevant art will recognize that the embodiments of the application can be practiced without one or more of the specific details, or with other devices, systems, components, methods, assemblies, materials, parts, etc. In other instances, well-known structures, materials, and / or operations have not been particularly shown or described in detail to avoid obscuring aspects of the embodiments of the application.

[0157] Accordingly, while the present application has been described herein with reference to its specific embodiments, modifications, various changes and substitutions are also within the above disclosure, and it should be understood that in some cases, some features of the present application will be employed without corresponding use of other features without departing from the scope and spirit of the claimed invention. Therefore, many modifications may be made to adapt a particular environment or material to the essential scope and spirit of the present application. The present application is not intended to be limited to the specific terms used in the following claims and / or the specific embodiments disclosed as the best mode contemplated for carrying out the present application, but the present application will include any and all embodiments and equivalents falling within the scope of the appended claims. Accordingly, the scope of the present application will be determined only by the appended claims.

Claims

1. A compressible access control method based on Q-learning in a large-scale Internet of Things, characterized in that The method includes: Receiving broadcast information broadcast by a satellite terminal, where the broadcast information includes the length value and sequence number of a data frame, the Q-learning window length, and a measurement matrix, and the measurement matrix includes a plurality of measurement vectors; Dividing a plurality of transmission time slots according to the length value of the data frame, and generating a two-dimensional Q-value table, where the two-dimensional Q-value table includes Q-function evaluation values between each measurement vector and each transmission time slot, and the initial value of each Q-function evaluation value is a preset value; Dividing the data frames into odd frames and even frames according to the sequence number, and selecting either odd frames or even frames for access; Determining a target measurement vector and a target transmission time slot in the data frame to be transmitted according to the plurality of Q-function evaluation values included in the two-dimensional Q-value table; Performing compressive measurement on a first data packet to be transmitted by using the target measurement vector to obtain a second data packet to be transmitted, and transmitting the second data packet to be transmitted to the satellite terminal at the target time slot in the data frame to be transmitted; Receiving feedback information for the second data packet to be transmitted broadcast by the satellite terminal, and updating the Q-function evaluation value corresponding to the target measurement vector and the target transmission time slot according to the feedback information to obtain an updated two-dimensional Q-value table; Determining a final measurement vector and a final transmission time slot according to the updated two-dimensional Q-value table; Accessing the satellite Internet of Things by using the final measurement vector and the final transmission time slot.

2. The method according to claim 1, wherein Determining a final measurement vector and a final transmission time slot according to the updated two-dimensional Q-value table specifically includes: Judging whether the total length of the already transmitted data frames is less than the Q-learning window length; If the judgment result is negative, determining the final measurement vector and the final transmission time slot, where the final measurement vector and the final transmission time slot are the measurement vector and the transmission time slot corresponding to the maximum Q-function evaluation value in the updated two-dimensional Q-value table; If the judgment result is positive, using the next-next data frame of the data frame to be transmitted as the data frame to be transmitted, and returning to execute the step of determining a target measurement vector and a target transmission time slot in the data frame to be transmitted according to the plurality of Q-function evaluation values included in the two-dimensional Q-value table until the total length of the already transmitted data frames is not less than the Q-learning window length.

3. The method according to claim 1, wherein The determining a target measurement vector and a target transmission time slot in the data frame to be transmitted according to the plurality of Q-function evaluation values included in the two-dimensional Q-value table includes: Selecting a target Q-function evaluation value from the plurality of Q-function evaluation values included in the two-dimensional Q-value table, where the target Q-function evaluation value is the maximum value among the Q-function evaluation values in the two-dimensional Q; Using the measurement vector and the transmission time slot corresponding to the target Q-function evaluation value as the target measurement vector and the target transmission time slot in the data frame to be transmitted.

4. The method according to claim 1, wherein The feedback information is NACK or ACK, where NACK indicates that the decoding of the data frame to be transmitted fails, and ACK indicates that the decoding of the data frame to be transmitted is successful; The updating the Q-function evaluation value corresponding to the target measurement vector and the target transmission time slot according to the feedback information to obtain an updated two-dimensional Q-value table includes: If the feedback information is NACK, decrease the Q - function evaluation value corresponding to the target measurement vector and the target transmission time slot to obtain the updated two - dimensional Q - value; If the feedback information is ACK, increase the Q - function evaluation value corresponding to the target measurement vector and the target transmission time slot to obtain the updated two - dimensional Q - value.

5. The method according to claim 4, characterized in that, Use the following formula to increase or decrease the Q - function evaluation value corresponding to the target measurement vector and the target transmission time slot: Among them, Q u () (t, m) represents the Q - function evaluation value corresponding to the updated target measurement vector and the target transmission time slot, represents the Q - function evaluation value corresponding to the target measurement vector and the target transmission time slot before the update, α represents the learning rate, and r represents the reward and punishment factor.

6. The method according to claim 4, characterized in that Use the target measurement vector to perform compressive measurement on the first data packet to be transmitted to obtain the second data packet to be transmitted, which is achieved by the following formula: Among them, is the second data packet to be sent, is the first data packet to be sent, is the target measurement vector.

7. The method according to any one of claims 1 to 6, characterized in that After determining the final measurement vector and the final transmission time slot, the method further includes: Obtain the length value of the updated data frame broadcast by the satellite side, and return to execute the step of dividing multiple transmission time slots according to the length value of the data frame and generating a two - dimensional Q - value table until the load of the entire system is stable.

8. The method according to claim 7, wherein The length value of the data frame is updated in the following way: Take the total number of elements in all time - slot support sets in the last data frame within the Q - learning window as the estimated number of active nodes, and determine whether the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes; If the estimated number of active nodes is greater than the upper limit value of the number of accessible nodes, increase the data frame length; If the estimated number of active nodes is less than the lower limit value of the number of accessible nodes, decrease the data frame length; If the estimated number of active nodes is between the upper limit value and the lower limit value of the number of accessible nodes, the load is stable.

9. The method according to claim 8, wherein The upper limit value of the number of accessible nodes is calculated by the following formula: S up = G * ·N·T; The lower limit value of the number of accessible nodes is calculated by the following formula: S down = (G * - Δ s )·(N - 1)·T; Among them, G * is the load threshold, N is the number of rows of the measurement matrix, T is the length value of the data frame, and Δ s is the adjustment buffer interval.

10. The method according to claim 8, characterized in that The length value of the data frame is increased by the following formula: T = 2T; The length value of the data frame is decreased by the following formula: where round() is the floor function, T is the length value of the data frame, S is the estimated number of active nodes, and N is the number of rows of the measurement matrix.

Citation Information

Patent Citations

  • Intelligent random access method in Satellite Internet of Things

    CN108924946A

  • satellite Internet of Things asynchronous random access method based on a Q learning algorithm

    CN109905165A