Data packet transmission method and electronic equipment
By using the trained main network in the node device, intelligent decision-making based on channel characteristics determines the backoff time, solving the problem that traditional backoff strategies cannot meet diversified needs and insufficient channel resource utilization, and achieving more efficient packet transmission.
Patent Information
- Application Number
- CN202510260882.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional backoff strategies cannot meet diverse transmission needs and cannot make full use of channel resources to improve overall transmission efficiency.
By obtaining the current channel characteristics of the node device, inputting them to the trained main network, obtaining the target action, and determining the backoff time to send the data packet based on the target action. The main network optimizes its model parameters to achieve intelligent decision-making through the empirical samples accumulated during training and the target network.
Compared with traditional fixed strategies, this method avoids resource waste or collision conflict problems, realizes efficient utilization of channel resources, and significantly improves the accuracy, stability and efficiency of transmission.
Smart Images

Figure CN120224480A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communication technologies, and particularly relates to a data packet transmission method and an electronic device. Background Art
[0002] In a wireless network environment, communication between node devices faces many challenges. On the one hand, channel resources are limited, and when multiple node devices use the channel simultaneously, conflicts are likely to occur. On the other hand, the channel is easily affected by various factors, resulting in continuous changes in state characteristics such as the bandwidth, signal-to-noise ratio, and delay of the channel, thereby affecting the transmission quality and efficiency of data packets.
[0003] Currently, fixed strategies are usually adopted to handle the channel and data packet transmission. For example, in the early Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) protocol, node devices usually adopt a fixed backoff algorithm to avoid conflicts.
[0004] However, the fixed backoff algorithm cannot be adjusted according to the actual state of the channel, which may lead to too long a backoff time when the channel is busy, wasting channel resources; or, too short a backoff time when the channel is idle, easily causing conflicts. That is, the traditional backoff strategy cannot meet diverse transmission requirements and cannot fully utilize channel resources to improve the overall transmission efficiency. Summary of the Invention
[0005] Embodiments of this application provide a data packet transmission method and an electronic device, which can solve the problem that the traditional backoff strategy cannot meet diverse transmission requirements and cannot fully utilize channel resources to improve the overall transmission efficiency.
[0006] In a first aspect, embodiments of this application provide a data packet transmission method, which includes:
[0007] For any node device, obtain the current first channel feature of the node device; the first channel feature is used to describe the network state of the node device and the information of the data packet to be transmitted;
[0008] Input the first channel feature into the trained main network to obtain the target action for the node device to transmit the data packet;
[0009] Send the data packet based on the backoff time corresponding to the target action.
[0010] In an embodiment, the main network is trained through the following steps, including:
[0011] Obtain the current second channel feature of the node device;
[0012] Based on the second channel feature and the main network to be trained, determine the first action for the node device to transmit the data packet;
[0013] Obtain the execution result after the node device executes the first action, as well as the third channel feature;
[0014] Based on the execution result, the second channel feature, and the third channel feature, determine the first reward; the first reward is used to characterize the effect of the node device executing the first action;
[0015] At least store the second channel feature and the first reward as a training sample in the experience pool;
[0016] Based on multiple training samples in the experience pool and the target network, train the model parameters of the main network to be trained to obtain the main network; the main network to be trained has the same structure as the target network.
[0017] In one embodiment, determining the first action of the node device to transmit a data packet based on the second channel feature and the main network to be trained includes:
[0018] Input the second channel feature into the main network to be trained to obtain the second action of the node device to transmit a data packet;
[0019] Determine a random action from the preset action space; the action space includes multiple actions for the node device to execute;
[0020] Determine one of the second action and the random action as the first action.
[0021] In one embodiment, determining one of the second action and the random action as the first action includes:
[0022] Generate a random value;
[0023] If the random value is less than or equal to the exploration rate corresponding to the current training process, determine the random action as the first action;
[0024] If the random value is greater than the exploration rate, determine the second action as the first action.
[0025] In one embodiment, obtaining the execution result after the node device executes the first action includes:
[0026] Based on the first action, determine the target backoff time for the node device to send a data packet;
[0027] Send the data packet after the target backoff time to obtain the execution result.
[0028] In one embodiment, determining the first reward based on the execution result, the second channel feature, and the third channel feature includes:
[0029] Based on the second channel feature and the third channel feature, determine the channel change;
[0030] Determine the second reward corresponding to the execution result, the third reward corresponding to the channel change, and the fourth reward corresponding to the first action respectively;
[0031] Determine the first reward based on the second reward, the third reward, and the fourth reward.
[0032] In one embodiment, training the model parameters of the main network to be trained based on multiple training samples in the experience pool and the target network to obtain the main network, including:
[0033] For any training sample, input the training sample into the main network to be trained to obtain the predicted value of the training sample; the predicted value is used to reflect the predicted reward obtained after the node device executes the action predicted by the main network to be trained under the second channel feature;
[0034] Input the training sample into the target network to obtain the target value of the training sample; the target value is used to reflect the expected reward obtained after the node device executes the action predicted by the target network under the second channel feature;
[0035] Update the model parameters based on the predicted values and target values corresponding to the multiple training samples respectively to obtain the main network model.
[0036] In one embodiment, inputting the training sample into the target network to obtain the target value of the training sample, including:
[0037] Input the second channel feature in the training sample into the target network to obtain the future reward of the training sample;
[0038] Based on the future reward and the first reward, obtain the target value.
[0039] In one embodiment, inputting the training sample into the target network to obtain the target value of the training sample, including:
[0040] Calculate the difference between the preset maximum reward and the first reward;
[0041] Determine the actual future reward as the product of the difference, the preset discount factor, and the future reward;
[0042] Determine the sum of the actual future reward and the first reward as the target value.
[0043] In a second aspect, an embodiment of the present application provides a data packet transmission device, and the device includes:
[0044] A first acquisition module, configured to acquire the current first channel feature of any node device; the first channel feature is used to describe the network state of the node device and the information of the data packet to be transmitted;
[0045] The first determination module is configured to input the first channel feature into the trained main network to obtain the target action of the node device for transmitting the data packet.
[0046] The sending module is configured to send the data packet based on the backoff time corresponding to the target action.
[0047] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to the first aspect above is implemented.
[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to the first aspect above is implemented.
[0049] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on an electronic device, the electronic device is caused to execute the method according to the first aspect above.
[0050] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: For any node device, the electronic device can first obtain the first channel feature that can reflect its network state and data packet information to comprehensively and accurately understand the transmission environment of the node device. Then, the electronic device can input the first channel feature into the trained main network and determine the target action of transmitting the data packet based on the intelligent decision-making ability of the trained main network, so as to determine the backoff time according to the target action. This method is more scientific and flexible than the traditional fixed strategy, and can avoid the resource waste or collision problem of the traditional backoff method, realizing the efficient utilization of channel resources. Finally, sending the data packet based on the backoff time can enable the node device in the whole process to adapt to the network dynamic changes in real time, dynamically adjust the backoff time in a complex network environment to ensure the reliable transmission of the data packet, and significantly improve the accuracy, stability and efficiency of the transmission. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flowchart of the implementation of a data packet transmission method provided by an embodiment of the present application;
[0053] Figure 2 It is a flowchart of the implementation of a main network training method provided by an embodiment of the present application;
[0054] Figure 3 It is a schematic diagram of the state space and action space in a data packet transmission method provided by an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of an implementation manner of updating model parameters in a main network training method provided by an embodiment of the present application;
[0056] Figure 5 It is a schematic diagram of an application scenario of a solution for training a main network provided by another embodiment of the present application;
[0057] Figure 6 It is a schematic diagram of the structure of a data packet transmission device provided by an embodiment of the present application;
[0058] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0059] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0060] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0061] It should be noted that the information collection process (such as the face image collection process, fingerprint information collection process, etc.) / feature extraction process involved in the present application is executed with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not belong to an act that harms the public interest.
[0062] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0063] In a wireless network environment, communication between node devices faces many challenges. On the one hand, channel resources are limited and shared. When multiple node devices compete to use the channel simultaneously, conflicts are likely to occur, resulting in packet transmission failures or low transmission efficiency. For example, in a wireless local area network (WLAN), multiple devices such as mobile phones and laptops need to transmit data through a wireless access point (AP). When they simultaneously attempt to send packets, collisions are likely to occur.
[0064] On the other hand, the channel usually changes dynamically and is easily affected by various factors such as interference in the environment, signal fading, and multipath effects. Usually, these factors cause the state characteristics of the channel, such as bandwidth, signal-to-noise ratio, and delay, to change continuously, thereby affecting the transmission quality and efficiency of packets. For example, in a mobile scenario, as the terminal device moves, the channel changes rapidly with the change of location.
[0065] Currently, fixed strategies are usually adopted to handle the channel and packet transmission. For example, in the early collision avoidance (CSMA / CA) protocol, node devices usually adopt a fixed backoff algorithm to avoid conflicts.
[0066] However, the fixed backoff algorithm cannot be adjusted according to the actual state of the channel. It may result in too long a backoff time when the channel is busy, wasting channel resources; or, too short a backoff time when the channel is idle, easily triggering conflicts. That is, the traditional backoff strategy cannot meet diverse transmission requirements and cannot fully utilize channel resources to improve the overall transmission efficiency.
[0067] In addition, the fixed backoff algorithm does not fully utilize the network state of node devices and the information of packets to be transmitted to optimize the transmission strategy. For example, for packets with different priorities and different sizes, the same backoff algorithm is adopted, which cannot meet diverse transmission requirements and cannot fully utilize channel resources to improve the overall transmission efficiency.
[0068] Based on this, in order to meet diverse transmission requirements and fully utilize channel resources to improve the overall transmission efficiency, the embodiments of the present application provide a packet transmission method. This method can be applied to electronic devices such as mobile phones, laptops, ultra-mobile personal computers (UMPCs), netbooks, switches, and routers. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.
[0069] Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a packet transmission method provided by the embodiments of the present application. The method includes the following steps:
[0070] S101. For any node device, obtain the current first channel feature of the node device; the first channel feature is used to describe the network state of the node device and the information of the data packet to be transmitted.
[0071] In one embodiment, the above-mentioned node device refers to a device that can send, receive, or forward data and is the basic unit constituting the network. For example, mobile phones, tablet computers, laptop computers, routers, base stations, etc. all belong to node devices. Multiple node devices are interconnected through network links to jointly complete data transmission and processing tasks. Among them, the electronic device that executes the above data packet transmission method can be this node device.
[0072] The above first channel feature can be a parameter set for comprehensively describing the network state of the node device and the information of the data packet to be transmitted. Exemplarily, the first channel feature includes but is not limited to features such as signal strength, signal-to-noise ratio, bit error rate, and channel occupancy rate that reflect the ability and reliability of the channel to transmit data, and may also include but is not limited to features such as data packet size, data packet priority, and data packet type that reflect the strategies and requirements for data packet transmission.
[0073] As an example, the first channel feature can be composed of the historical channel occupancy rate (for example, the channel occupancy rates of the last 8 times) and the priority queue length (for example, 8 priority queues). At this time, the vector of the first channel feature can be defined as:
[0074] state = [u1, u2,..., u8, q1, q2,..., q8];
[0075] where u i represents the channel occupancy rate of the i-th time, and q i represents the length of the i-th priority queue.
[0076] In one embodiment, the channel occupancy rate refers to the ratio of the time the channel is occupied to the total time within a certain period, usually expressed as a percentage. Among them, the channel occupancy rate can reflect the busyness of the channel and is an important indicator for measuring the usage of network resources.
[0077] For example, a higher channel occupancy rate may indicate a heavier network load, and the node device needs to wait longer to obtain the channel for data transmission, which will lead to problems such as an increase in data packet transmission delay and an increase in the packet loss rate. On the contrary, a lower channel occupancy rate indicates that the channel resources are not fully utilized, and there may be a situation of waste of network resources.
[0078] Moreover, the priority queue length refers to the number of data packets waiting to be processed or transmitted in each queue divided according to different priorities in a system with a priority management mechanism. In network communication, to ensure the priority transmission of critical services or important data, data packets are usually assigned to different priority queues according to their importance or urgency.
[0079] It should be noted that the priority queue length can help the electronic device understand the load conditions of different priority services. By monitoring the priority queue length, the electronic device can dynamically adjust the resource allocation and scheduling strategy according to the changes in the queue. For example, when the length of the high-priority queue continues to increase, it indicates that the load of high-priority services is heavy, and the electronic device may need to preferentially allocate more resources to process the data packet to ensure the quality of service of high-priority services. Moreover, for the low-priority queue, if the length is too long, the processing can be delayed to ensure the smooth progress of high-priority services.
[0080] S102: Input the first channel feature into the trained main network to obtain the target action for the node device to transmit the data packet.
[0081] In one embodiment, the above-mentioned main network is usually a model based on machine learning or deep learning, such as a neural network. Generally, in the training phase, the network can learn through a large amount of training data, adjust its own parameters, and establish a mapping relationship between the channel feature and the optimal action. Generally, the training data generally includes different channel features and the corresponding optimal action selections. Through continuous iterative optimization, the main network can predict the appropriate action according to the input channel feature. After training, the main network has the ability to output the target action under the given channel feature.
[0082] In one embodiment, the above-mentioned target action refers to the optimal action that the node device should take to achieve efficient and reliable data packet transmission under the current first channel feature. The target action can cover various decisions, such as selecting the appropriate transmission power, modulation and coding method, channel access time, backoff time for data packet transmission, whether to perform data packet fragmentation, etc. That is to say, the target action is the result obtained by the main network through analysis and judgment according to the input first channel feature.
[0083] As an example, the above-mentioned main network can be a network for estimating the Q value of the state (first channel feature)-action pair, and then, the action corresponding to the maximum value of the Q value is determined as the target action. Moreover, the target action can be the backoff time for data packet transmission, so that when the node device sends the data packet based on the backoff time corresponding to the target action, data packet collision can be avoided. That is to say, the main network can output the Q value corresponding to each action under the first channel feature, and then, the action corresponding to the maximum value among the multiple Q values is determined as the target action.
[0084] In this embodiment, since the first channel feature can accurately characterize the current network condition of the node device and the characteristics of the data packet to be transmitted, and the trained main network can mine the mapping relationship between the channel feature and the optimal action. Therefore, by inputting the first channel feature into the main network for processing, the electronic device can screen out the action from many possible actions that can maximize the data packet transmission efficiency and reduce conflicts.
[0085] S103. Transmit the data packet based on the backoff time corresponding to the target action.
[0086] In one embodiment, when the target action is the backoff time, the electronic device can directly transmit the data packet based on the backoff time corresponding to the target action.
[0087] In another embodiment, in order to enable the backoff time determined by the node device to maintain good adaptability and flexibility in different application scenarios, it can be calculated according to the following formula:
[0088] backoff_time = min_backoff + (max_backoff - min_backoff) * action / (action_space - 1);
[0089] Wherein, backoff_time represents the backoff time, min_backoff represents the minimum backoff time preset by the node device, max_backoff represents the maximum backoff time preset by the node device, action represents the action (for example, the target action), and action_space represents the preset action space dimension.
[0090] Based on the above formula, it can be understood that the minimum backoff time and the maximum backoff time limit the value range of the backoff time, and this range can be flexibly set according to the specific network scenario and requirements. Furthermore, in practical applications, the electronic device can adjust the maximum and minimum backoff times in scenarios with different channel busy degrees, and calculate the appropriate backoff time in combination with the action, so that the setting of the backoff time can better adapt to different network conditions and improve the success rate of data packet transmission and network efficiency.
[0091] In this embodiment, for any node device, the electronic device can first obtain the first channel feature that can reflect its network state and packet information, so as to comprehensively and accurately understand the transmission environment of the node device. Then, the electronic device can input the first channel feature into the trained main network, and determine the target action of transmitting the data packet based on the intelligent decision-making ability of the trained main network, so as to determine the backoff time according to the target action. This method is more scientific and flexible than the traditional fixed strategy, and can avoid the resource waste or collision problems of the traditional backoff method, realizing the efficient utilization of channel resources. Finally, sending the data packet based on the backoff time can enable the node devices in the whole process to adapt to the dynamic changes of the network in real time, dynamically adjust the backoff time in a complex network environment, so as to ensure the reliable transmission of the data packet, and significantly improve the accuracy, stability and efficiency of the transmission.
[0092] In one embodiment, the main network in S102 can be trained through the steps S201-S206 as Figure 2 shown. Details are as follows:
[0093] S201. Obtain the current second channel feature of the node device.
[0094] In one embodiment, the above second channel feature is similar to the first channel feature, and the only difference is that the second channel feature is the data in the training process, and the first channel feature is the data in the actual application scenario, which will not be elaborated here.
[0095] S202. Based on the second channel feature and the main network to be trained, determine the first action of the node device to transmit the data packet.
[0096] In one embodiment, the above main network to be trained can be an initialized network or a network that needs to be updated. Among them, the network that needs to be updated can be considered to have a certain prediction ability. At this time, due to the update of the training samples, it is necessary to train this network again to improve the prediction ability.
[0097] Moreover, for the initialized network, it can be considered as a network model that has not completed the parameter optimization and training process. Different from the trained main network, the initial parameters of the main network to be trained (such as weights and biases in the neural network) are usually randomly initialized or set based on some prior knowledge. Before the start of the training process, it usually does not have the ability to accurately predict the appropriate action.
[0098] As an example, taking the main network to be trained as DQN (Deep Q-Networks, a deep reinforcement learning algorithm) as an example, the DQN agent initialization hyperparameters include: the number of node priorities (num_priorities), the historical channel length (history_depth), the minimum backoff time (min_backoff), the maximum backoff time (max_backoff), the state space dimension (state_space), the action space dimension (action_space), the learning rate (learning_rate), the discount factor (gamma), the decay coefficient (rho), the exploration rate (epsilon), the batch size (batch_size), the experience replay buffer size (buffer_size), the target network update frequency (memory_capacity), etc.
[0099] Among them, the above-mentioned minimum backoff time (min_backoff) is the minimum waiting time for the node device after encountering a collision or other problems when sending data. The above-mentioned maximum backoff time (max_backoff) corresponds to the minimum backoff time and stipulates the longest waiting time for the node device during backoff.
[0100] The above-mentioned state space dimension (state_space) represents the number or dimension of the characteristics of the environmental state faced by the DQN network. These characteristics can be considered as the various channel characteristics that appear above (for example, the first channel characteristic, the second channel characteristic, etc.). The state space dimension determines the complexity of the information that the network needs to process. A larger dimension means that the network needs to process more information to make a decision, but it may also provide a richer description of the environment, helping the network learn a better strategy.
[0101] The above-mentioned action space dimension (action_space) represents the number or dimension of different actions that the node device can take. For example, the possible actions of the node device include but are not limited to sending data, backoff time, adjusting the transmission power, etc. The action space dimension determines the decision-making range of the network. More actions mean that the network has more choices to adapt to different environmental states, but it also increases the difficulty and complexity of learning.
[0102] Among them, the state space dimension and the dynamic space dimension can be determined by the number of node priorities and the historical channel length. For example, both the state space dimension and the action space dimension can be the sum of the number of node priorities and the historical channel length respectively.
[0103] Exemplarily, referring to Figure 3 , Figure 3It is a schematic diagram of the state space and action space in a data packet transmission method provided by an embodiment of the present application. Among them, the state space is formed by multiple historical channel occupancy rates (for example, 0.9, 0.9, 0.8,...) and multiple priority queue lengths (0, 1, 0, 1,...). The action space is formed by multiple actions (for example, backoff time).
[0104] The above learning rate controls the speed at which the neural network updates parameters during training. A larger learning rate can make the network converge faster, but may lead to unstable training; a smaller learning rate will make the training process slow, but can make the network converge more stably to a better solution.
[0105] The discount factor (gamma) is used to measure the importance of future rewards. In reinforcement learning, the goal of DQN is to maximize the long-term cumulative reward. The discount factor determines the weight of future rewards in the current decision-making, and its value range is usually between 0 and 1. When gamma is close to 1, it means that the network to be trained pays more attention to long-term rewards and will consider the consequences of multiple future steps to make the current decision; when gamma is close to 0, DQN pays more attention to immediate rewards and mainly acts based on the rewards of the current step.
[0106] The decay coefficient (rho) is usually used to control the decay rate of certain parameters over time or the number of training steps. For example, the exploration rate may gradually decay as training progresses to reduce random exploration and make more use of the learned strategies.
[0107] The exploration rate (epsilon) is used to balance exploration and exploitation in the training process of DQN. In reinforcement learning, DQN needs to trade off between trying new actions (exploration) and choosing known optimal actions (exploitation).
[0108] The batch size refers to the number of data samples taken from the experience replay buffer each time to update the network parameters during the training of the neural network. A larger batch size can use more data information to update the network and make the training more stable, but may cause excessive memory occupation and slower training speed; a smaller batch size may make the training process more unstable, but can update the network parameters faster.
[0109] The experience pool (buffer_size) is used to store the experience data accumulated by DQN during the interaction with the environment. These experience data include, but are not limited to, the current state (for example, the current channel characteristics), actions, rewards, and the next (for example, the channel characteristics at the next moment) state and other information. DQN can randomly sample data from this experience pool for learning during training.
[0110] The target network update frequency (memory_capacity) refers to the target network used to stabilize the learning process during the training of the DQN network. The target network update frequency indicates how many steps or time intervals between copying the parameters of the main network to the target network.
[0111] In one embodiment, the above first action is similar to the above target action, with the only difference being that the first action is data generated during the training process, and the target action is data generated in the actual application scenario, which will not be elaborated here.
[0112] It should be noted that in order to balance exploration and exploitation when the electronic device selects an action during the training process, the electronic device can also input the second channel feature into the main network to be trained to obtain the second action of the node device for transmitting data packets. Then, a random action is determined from the preset action space; the action space includes multiple actions for the node device to execute. Finally, one of the second action and the random action is determined as the first action.
[0113] In one embodiment, the above second action is similar to the target action, which will not be elaborated here. Among them, the electronic device can randomly determine the first action from the second action and the random action. Alternatively, the electronic device can generate a random value. Then, when the random value is less than or equal to the exploration rate corresponding to the current training process, the random action is determined as the first action. Otherwise, when the random value is greater than the exploration rate, the second action is determined as the first action.
[0114] In one embodiment, the above exploration rate can be a fixed value or a value that decays linearly or exponentially in each round of training. In this embodiment, the setting of the exploration rate is not limited.
[0115] It should be noted that when determining the first action in the above manner, the balance between exploitation and exploration can be achieved. First, inputting the second channel feature into the main network to be trained to obtain the second action enables the main network to learn and output actions that conform to the actual situation based on the current channel state information, reflecting the exploitation of environmental information and guiding the main network to be trained to learn in the direction of a better transmission strategy. At the same time, determining a random action from the preset action space can introduce an exploration mechanism for the action selection of the main network to be trained, avoiding falling into a local optimal solution due to over-reliance on the currently imperfect output of the main network to be trained, helping to discover new and potentially better action strategies, and enriching the diversity of action selection of the main network to be trained. Furthermore, the method of determining the first action from the second action and the random action balances the exploitation of existing knowledge (the second action output by the main network to be trained) and the exploration of the unknown (the random action), enabling the main network to be trained to make decisions by utilizing the current understanding of the channel state to a certain extent and continuously explore new action possibilities, enhancing the ability of the main network to be trained to adapt to complex and changing network environments and promoting the more comprehensive learning and optimization of the main network to be trained.
[0116] S203. Obtain the execution result after the node device executes the first action, and the third channel feature.
[0117] In one embodiment, the above execution result can be divided into a successful result where the data packet is successfully sent and a failure result where the data packet is sent unsuccessfully. And the above third channel feature is the channel feature after executing the first action. Generally, when the data packet is successfully sent or fails, the channel state usually changes.
[0118] Based on this, obtaining the third channel feature helps to update the node device's understanding of the network state, can understand the current latest channel situation, so as to provide accurate state information for subsequent possible operations such as data packet sending. And by combining the execution result and the new third channel feature, the actual effect of the previously executed first action can be more comprehensively evaluated, such as whether it truly optimizes the transmission process and whether it has the expected impact on channel utilization, etc., thereby providing more effective feedback for subsequent action selection.
[0119] As an example, the electronic device can determine the target backoff time for the node device to send a data packet based on the first action. Then, after the target backoff time, execute the first action to obtain the execution result.
[0120] Among them, the method of determining the target backoff time based on the first action can refer to the example description in step S104 above, and will not be explained here.
[0121] S204. Determine the first reward based on the execution result, the second channel feature, and the third channel feature.
[0122] In one embodiment, the above first reward is used to characterize the effect of the node device executing the first action. The first reward can be a numerical index for quantitatively characterizing the effect of the node device executing the first action, which is comprehensively obtained based on the execution result, the second channel feature, and the third channel feature, and can reflect the advantages and disadvantages of the action in achieving communication objectives, optimizing channel utilization, etc., providing a feedback basis for policy optimization in reinforcement learning.
[0123] In one embodiment, when determining the first reward, when the execution result is a successful result, a basic positive reward can be given, indicating that the action is successful in completing the basic task. If the execution result is a failure result, a negative reward can be given.
[0124] In addition, the electronic device can also compare the second channel feature and the third channel feature. If the third channel feature is optimized compared to the second channel feature (i.e., the channel change is positive), such as an increase in signal-to-noise ratio, a reduction in interference, etc., it can be considered that the first action has a positive impact on the channel, and the reward value can be increased. Conversely, if the channel feature deteriorates (i.e., the channel change is negative), the reward value is correspondingly reduced. Finally, the respective reward values can be weighted to obtain the final first reward.
[0125] As an example, the electronic device can determine the channel change based on the second channel feature and the third channel feature. Then, respectively determine the second reward corresponding to the execution result, the third reward corresponding to the channel change, and the fourth reward corresponding to the target backoff time. Finally, determine the first reward based on the second reward, the third reward, and the fourth reward.
[0126] In one embodiment, the electronic device can preset the reward values corresponding to each execution result, each channel change, and each backoff time respectively to determine the above second reward, third reward, and fourth reward.
[0127] Exemplarily, when the execution result is a successful result, the second reward can be 10, and when the execution result is a failure result, the second reward can be -1. In addition, when the backoff time corresponding to the first action (which can be the target backoff time or the backoff time corresponding to the first action itself) is the minimum backoff time (for example, 0), the third reward can be 15; when the backoff time corresponding to the first action is less than half of the maximum backoff time, the third reward can be 5; when the backoff time corresponding to the first action is greater than or equal to half of the maximum backoff time, the third reward can be -1. In addition, when the change in the priority queue length in the channel change is a decrease, the fourth reward is a positive reward, for example, 1; when the change in the priority queue length is an increase, the fourth reward is a negative reward, for example, -1.
[0128] In one embodiment, the electronic device may weight the second reward, the third reward, and the fourth reward to obtain the first reward above. The weight values corresponding to each reward can be set in advance, and can be the same or different, which is not limited herein.
[0129] For example, the weight value corresponding to the second reward is 60%, the weight value corresponding to the third reward is 20%, and the weight value corresponding to the fourth reward is 20%, so as to comprehensively and comprehensively evaluate the effect of the first action.
[0130] In another embodiment, the amplitude of reward adjustment can also be determined according to the degree of change in channel characteristics. The more obvious the change, the greater the amplitude of reward adjustment. Or, fine-tuning of the first-line reward is performed according to some additional indicators such as transmission rate and delay when transmission is successful. For example, a high transmission rate or low delay can increase the reward value, and vice versa. In this embodiment, the method for determining the first reward is not limited.
[0131] S205. Store at least the second channel characteristic and the first reward as a training sample in the experience pool.
[0132] In one embodiment, the electronic device may determine the second channel characteristic and the first reward as a training sample, or may also determine information such as the second channel characteristic, the first reward, the third channel characteristic, and the first action as a training sample, which is not limited herein. And, the experience pool has been explained above, and will not be described again here.
[0133] S206. Train the model parameters of the main network to be trained based on multiple training samples in the experience pool and the target network to obtain the main network.
[0134] In one embodiment, the electronic device may train the main network to be trained at preset time intervals, or when the number of training samples in the experience pool reaches a preset batch size, which is not limited herein.
[0135] In this embodiment, taking the training when the number of training samples in the experience pool reaches the batch size as an example, when the number of training samples in the experience pool reaches the predetermined batch size, the electronic device may randomly select multiple training samples from the experience pool for training. Furthermore, the correlation between training samples can be broken, and the training efficiency and stability can be improved.
[0136] In one embodiment, the main network to be trained has the same structure as the target network. And, the model parameters of the target network are usually copied from the main network during initialization, so that the target network and the main network have the same structure and parameter values initially.
[0137] Among them, the model parameters of the target network are usually not updated frequently with those of the main network. Instead, at regular time steps (e.g., 100 steps) or after a certain number of training iterations, the parameters of the main network are copied to the target network for update. That is to say, the role of the target model can be considered as providing a stable target value for the currently to-be-trained main network model.
[0138] It should be noted that adopting the above update method can enable the target network to gradually keep up with the learning progress of the main network while maintaining a certain degree of stability, and obtain the new information learned by the main network.
[0139] Specifically, when the to-be-trained main network is being trained, it needs to continuously update its parameters to optimize the strategy. At this time, if the output of the to-be-trained main network itself is directly used as the target for update, it is easy for the target to change frequently due to the change of parameters, resulting in an unstable learning process and easy oscillation. However, the target network is updated relatively slowly, which can provide a relatively stable reference for calculating the target value. When the to-be-trained main network calculates the loss function to update the parameters, it can utilize the relatively stable target value output by the target network to effectively reduce the estimation bias, enabling the to-be-trained main network to focus more on learning the long-term trends and patterns in the data under the guidance of the relatively stable target value.
[0140] As an example, the electronic device can obtain the main network model according to the steps S401 - S403 as shown in Figure 4 which are detailed as follows:
[0141] S401: For any training sample, input the training sample into the to-be-trained main network to obtain the predicted value of the training sample.
[0142] In one embodiment, the predicted value is used to reflect the predicted reward obtained after the node device executes the action predicted by the to-be-trained main network under the second channel feature.
[0143] In one embodiment, the above predicted value can be considered as the Q value estimated by the to-be-trained main network for the action corresponding to the second channel feature. Specifically, the explanation in the above S102 can be referred to. That is, the main network can output the Q values corresponding to each action executed under the first channel feature, and then determine the action corresponding to the maximum value among the multiple Q values as the target action. Based on this, the electronic device can determine the Q value of the action corresponding to the second channel feature as the above predicted value.
[0144] S402: Input the training sample into the target network to obtain the target value of the training sample.
[0145] In one embodiment, the above target value is used to reflect the expected reward obtained after the node device executes the action predicted by the target network under the second channel characteristic. The target value may be similar to the corresponding Q value mentioned above, and no detailed description will be given here.
[0146] However, it should be noted that the reward directly output by the target network based on the second channel characteristic is usually a future reward, and this reward has uncertainty. Furthermore, it is likely to cause the main network to be trained to be difficult to evaluate the true value of the action.
[0147] Based on this, in order to accurately evaluate the true value of the action, the electronic device may input the second channel characteristic in the training sample into the target network to obtain the future reward of the training sample. Then, based on the future reward and the first reward, the target value is obtained.
[0148] Among them, the first reward, as the immediate reward corresponding to the second channel characteristic, can provide a short-term evaluation of the action. It can be understood that due to the uncertainty and delay of the future reward, it is difficult to judge the current behavior only by it, while the immediate reward can enable the main network to be trained to adjust quickly. In addition, the immediate reward plays a key role in balancing exploration and exploitation. It can encourage the main network to be trained to explore new actions and also prompt it to use existing experience to obtain more rewards.
[0149] As an example, the electronic device may add the future reward and the first reward to obtain the target value.
[0150] In another embodiment, in order to improve the accuracy of the calculated target value, the electronic device may calculate the difference between the preset maximum reward and the first reward. Then, the product of the difference, the preset discount factor, and the future reward is determined as the actual future reward. Finally, the sum of the actual future reward and the first reward is determined as the target value.
[0151] Among them, the above preset discount factor may be the discount factor described above, and the above preset maximum reward may be set according to the actual situation, and no limitation will be made here. Exemplarily, the above preset maximum reward may be 1.
[0152] It should be noted that calculating the target value in the above manner can balance short-term rewards and long-term rewards. Adding the first reward and the actual future reward obtained through complex calculations can prompt the main network to be trained to take into account both immediate benefits and long-term development. And by considering the difference between the preset maximum reward and the first reward to dynamically calculate the future reward, the reward mechanism can flexibly reflect the current performance of the main network to be trained and environmental changes, enhancing adaptability. At the same time, the introduction of the discount factor can reflect the importance of the future reward, enabling the finally calculated target value to fit the uncertainty and time value characteristics of the future reward and improving the learning efficiency.
[0153] S403. Update the model parameters based on the predicted values and target values respectively corresponding to multiple training samples to obtain the main network model.
[0154] In one embodiment, the electronic device can calculate the loss value between the predicted value and the target value to update the model parameters based on the loss value. Among them, the loss function for calculating the loss value includes but is not limited to the mean squared error loss function, the mean absolute error loss function, the cross-entropy loss function, and there is no limitation on this. And, the ways to update the model parameters include but are not limited to gradient descent update, adaptive learning rate update, etc., and there is no limitation on this.
[0155] Exemplarily, the electronic device can use the mean squared error loss function to calculate the loss values of multiple predicted values and target values. Then, use the gradient descent update method to update the model parameters based on the loss value to obtain the main network model.
[0156] In another embodiment, during each training process, the electronic device can also record the loss value for subsequent analysis and drawing of the loss curve. Through the above steps, the main network can gradually optimize its decision-making strategy and improve the efficiency and accuracy of the backoff time selection in wireless communication.
[0157] To more clearly illustrate the solution for training the main network in this application, the following uses specific embodiments to elaborate on the solution in this application. For details, refer to Figure 5 , Figure 5 FIG. is a schematic diagram of an application scenario of a solution for training the main network provided in another embodiment of this application. For ease of explanation, taking a system composed of multiple node devices as an example, each node device is deployed with an initialized main network to be trained (DQN) and a target network. Among them, the initialized hyperparameters can be as described in S202 above.
[0158] For any node device, the node device can sense the historical channel occupancy rate and the priority queue length to generate a second channel feature, and input the second channel feature into the main network to be trained. Then, the main network to be trained can output the Q values corresponding to performing each action, and determine the action corresponding to the maximum value among the multiple Q values as the first action.
[0159] After that, the electronic device can send a data packet based on the backoff time corresponding to the first action to obtain an execution result, and the third channel feature at this time. And, based on the channel change between the execution result, the second channel feature and the third channel feature, and based on the target backoff time calculated by the first action, determine the first reward, and determine information such as the second channel feature, the first reward, the third channel feature, and the first action as a training sample, and store it in the experience pool.
[0160] Subsequently, the electronic device can input multiple training samples in the experience pool into the target network and the main network to be trained respectively, obtain the corresponding predicted values and target values, and calculate the training loss to train the main network to be trained to obtain the main network.
[0161] Finally, each node device can output the corresponding target action based on the trained main network and the first channel feature deployed by itself, and send the data packet based on the backoff time corresponding to the target action.
[0162] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a data packet transmission device provided by an embodiment of the present application. In this embodiment, each module included in the data packet transmission device is used to execute Figures 1 to 5 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figures 1 to 5 and Figures 1 to 5 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 6 , the data packet transmission device 600 may include: a first acquisition module 610, a first determination module 620, and a sending module 630, where:
[0163] The first acquisition module 610 is configured to acquire, for any node device, the current first channel feature of the node device; the first channel feature is used to describe the network state of the node device and the information of the data packet to be transmitted.
[0164] The first determination module 620 is configured to input the first channel feature into the trained main network to obtain the target action of the node device for transmitting the data packet.
[0165] The sending module 630 is configured to send the data packet based on the backoff time corresponding to the target action.
[0166] In one embodiment, the data packet transmission device 600 further includes the following modules to train the main network:
[0167] The second acquisition module is configured to acquire the current second channel feature of the node device.
[0168] The second determination module is configured to determine the first action of the node device for transmitting the data packet based on the second channel feature and the main network to be trained.
[0169] The third acquisition module is configured to acquire the execution result after the node device executes the first action, and the third channel feature.
[0170] The third determination module is configured to determine the first reward based on the execution result, the second channel feature, and the third channel feature; the first reward is used to characterize the effect of the node device executing the first action.
[0171] A storage module for storing at least the second channel feature and the first reward as a training sample in an experience pool.
[0172] A training module for training the model parameters of the main network to be trained based on multiple training samples in the experience pool and a target network to obtain the main network; the main network to be trained has the same structure as the target network.
[0173] In one embodiment, the second determination module is further configured to:
[0174] Input the second channel feature into the main network to be trained to obtain a second action for the node device to transmit a data packet; determine a random action from a preset action space; the action space includes multiple actions for the node device to execute; determine one of the second action and the random action as the first action.
[0175] In one embodiment, the second determination module is further configured to:
[0176] Generate a random value; if the random value is less than or equal to the exploration rate corresponding to the current training process, determine the random action as the first action; if the random value is greater than the exploration rate, determine the second action as the first action.
[0177] In one embodiment, the third acquisition module is further configured to:
[0178] Determine a target backoff time for the node device to send a data packet based on the first action; send the data packet after the target backoff time to obtain an execution result.
[0179] In one embodiment, the third determination module is further configured to:
[0180] Determine a channel change based on the second channel feature and the third channel feature; respectively determine a second reward corresponding to the execution result, a third reward corresponding to the channel change, and a fourth reward corresponding to the first action; determine the first reward based on the second reward, the third reward, and the fourth reward.
[0181] In one embodiment, the training module is further configured to:
[0182] For any training sample, input the training sample into the main network to be trained to obtain a predicted value of the training sample; the predicted value is used to reflect the predicted reward obtained after the node device executes the action predicted by the main network to be trained under the second channel feature; input the training sample into the target network to obtain a target value of the training sample; the target value is used to reflect the expected reward obtained after the node device executes the action predicted by the target network under the second channel feature; update the model parameters based on the predicted values and target values respectively corresponding to multiple training samples to obtain the main network model.
[0183] In one embodiment, the training module is further configured to:
[0184] Input the second channel feature in the training sample into the target network to obtain the future reward of the training sample; based on the future reward and the first reward, obtain the target value.
[0185] In one embodiment, the training module is further configured to:
[0186] Calculate the difference between the preset maximum reward and the first reward; determine the product of the difference, the preset discount factor, and the future reward as the actual future reward; determine the sum of the actual future reward and the first reward as the target value.
[0187] It should be understood that Figure 6 In the structural schematic diagram of the data packet transmission device shown, each module is used to execute Figures 1 to 5 the respective steps in the corresponding embodiment, and for Figures 1 to 5 the respective steps in the corresponding embodiment have been explained in detail in the above embodiments. For details, please refer to Figures 1 to 5 and Figures 1 to 5 the relevant descriptions in the corresponding embodiment, which will not be elaborated here.
[0188] Figure 7 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device 700 in this embodiment includes: a processor 710, a memory 720, and a computer program 730 stored in the memory 720 and executable on the processor 710, such as a program for the data packet transmission method. When the processor 710 executes the computer program 730, it implements the steps in each of the above embodiments of the data packet transmission method, such as Figure 1 S101 to S103 shown. Alternatively, when the processor 710 executes the computer program 730, it implements the functions of each module in the above Figure 6 corresponding embodiment, for example, Figure 6 the functions of each module shown. For details, please refer to Figure 6 the relevant descriptions in the corresponding embodiment.
[0189] Exemplarily, the computer program 730 can be divided into one or more modules. One or more modules are stored in the memory 720 and executed by the processor 710 to implement the data packet transmission method provided by the embodiments of the present application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 730 in the electronic device 700. For example, the computer program 730 can implement the data packet transmission method provided by the embodiments of the present application.
[0190] The electronic device 700 may include, but is not limited to, a processor 710 and a memory 720. Those skilled in the art can understand,Figure 7 This is only an example of the electronic device 700, which does not constitute a limitation on the electronic device 700. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0191] The so-called processor 710 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0192] The memory 720 may be an internal storage unit of the electronic device 700, such as the hard disk or memory of the electronic device 700. The memory 720 may also be an external storage device of the electronic device 700, such as a plug-in hard disk, smart memory card, flash memory card, etc. equipped on the electronic device 700. Further, the memory 720 may also include both the internal storage unit and the external storage device of the electronic device 700.
[0193] The embodiments of the present application provide a computer-readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the data packet transmission method in the above-mentioned various embodiments is implemented.
[0194] The embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute the data packet transmission method in the above-mentioned various embodiments.
[0195] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A data packet transmission method, characterized in that: The method comprises: For any node device, obtaining a current first channel characteristic of the node device; the first channel characteristic is used to describe the network state of the node device and information of a data packet to be transmitted; Inputting the first channel feature into the trained main network to obtain a target action of the node device transmitting a data packet; The data packet is sent based on a backoff time corresponding to the target action.
2. The method according to claim 1, characterized in that The main network is trained by the following steps, including: Acquire the current second channel characteristics of the node device; Based on the second channel characteristics and the main network to be trained, determining a first action of the node device to transmit a data packet; Acquire an execution result after the node device executes the first action, and a third channel feature; Determine a first reward based on the execution result, the second channel characteristic, and the third channel characteristic; the first reward is used to characterize the effect of the node device executing the first action; At least taking the second channel feature and the first reward as a training sample, and storing them in an experience pool; Based on the plurality of training samples in the experience pool and the target network, the model parameters of the main network to be trained are trained to obtain the main network; the main network to be trained has the same structure as the target network.
3. The method according to claim 2, characterized in that The step of determining a first action for transmitting a data packet by the node device based on the second channel characteristic and the main network to be trained includes: Inputting the second channel feature into the main network to be trained to obtain a second action of the node device transmitting a data packet; Determine a random action from a preset action space; the action space includes a plurality of actions for the node device to execute; One of the second action and the random action is determined as the first action.
4. The method according to claim 3, characterized in that The step of determining one of the second action and the random action as the first action includes: Generate random values; If the random value is less than or equal to the exploration rate corresponding to the current training process, the random action is determined as the first action; If the random value is greater than the exploration rate, the second action is determined as the first action.
5. The method according to claim 2, characterized in that: The obtaining the execution result of the node device after executing the first action includes: Determine a target backoff time for the node device to send a data packet based on the first action; The data packet is sent after the target backoff time to obtain the execution result.
6. The method according to claim 5, characterized in that The determining a first reward based on the execution result, the second channel characteristic, and the third channel characteristic comprises: determining a channel change based on the second channel characteristic and the third channel characteristic; respectively determining a second reward corresponding to the execution result, a third reward corresponding to the channel change, and a fourth reward corresponding to the first action; The first reward is determined based on the second reward, the third reward, and the fourth reward.
7. The method according to claim 2, characterized in that: The step of training the model parameters of the main network to be trained based on the plurality of training samples and the target network in the experience pool to obtain the main network comprises: For any of the training samples, the training sample is input into the main network to be trained to obtain a predicted value of the training sample; the predicted value is used to reflect the predicted reward obtained by the node device after executing the action predicted by the main network to be trained under the second channel characteristics; Inputting the training sample into the target network to obtain a target value of the training sample; the target value is used to reflect the expected reward obtained by the node device after executing the action predicted by the target network under the second channel characteristics; The model parameters are updated based on the predicted values and the target values respectively corresponding to the plurality of training samples to obtain the main network model.
8. The method according to claim 7, characterized in that The step of inputting the training sample into the target network to obtain a target value of the training sample includes: Inputting the second channel feature in the training sample into the target network to obtain a future reward for the training sample; The target value is obtained based on the future reward and the first reward.
9. The method according to claim 8, characterized in that The step of inputting the training sample into the target network to obtain a target value of the training sample includes: Calculate the difference between the preset maximum reward and the first reward; Determine the actual future reward by multiplying the difference, the preset discount factor and the future reward; The sum of the actual future reward and the first reward is determined as the target value.
10. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method as claimed in any one of claims 1 to 9.