Airport environment communication concurrency control method and device and electronic equipment
By using a policy neural network with reinforcement learning model in the airport environment communication concurrency control system, the communication strategy of airport equipment is dynamically adjusted, and the problem of low resource utilization efficiency of existing systems is solved, and more efficient resource utilization and communication strategy adjustment is achieved.
Patent Information
- Application Number
- CN202510200728.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing airport environmental communication concurrency control system cannot make full use of available resources, resulting in low resource utilization efficiency, especially during peak periods, which is prone to resource waste or bottlenecks.
The strategy neural network in the reinforcement learning model is adopted to select and perform the target actions of the airport device based on sample data, calculate reward values through device data, optimize the target neural network, and dynamically adjust the communication strategy to adapt to environmental changes.
It improves the resource utilization efficiency of concurrent control of airport environment communications, can quickly adapt to changes in the communication environment, adjust communication strategies in a timely manner, and avoid resource waste and bottlenecks.
Smart Images

Figure CN120075261A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of airport communication, reinforcement learning, machine learning, and airport Internet of Things. Specifically, it relates to a method, device, and electronic device for concurrent control of airport environment communication. Background Art
[0002] With the rapid development and technological progress of the current aviation industry, modern airports are facing an increasingly complex communication environment. Airports usually need to conduct communication management and resource scheduling for numerous key devices (such as radar systems, navigation equipment, ground support equipment, etc.). The traditional concurrent control system for airport environment communication mainly relies on preset rules and manual intervention. Although this method can meet the requirements to a certain extent, the existing systems often cannot make full use of available resources, such as bandwidth, computing power, etc. Especially during peak hours, the limited resources cannot be allocated most effectively, resulting in resource waste or bottleneck phenomena. Therefore, the resource utilization efficiency of the current concurrent control of airport environment communication is relatively low. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a method, device, and electronic device for concurrent control of airport environment communication, which is used to improve the problem of relatively low resource utilization efficiency of concurrent control of airport environment communication.
[0004] An embodiment of the present application provides an airport environment communication concurrency control method, including: selecting and executing a target action of airport equipment based on sample data through a policy neural network in a reinforcement learning model, and obtaining device data collected by the airport equipment in the current communication cycle, where the sample data is collected by the airport equipment in the previous communication cycle; calculating a reward value corresponding to the target action according to the sample data and the device data; optimizing a target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and the reward value corresponding to the target action, where the network architecture of the target neural network is the same as that of the policy neural network; assigning the model parameters of the target neural network to the policy neural network to obtain a policy network model; using the policy network model to control the communication policy of the airport equipment, where the communication policy is to select a target action to be executed from an action space, and the action space includes multiple actions. In the implementation process of the above solution, by using the policy neural network of the reinforcement learning model to execute an action based on the sample data collected in the previous communication cycle to obtain the device data in the current communication cycle, and calculating the reward value according to the sample data and the device data, so as to optimize the target neural network and assign and update its model parameters to the policy neural network, enabling the reinforcement learning model to make optimal decisions based on historical experience and the current environmental state, and continuously collecting the latest device data and calculating the reward value. This data-driven decision optimization and reward mechanism guidance enables the reinforcement learning model to accurately select the optimal action from the action space, enabling the system to quickly adapt to changes in the communication environment, timely adjust the communication policy to cope with different demand fluctuations, thereby dynamically adapting to changes in device status and effectively improving the resource utilization efficiency of airport environment communication concurrency control.
[0005] Optionally, in an embodiment of the present application, before selecting and executing a target action of airport equipment based on sample data through a policy neural network in the reinforcement learning model, it further includes: if the policy neural network in the reinforcement learning model has not been initialized, initializing the policy neural network and randomly selecting a target action of the airport equipment from the action space. In the implementation process of the above solution, by initializing the policy neural network when the policy neural network in the reinforcement learning model has not been initialized and randomly selecting a target action of the airport equipment from the action space, this random selection of actions enables the system to quickly start in the initial stage, thereby providing diverse data samples for the model in the initial stage. This diverse data input helps the model better understand the environment and make better decisions when facing different communication environment changes, avoiding the model falling into a local optimal solution prematurely, thereby enhancing the adaptability of the model.
[0006] Optionally, in the embodiments of the present application, the policy neural network in the reinforcement learning model is used to select and execute the target action of the airport equipment based on the sample data, including: if the policy neural network in the reinforcement learning model has been initialized, the sample data is input into the policy neural network of the reinforcement learning model, so that the policy neural network selects the target action of the airport equipment from the action space; the target action is sent to the airport equipment so that the airport equipment executes the target action. In the implementation process of the above solution, the policy neural network in the reinforcement learning model is used to select and execute the target action of the airport equipment based on the sample data. This dynamic adjustment strategy of the target action makes the action selection of the airport equipment more flexible and adaptive. Compared with the traditional static rules or fixed policies, the reinforcement learning model can automatically optimize the action selection according to the environmental changes and equipment status, reducing the dependence on manual decision-making. This not only reduces the labor cost but also avoids the errors and inconsistencies that may be brought by human decision-making, improving the stability and reliability of the operation of the airport equipment.
[0007] Optionally, in the embodiments of the present application, calculating the reward value corresponding to the target action according to the sample data and the equipment data includes: if it is determined according to the sample data and the equipment data that the target action is completed on time, and the task execution efficiency of the target action is greater than the preset efficiency threshold, and the resource utilization rate of the airport equipment is greater than the preset utilization rate threshold, then a preset score is added to the original reward value to obtain the reward value corresponding to the target action. In the implementation process of the above solution, determining the increase or decrease of the reward value through conditions such as the task completion situation, task execution efficiency, and the resource utilization rate of the airport equipment can more accurately evaluate the effect of task execution. This multi-dimensional evaluation mechanism not only rewards the task completed on time but also optimizes the resource utilization efficiency, thereby improving the overall system operation efficiency.
[0008] Optionally, in the embodiments of the present application, calculating the reward value corresponding to the target action according to the sample data and the equipment data includes: if it is determined according to the sample data and the equipment data that the target action is completed late, or the task execution efficiency of the target action is less than the preset efficiency threshold, or the resource utilization rate of the airport equipment is less than the preset utilization rate threshold, then a preset score is subtracted from the original reward value to obtain the reward value corresponding to the target action. In the implementation process of the above solution, by introducing the evaluation of the target action delay, task execution efficiency, and resource utilization rate, the reward value is no longer static but dynamically adjusted according to the actual execution situation. This dynamic adjustment mechanism can more accurately reflect the actual effect of the target action, thereby being able to more precisely evaluate the effect of task execution. This multi-dimensional evaluation mechanism not only rewards the task completed on time but also optimizes the resource utilization efficiency, thereby improving the overall system operation efficiency.
[0009] Optionally, in the embodiments of the present application, optimizing the target neural network in the reinforcement learning model based on sample data, device data, target actions, and the reward values corresponding to the target actions includes: generating a scheduling event according to the sample data, device data, target actions, and the reward values corresponding to the target actions; storing the scheduling event in the replay memory of the reinforcement learning model; and optimizing the target neural network in the reinforcement learning model through the scheduling events in the replay memory. In the implementation process of the above solution, traditional reinforcement learning models usually rely on real-time data for training. By using the replay memory, historical scheduling events can be stored and reused in subsequent training, which not only reduces the cost of data collection but also improves the utilization efficiency of data. Further, the introduction of the replay memory enables the model to support both offline learning and online learning. In the offline stage, historical data can be used for pre-training; in the online stage, the memory can be updated in real time and the model can be optimized. This combination can significantly improve the adaptability and real-time decision-making ability of the model.
[0010] Optionally, in the embodiments of the present application, optimizing the target neural network in the reinforcement learning model based on sample data, device data, target actions, and the reward values corresponding to the target actions includes: inputting the sample data and device data into a machine learning model to enable the machine learning model to predict the multi-source sensor data of airport equipment in the next communication cycle and obtain prediction data; generating a scheduling event according to the prediction data, sample data, device data, target actions, and the reward values corresponding to the target actions; storing the scheduling event in the replay memory of the reinforcement learning model; and optimizing the target neural network in the reinforcement learning model through the scheduling events in the replay memory. In the implementation process of the above solution, by inputting the sample data and device data into a machine learning model to predict the multi-source sensor data of airport equipment in the next communication cycle, the operating state of the equipment and environmental changes can be predicted in advance. This prediction ability enables the reinforcement learning model to more accurately evaluate the effects of target actions, thereby optimizing the decision-making process. Further, optimizing the target neural network through the scheduling events in the replay memory can effectively reduce the overfitting phenomenon in the model training process. This optimization method enables the model to maintain good performance when facing new and unseen data, improving the generalization ability of the model.
[0011] Optionally, in the embodiments of the present application, the multiple actions include at least one of the following: disabling communication actions, enabling communication actions, adjusting the execution priority of communication actions, adjusting communication parameters (such as signal strength, frequency band), and changing the communication mode (such as switching to an alternate channel). In the implementation process of the above solution, through technical means such as disabling communication actions, enabling communication actions, adjusting communication actions, adjusting communication parameters (such as signal strength, frequency band), and changing the execution priority of the communication mode (such as switching to an alternate channel), the allocation and scheduling efficiency of communication resources are improved, thereby effectively improving the overall performance and response speed of the system.
[0012] The embodiments of the present application also provide an airport environment communication concurrency control device, including: an action selection and execution module, configured to select and execute a target action of airport equipment based on sample data through a policy neural network in a reinforcement learning model, and obtain device data collected by the airport equipment during the current communication cycle, where the sample data is collected by the airport equipment during the previous communication cycle; an action reward calculation module, configured to calculate a reward value corresponding to the target action according to the sample data and the device data; a neural network optimization module, configured to optimize a target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and the reward value corresponding to the target action, where the network architecture of the target neural network is the same as that of the policy neural network; a policy network assignment module, configured to assign the model parameters of the target neural network to the policy neural network to obtain a policy network model; and a communication policy control module, configured to use the policy network model to control the communication policy of the airport equipment, where the communication policy is to select a target action to be executed from an action space, and the action space includes multiple actions.
[0013] Optionally, in the embodiments of the present application, the airport environment communication concurrency control device further includes: a policy network initialization module, configured to initialize the policy neural network in the reinforcement learning model if it has not been initialized, and randomly select a target action of the airport equipment from the action space.
[0014] Optionally, in the embodiments of the present application, the action selection and execution module includes: a target action selection sub-module, configured to input the sample data into the policy neural network of the reinforcement learning model if the policy neural network in the reinforcement learning model has been initialized, so that the policy neural network selects a target action of the airport equipment from the action space; and a target action execution sub-module, configured to send the target action to the airport equipment so that the airport equipment executes the target action.
[0015] Optionally, in the embodiments of the present application, the action reward calculation module includes: a reward score increase sub-module, configured to increase a preset score to the original reward value if it is determined according to the sample data and the device data that the target action is completed on time, the task execution efficiency of the target action is greater than a preset efficiency threshold, and the resource utilization rate of the airport device is greater than a preset utilization rate threshold, so as to obtain the reward value corresponding to the target action.
[0016] Optionally, in the embodiments of the present application, the action reward calculation module includes: a reward score decrease sub-module, configured to decrease a preset score from the original reward value if it is determined according to the sample data and the device data that the target action is completed late, or the task execution efficiency of the target action is less than a preset efficiency threshold, or the resource utilization rate of the airport device is less than a preset utilization rate threshold, so as to obtain the reward value corresponding to the target action.
[0017] Optionally, in the embodiments of the present application, the neural network optimization module includes: a scheduling event generation sub-module, configured to generate a scheduling event according to the sample data, the device data, the target action, and the reward value corresponding to the target action; a scheduling event storage sub-module, configured to store the scheduling event into a replay memory in the reinforcement learning model; and a target network optimization sub-module, configured to optimize the target neural network in the reinforcement learning model through the scheduling events in the replay memory.
[0018] Optionally, in the embodiments of the present application, the neural network optimization module includes: a sensing data prediction sub-module, configured to input the sample data and the device data into a machine learning model, so that the machine learning model predicts the multi-source sensor data of the airport device in the next communication cycle to obtain prediction data; a scheduling event generation sub-module, configured to generate a scheduling event according to the prediction data, the sample data, the device data, the target action, and the reward value corresponding to the target action; a scheduling event storage sub-module, configured to store the scheduling event into a replay memory in the reinforcement learning model; and a target network optimization sub-module, configured to optimize the target neural network in the reinforcement learning model through the scheduling events in the replay memory.
[0019] Optionally, in the embodiments of the present application, the multiple actions include at least one of the following: disabling a communication action, enabling a communication action, adjusting the execution priority of a communication action, adjusting communication parameters (such as signal strength, frequency band), and changing a communication mode (such as switching to a standby channel).
[0020] The embodiments of the present application further provide an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are run by the processor, the above-described method is executed.
[0021] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the method described above.
[0022] An embodiment of the present application also provides a computer program product, including: a computer program or computer instructions, and when the computer program or computer instructions are run by a processor, they execute the method described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0024] Figure 1 A flowchart showing the process of the airport environment communication concurrency control method provided by the embodiment of the present application;
[0025] Figure 2 A communication strategy diagram showing data communication between airport equipment and a cloud server provided by the embodiment of the present application;
[0026] Figure 3 A structural diagram showing the airport environment communication concurrency control device provided by the embodiment of the present application;
[0027] Figure 4 A structural diagram showing the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. It should be understood that the drawings in the embodiments of the present application are only for the purpose of illustration and description, and are not used to limit the protection scope of the embodiments of the present application. In addition, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in the embodiments of the present application show the operations implemented according to some embodiments of the embodiments of the present application. It should be understood that the operations of the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the embodiments of the present application.
[0029] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the figures is not intended to limit the scope of the claimed embodiments of the present application, but merely represents selected embodiments of the present application.
[0030] It can be understood that "first" and "second" in the embodiments of the present application are used to distinguish similar objects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit to be different. In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects. The term "plurality" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups).
[0031] It should be noted that the airport environment communication concurrency control method provided by the embodiments of the present application can be executed by an electronic device, where the electronic device refers to a device terminal or a server with the function of executing a computer program. The device terminal is, for example: a smart phone, a personal computer, a tablet computer, a personal digital assistant or a mobile Internet device, etc. A server refers to a device that provides computing services through a network. Servers include, for example: x86 servers and non-x86 servers. Non-x86 servers include: mainframes, minicomputers and UNIX servers.
[0032] In the safety management and operation of modern airports, with the rapid development of the aviation industry and technological progress, modern airports are facing an increasingly complex communication environment. Airports usually need to conduct communication management and resource scheduling for many key devices (such as radar systems, navigation devices, ground support devices, etc.). However, traditional airport environment communication concurrency control systems mainly rely on preset rules and manual intervention. Although this method can meet the requirements to a certain extent, with the increase in the number of devices and resource limitations, it often cannot meet the requirements of resource dynamic scheduling for multi-tasking and concurrent execution of airport devices. Relying on preset rules and manual intervention cannot make full use of available resources, such as bandwidth, computing power, etc. Especially during peak hours, limited resources cannot be most effectively allocated, resulting in resource waste or bottleneck phenomena. Therefore, the current resource utilization efficiency of airport environment communication concurrency control is relatively low.
[0033] To improve the above problems, please refer to Figure 1Schematic flowchart of the airport environment communication concurrency control method provided by the embodiments of the present application; the main idea of the airport environment communication concurrency control method is that the policy neural network in the reinforcement learning model selects and executes the target actions of airport devices based on historical sample data, obtains the device data within the current communication cycle, calculates the reward value of the target actions based on the sample data and the device data, and then optimizes the target neural network in the reinforcement learning model. Finally, by assigning the model parameters of the target neural network to form a new policy network model, the communication policies of airport devices are dynamically adjusted, and the optimal target actions are selected from the action space to optimize the resource utilization efficiency and communication efficiency of airport device control. The implementation manners of the above airport environment communication concurrency control method may include:
[0034] Step S110: The policy neural network in the reinforcement learning model selects and executes the target actions of airport devices based on the sample data, and obtains the device data collected by the airport devices within the current communication cycle. The sample data is the data collected by the airport devices in the previous communication cycle.
[0035] The reinforcement learning model is a machine learning model that learns policies by interacting with the environment to maximize a certain cumulative reward. The above reinforcement learning model can adopt a deep Q-learning network (DQN) model, etc. In airport device management, the above reinforcement learning model is used to determine the best actions that devices should take in different states. Suppose there is a baggage conveyor system at an airport. The reinforcement learning model can learn how to adjust the conveyor speed during peak and off-peak hours to optimize energy consumption and baggage handling efficiency.
[0036] The policy neural network is a component in the reinforcement learning model, mainly responsible for the neural network that selects target actions according to the current state. It generates a policy by learning the relationship between states and actions in the environment, that is, the rule for selecting the optimal action in different states. The input of the above policy neural network can be the sample data collected by the environment as the state, and the output can be the probability distribution of actions. For example, in the lighting system of an airport terminal, the policy neural network can decide whether to turn on or off the lights in certain areas according to the current light intensity and flight schedule.
[0037] Sample Data refers to the data collected from airport equipment through multi-source sensors (such as cameras, radars, temperature and humidity sensors, etc.) in the complex communication environment at the airport during the past communication cycle. These data can include the number of device connections, bandwidth, memory, utilization rate of the Central Processing Unit (CPU), utilization rate of the Graphic Process Unit (GPU), etc., and can be used to train and update the policy neural network to enable it to better predict and select actions. For example, in the baggage conveyor system at the airport, the sample data can include the number of bags transported by the turntable conveyor belt per hour, the transfer time, and the failure rate of the turntable equipment within the past week.
[0038] The target action is the action selected by the policy neural network during the current communication cycle, aiming to optimize the performance or efficiency of airport equipment. For example, in the lighting system of the airport terminal, the target action can be to turn on the lights in some areas during a specific time period to facilitate customers to go through the formalities or board the plane, or, to turn off the lights in other areas during other time periods to save electricity.
[0039] Device Data is the data collected from airport equipment during the current communication cycle after the target action is executed, which reflects the state and performance of the equipment after the action is executed. For example, in the baggage conveyor system at the airport, the device data can be the number of bags transported by the turntable conveyor belt, the transfer time, and the failure rate of the turntable equipment after the instruction action of increasing the number or speed of the turntable of the baggage conveyor belt is executed.
[0040] The Communication Cycle refers to the time interval for data communication between airport equipment and the cloud server. During each communication cycle, multi-source sensors around the airport equipment will collect data and send the collected data to the cloud server. Then, the cloud server selects the target action to be executed according to the collected data through the policy neural network in the reinforcement learning model. Finally, the target action is sent to the edge device or airport equipment (the action sending process is not shown in the figure), and the edge device or airport equipment executes the target action. In the lighting system of the airport, the communication cycle may be once every 5 minutes, and the system will decide whether to turn on or off the lights in certain areas according to the current passenger flow, current light intensity, and flight schedule.
[0041] Step S120: Calculate the reward value corresponding to the target action according to the sample data and the device data.
[0042] It can be understood that the reward value corresponding to the target action is usually a value calculated based on sample data and device data, with the aim of optimizing the system's performance or achieving a specific goal. For example, in an elevator system at an airport, to reduce the waiting time of passengers, the elevator system equipment can be made to perform the target action of increasing the operating frequency of the elevator. The above-mentioned reward value is a quantitative indicator used to evaluate the effect or quality of the target action. The reward value is usually calculated according to a certain reward function preset in the reinforcement learning model, reflecting the contribution degree of the target action to the system goal.
[0043] Step S130: Optimize the target neural network in the reinforcement learning model based on the sample data, device data, target action, and the reward value corresponding to the target action. The network architecture of the target neural network is the same as that of the policy neural network.
[0044] The target neural network is also a component in the reinforcement learning model and is a neural network that generates target values or reference values. These target values or reference values are mainly used to estimate the expected value of future rewards or the improvement direction of the policy. For example, in a Deep Q-Network (DQN), this target neural network is used to stabilize the training process and reduce fluctuations in training through regular updates. The policy neural network is another component in the reinforcement learning model and is used to directly output the probability distribution of actions. The reinforcement learning model can select actions according to this distribution.
[0045] It can be understood that the network architecture of the target neural network is the same as that of the policy neural network, which means they have the same number of layers, the same number of neurons, and the same activation functions and other structures. In this way, the target neural network can better estimate future rewards or the improvement direction of the policy, thereby improving the performance and stability of the model.
[0046] Step S140: Assign the model parameters of the target neural network to the policy neural network to obtain the policy network model.
[0047] Model parameters refer to the learnable weights and biases in the neural network. These learnable weights and biases determine the output of the target neural network and / or the policy neural network. During the training process, these parameters are adjusted through optimization algorithms (such as gradient descent) to minimize the loss function. The so-called assignment usually means copying the value of one variable to another variable. The above-mentioned assignment operation means copying the parameters of the target neural network to the policy neural network to ensure the stability and consistency of the policy. After the assignment operation on the policy network model, the parameters of the policy neural network are consistent with those of the target neural network, thus forming a new policy network model. This new policy network model will be used for subsequent policy optimization or action selection.
[0048] Step S150: Use the policy network model to control the communication policy of the airport equipment. The communication policy is to select the target action to be executed from the action space, and the action space includes multiple actions.
[0049] Please refer to Figure 2 The schematic diagram of the communication policy for data communication between the airport equipment and the cloud server provided by the embodiment of the present application shown; the communication policy of the airport equipment refers to the rules and methods followed when the airport equipment and the cloud server conduct data communication in the complex communication environment of the airport to ensure the efficiency and reliability of communication. In the schematic diagram of the communication policy for data communication between the airport equipment and the cloud server, the airport equipment is located on the right side of the figure, representing the actual physical equipment to be monitored and controlled. The external environment of the airport equipment represents the environment that affects or is collected by the sensor. Multiple source sensors (labeled "sensor") are distributed around the airport equipment to collect data on the airport equipment and its external environment. Each sensor is responsible for collecting data from the airport equipment and the external environment and transmitting this data to the edge device. These multiple source sensor devices require a large amount of computing and bandwidth support. Multiple edge devices (labeled "edge device") receive data from the sensors. The edge devices perform preliminary processing on the data and upload the processed data to the central server. The central server receives data from multiple edge devices, performs further processing and storage, and is also responsible for uploading the processed data to the cloud server. The cloud server is located at the top of the figure and is the highest layer of the entire system. The cloud server receives data from the central server and may perform higher-level data analysis and decision-making, such as using the policy network model to control the communication policy of the airport equipment, including selecting the actions that the edge device needs to execute from the action space, and / or selecting the target action that the airport equipment needs to execute from the action space.
[0050] The action space refers to the set of all possible actions or operations in a certain decision-making problem. This action space defines the range of actions that the policy network model can select. Specifically, for example: during the communication process of the airport equipment, the above-mentioned action space may include actions such as "increasing the signal strength", "decreasing the signal strength", "switching the communication channel", "maintaining the current state", etc.
[0051] Suppose in the wireless communication network of an airport, the policy network model is monitoring the communication status of multiple airport devices. The current status shows that the signal strength in a certain area is weak and there is a certain degree of interference. Based on this status information, the policy network model selects "increase signal strength" and "switch communication channel" from the action space as potential target actions. After evaluation, the model decides to select "switch communication channel" as the target action because this can more effectively reduce interference and improve communication quality. In this way, the policy network model can dynamically adjust the communication strategies of airport devices to adapt to the changing communication environment and ensure efficient and reliable communication.
[0052] In the implementation process of the above solution, after the policy neural network of the reinforcement learning model performs actions based on the sample data collected in the previous communication cycle to obtain the device data of the current communication cycle, the reward value is calculated according to the sample data and the device data. Then, the target neural network is optimized and its model parameters are assigned to update the policy neural network, so that the reinforcement learning model can make optimal decisions based on historical experience and the current environmental state, and continuously collect the latest device data and calculate the reward value. This data-driven decision optimization and reward mechanism guidance enables the reinforcement learning model to accurately select the optimal action from the action space, allowing the system to quickly adapt to changes in the communication environment, timely adjust the communication strategy to cope with different demand fluctuations, and thus dynamically adapt to changes in device status, effectively improving the resource utilization efficiency of communication concurrency control in the airport environment.
[0053] As an alternative implementation of the above method for communication concurrency control in the airport environment, before the policy neural network in the reinforcement learning model selects and executes the target action of the airport device based on the sample data, it further includes:
[0054] Step S101: If the policy neural network in the reinforcement learning model has not been initialized, initialize the policy neural network and randomly select the target action of the airport device from the action space.
[0055] It can be understood that the structure of the above policy neural network can be a multi-layer perceptron (MLP) or a convolutional neural network (CNN). The specific structure depends on the complexity of the action space and state space of the airport device. The network structure can include an input layer, a hidden layer, and an output layer. In addition, an appropriate activation function can be selected for each layer of the policy neural network. Common activation functions include ReLU, Sigmoid, Tanh, etc. The choice of activation function will affect the non-linear expression ability of the network.
[0056] For example, in the implementation of step S101 above: If the policy neural network in the reinforcement learning model has not been initialized, the weights and biases of the policy neural network can be initialized. Common initialization methods include random initialization (such as random sampling from a uniform distribution or a normal distribution), Xavier initialization, or other initializations, which help avoid the problems of gradient vanishing or gradient explosion and ensure that the network can effectively learn in the initial stage of training. After the policy neural network is initialized, since the network has not been trained yet, it is impossible to select the optimal action based on the current state. Therefore, the system randomly selects an action from the action space as the target action of the airport equipment. The above random selection can be achieved through a uniform distribution or a weighted distribution, specifically depending on the nature of the action space. The above action space refers to the set of all possible actions that the airport equipment can perform. For example, for the lighting control system of the airport, the action space may include "turn on the lights", "turn off the lights", "adjust the brightness", etc. The policy neural network of the reinforcement learning model can be run on the cloud server. After the policy neural network is initialized, it can select the target action from the action space. Then, the policy neural network can send the target action to the airport equipment for execution. After the airport equipment executes the action, the system observes the changes in the external environment through the data collected by the sensors and records information such as the state and reward after the action is executed. These information will be used for the subsequent training of the policy neural network.
[0057] In the process of implementing reinforcement learning in the above solution, exploration and exploitation are two key concepts. Exploration refers to trying new actions to discover better strategies, while exploitation refers to selecting the current optimal action based on existing knowledge. In the initial stage, randomly selecting actions helps the system conduct sufficient exploration and avoid falling into local optimal solutions prematurely. As the model is continuously optimized, the system can gradually shift from exploration to exploitation, thereby achieving better performance in the long run. Therefore, by randomly selecting actions in the initial stage, the system can accumulate sufficient data in the initial stage, avoid the cold start problem, and provide a solid foundation for subsequent model optimization and policy adjustment. Among them, the cold start problem refers to the situation that the model cannot effectively learn due to lack of sufficient data in the initial stage.
[0058] As an alternative implementation of step S110 above, the implementation of selecting and executing the target action of the airport equipment based on sample data may include:
[0059] Step S111: Determine whether the policy neural network in the reinforcement learning model has been initialized.
[0060] For example, the implementation of the above step S111 is as follows: Use an executable program compiled by a preset programming language to determine whether the policy neural network has been initialized. The initialization operation may include setting the weights and biases of the neural network. These parameters are the basis for the neural network to learn. If the network has not been initialized, the system will execute the initialization step (as described in step S101).
[0061] Step S112: If the policy neural network in the reinforcement learning model has been initialized, input the sample data into the policy neural network of the reinforcement learning model so that the policy neural network selects the target action of the airport equipment from the action space.
[0062] For example, the implementation of the above step S112 is as follows: If it is confirmed that the policy neural network in the reinforcement learning model has been initialized, use the sample data collected in the previous communication cycle as the input and transfer it to the policy neural network of the reinforcement learning model. The policy neural network processes these input data and selects a target action from the action space according to the learned policy. The policy neural network usually outputs the probability distribution of the action, and the system can select the most appropriate action according to these probabilities. For example, using the ε-greedy greedy policy, the action with the highest probability can be selected most of the time, and other actions can also be randomly selected occasionally to explore new possibilities. In the implementation process of the above solution, since the policy neural network of the reinforcement learning model has strong generalization ability and can adapt to different airport equipment and scenarios, it can quickly process the input data and output the target action, enabling the airport equipment to respond to environmental changes in real time. This real-time performance is particularly important in high-dynamic and high-complexity scenarios such as airports, which can effectively handle emergencies or demand changes and improve the adaptability of various airport equipment communication scenarios.
[0063] Step S113: Send the target action to the airport equipment so that the airport equipment executes the target action.
[0064] For example, the detailed implementation of the above step S113 is as follows: After the policy neural network of the reinforcement learning model selects the target action of the airport equipment from the action space, send this action to the corresponding airport equipment or edge device for execution. After these devices execute the action, they can collect new data, which will be used in the subsequent learning and optimization process. The execution of these actions may involve direct communication with the device, such as through API calls or sending control commands. After the action is executed, the status changes of the airport equipment or edge device can be monitored, and new data of the device can be collected, which will be used as the input for the next round of learning. Through these steps, the above reinforcement learning model can continuously learn from the environment, optimize its decision-making strategy, thereby improving the operation efficiency and performance of the airport equipment. This data-driven method enables the system to adapt to changing environmental conditions and achieve dynamic optimization and adjustment.
[0065] As an alternative implementation of the above step S120, the implementation of calculating the reward value corresponding to the target action based on the sample data and the device data may include:
[0066] Step S121: If it is determined according to the sample data and the device data that the target action is completed on time, and the task execution efficiency of the target action is greater than the preset efficiency threshold, and the resource utilization rate of the airport device is greater than the preset utilization threshold, then increase the preset score on the original reward value to obtain the reward value corresponding to the target action.
[0067] The implementation of the above step S121 includes: In the design of the reward function for calculating the reward value, factors such as the on-time completion rate of the target action, the task execution efficiency of the target action, and the resource utilization rate of the airport device can be considered for reward or punishment, so as to guide the system to make optimal decisions in a complex communication environment. It can be understood that the input of the policy neural network in the above reinforcement learning model can be the sample data collected in the previous communication cycle and the real-time device data after executing the target action in the current communication cycle. The sample data may include the status information of the airport device (such as the working status of the device, resource usage information, etc.) and the environmental information (such as weather, flight schedule, etc.). The above real-time device data may include the task completion time, resource consumption, device status change, etc.
[0068] To determine whether the target action is completed within the specified time according to the sample data and the device data, it can be judged by comparing the actual completion time of the target action with the preset time threshold. After analyzing the actual completion time of the target action based on the sample data and the device data, if the actual completion time exceeds the preset time threshold, it is confirmed that the target action is not completed on time. Similarly, if the actual completion time does not exceed the preset time threshold, it is confirmed that the target action is completed on time. To calculate the task execution efficiency according to the sample data and / or the device data, the task execution efficiency can be calculated by the ratio of the task completion time to the resource consumption in the sample data and / or the device data. To calculate the resource utilization rate of the airport device after executing the target action according to the sample data and / or the device data, the resource utilization rate after executing the target action can be calculated by the ratio of the device resource usage amount to the total resource amount.
[0069] The above-mentioned basic reward value is calculated based on the task completion situation, task execution efficiency, and resource utilization rate. This basic reward value can be calculated according to a preset scoring rule. For example: the target action is completed on time +10 points, the task execution efficiency is greater than the preset efficiency threshold: +5 points, the resource utilization rate is greater than the preset utilization threshold: +5 points. If the target action is completed on time, and the task execution efficiency of the target action is greater than the preset efficiency threshold, and the resource utilization rate of the airport equipment is greater than the preset utilization threshold, an additional preset score (such as +10 points) can be added to the basic reward value to obtain the final reward value corresponding to the target action.
[0070] As an alternative implementation of the above step S120, the implementation of calculating the reward value corresponding to the target action based on the sample data and device data may include:
[0071] Step S122: If it is determined according to the sample data and device data that the target action is completed late, or the task execution efficiency of the target action is less than the preset efficiency threshold, or the resource utilization rate of the airport equipment is less than the preset utilization threshold, then a preset score is reduced from the original reward value to obtain the reward value corresponding to the target action.
[0072] It can be understood that the implementation of the above step S122 is similar to the implementation of the above step S121. For unclear parts, refer to the implementation of the above step S121. Here is a simple example to illustrate. For example: the target action is completed late: -5 points, the task execution efficiency is less than the preset efficiency threshold: -5 points, the resource utilization rate is less than the preset utilization threshold: -5 points. If the target action is completed late (-5 points), or the task execution efficiency of the target action is less than the preset efficiency threshold (-5 points), or the resource utilization rate of the airport equipment is less than the preset utilization threshold (-5 points), an additional preset score (such as -5 points) can be reduced from the basic reward value to obtain the final reward value corresponding to the target action. By introducing the resource utilization rate as the basis for reward calculation, the system can effectively avoid resource waste and ensure the full utilization of resources such as airport equipment. This mechanism helps to maximize the resource utilization efficiency on the premise of ensuring task completion. Further, by comprehensively considering factors such as delay, efficiency, and resource utilization rate, the system can better handle various abnormal situations and enhance the robustness of the system. This robustness helps to improve the stability and reliability of the system.
[0073] As an alternative implementation of the above step S130, the implementation of optimizing the target neural network in the reinforcement learning model based on the sample data, device data, target action, and the reward value corresponding to the target action may include:
[0074] Step S131: Generate a scheduling event based on sample data, device data, a target action, and a reward value corresponding to the target action.
[0075] It can be understood that after the execution of the target action, the system will collect relevant sample data, device data, the executed target action, and the reward value corresponding to this action. These data may include: sample data such as the status information of airport equipment (equipment working status, resource usage, etc.) and environmental information (weather, flight scheduling, etc.), device data such as task completion time, resource consumption, device status changes, etc., the specific target actions executed by airport equipment, such as "turn on the lights", "turn off the lights", etc., and the reward value calculated based on task completion, task execution efficiency, and resource utilization rate. It can be understood that by combining sample data and device data for calculating the reward value, the system can adaptively adjust the reward strategy according to the actual situation. This adaptability enables the system to remain efficient and stable in different operating environments and has strong robustness.
[0076] The implementation manner of the above step S131 is as follows: For example, after the execution of the target action, the sample data after the execution of the target action, the device data collected within the current communication cycle, the executed target action, and the reward value obtained for the target action can be stored in the replay memory (Replay Memory) in the reinforcement learning model. Alternatively, a tuple of a scheduling event can be generated based on the sample data, device data, target action, and the reward value corresponding to the target action. This tuple may include: the environmental state (State) before the execution of the action analyzed from the sample data and / or device data, the target action (Action), the reward value (Reward) obtained after the execution of the action, the environmental state after the execution of the action, and whether the current state is a termination state (i.e., whether the task is completed), etc. The above replay memory is also known as the experience replay buffer (ReplayBuffer). The data stored in the replay memory is mainly used to train the target neural network in the reinforcement learning model later. The above experience replay is to train the target neural network by storing and randomly sampling historical data, which helps to break the correlation between data and improve the stability and efficiency of training.
[0077] Step S132: Store the scheduling event in the replay memory in the reinforcement learning model.
[0078] The implementation of the above step S132 may include: storing the generated scheduling events in a replay memory, which is usually a fixed-size queue or buffer. The replay memory can adopt a first-in-first-out (FIFO) strategy to manage the stored events. When the memory reaches its maximum capacity, new scheduling events will overwrite the oldest events. In the specific practice process, the maximum capacity of the memory can also be set to ensure that it will not grow infinitely. During the training process, a batch of scheduling events (referred to as mini-batch) can be randomly sampled from the replay memory, and these data are used to update the parameters of the target neural network. Random sampling helps to break the correlation between data and improve the stability and efficiency of training. As the system runs, new scheduling events will be continuously added to the replay memory, and old events will be gradually replaced. Through the above steps, the system can effectively utilize historical data to optimize the reinforcement learning model, thereby improving the operation efficiency and performance of airport equipment. The use of the replay memory not only improves the data utilization rate but also enhances the generalization ability of the model, enabling it to better adapt to the complex and changeable airport environment.
[0079] Step S133: Optimize the target neural network in the reinforcement learning model through the scheduling events in the replay memory.
[0080] The implementation of the above step S133 is, for example: reading the scheduling events from the replay memory in a random sampling manner, and parsing out the sample data, device data, target actions, and the reward values corresponding to the target actions from the scheduling events. Then, the sample data, device data, target actions, and the reward values corresponding to the target actions are fed back to the target neural network to adjust the weights and biases of the target neural network, thereby optimizing the target neural network in the reinforcement learning model. In this way, the reinforcement learning model can continuously learn from the environment and gradually improve the accuracy and efficiency of decision-making.
[0081] In the above solution process, by storing and replaying diverse scheduling events in the replay memory, the model can be exposed to more different state-action pairs and reward value combinations. This diverse training data helps the model better learn the complex rules in the environment, thereby improving its generalization ability in practical applications. Then, the target neural network is optimized by randomly sampling and reading historical scheduling events, avoiding the model's over-reliance on recent data or data in a specific sequence. This random sampling mechanism can effectively reduce the correlation during the training process, thereby improving the training stability and convergence speed of the model.
[0082] As an alternative implementation of the above step S130, optimizing the target neural network in the reinforcement learning model based on the sample data, device data, target actions, and the reward values corresponding to the target actions may include:
[0083] Step S134: Input the sample data and device data into the machine learning model so that the machine learning model predicts the multi-source sensor data of the airport device in the next communication cycle to obtain prediction data.
[0084] The implementation of the above Step S134 may include: The electronic device can read the sorted sample data and device data, and preprocess the sample data and device data, such as normalization, denoising, filling missing values, etc., to ensure data quality. Then, after the above machine learning model is trained, input the sample data and device data into machine learning models such as linear regression, decision tree, support vector machine, random forest regression model, etc., and the above machine learning model can be selected according to the characteristics of the data and the prediction task. Input the sample data and device data collected by the multi-source sensor in the current communication cycle into the trained machine learning model so that the machine learning model predicts the multi-source sensor data of the airport device in the next communication cycle to obtain prediction data, and these prediction data may include device status, resource usage, task completion time, etc. Through the above steps, the machine learning model can effectively predict the multi-source sensor data of the airport device in the next communication cycle, provide more accurate environmental state information for the reinforcement learning model, and thus optimize the operation efficiency and performance of the airport device. The prediction data can also help to identify potential communication bottlenecks or device failures in advance, so as to adjust the communication strategy in advance and avoid resource waste and communication interruption.
[0085] Step S135: Generate a scheduling event according to the prediction data, sample data, device data, target action, and the reward value corresponding to the target action.
[0086] It can be understood that by combining the prediction data and historical data to generate scheduling events, the reinforcement learning model can adapt to environmental changes more quickly and make real-time adjustments. This enhancement of real-time performance and adaptability is particularly suitable for scenarios such as airport devices that require high reliability and fast response.
[0087] Step S136: Store the scheduling event into the replay memory in the reinforcement learning model.
[0088] It can be understood that scheduling events are generated according to the prediction data, sample data, device data, target action, and reward value, and these events are stored in the replay memory of the reinforcement learning model. This process not only enhances the model's ability to utilize historical data, but also enables the model to learn the correlation between prediction data and historical data from past experience based on the prediction data, further improving the accuracy and robustness of decision-making.
[0089] Step S137: Optimize the target neural network in the reinforcement learning model by replaying scheduling events in the memory.
[0090] It can be understood that the implementation manners of the above steps S135 to S137 are similar to those of the above steps S131 to S133. The difference is that the specific data included in the scheduling events is different. The scheduling events of steps S135 to S137 include prediction data, sample data, device data, target actions, and the reward values corresponding to the target actions, while the scheduling events of the above steps S131 to S133 include sample data, device data, target actions, and the reward values corresponding to the target actions. In the implementation process of the above solution, the reinforcement learning model is optimized through prediction data, enabling the reinforcement learning model to allocate resources more reasonably and reduce unnecessary resource waste. For example, in the scheduling of airport equipment, the operating mode of the equipment can be adjusted in advance according to the predicted device status and environmental changes, thereby reducing energy consumption and maintenance costs.
[0091] As an alternative implementation manner of the above method, the multiple actions include at least one of the following: disabling communication actions, enabling communication actions, adjusting the execution priority of communication actions, adjusting communication parameters (such as signal strength, frequency band), and changing communication modes (such as switching to a backup channel). By dynamically disabling or enabling communication actions, the system can flexibly adjust the allocation of communication resources according to the current load and requirements, avoiding resource waste or over-occupation, thereby improving resource utilization. In addition, by adjusting the execution priority of communication actions, the system can ensure that critical tasks or high-priority tasks can obtain communication resources first, reducing latency and enhancing the real-time performance and reliability of the system. Further, disabling relevant actions when communication is not required can reduce unnecessary system overheads (such as power consumption, bandwidth occupancy, etc.), thereby extending the service life of the device or reducing operating costs. In complex communication scenarios (such as multi-task concurrency, network fluctuations, etc.), the policy neural network model can improve the allocation and scheduling efficiency of communication resources through technical means such as disabling communication actions, enabling communication actions, and / or adjusting the execution priority of communication actions, thereby effectively improving the overall performance and response speed of the system.
[0092] Please refer to Figure 3 the structural schematic diagram of the communication concurrency control device in the airport environment provided by the embodiment of the present application shown; The embodiment of the present application provides a communication concurrency control device 200 for an airport environment, including:
[0093] An action selection and execution module 210, configured to select and execute a target action of an airport device based on sample data through a policy neural network in a reinforcement learning model, and obtain device data collected by the airport device during the current communication cycle, where the sample data is collected by the airport device during the previous communication cycle.
[0094] An action reward calculation module 220, configured to calculate a reward value corresponding to a target action according to sample data and device data.
[0095] A neural network optimization module 230, configured to optimize a target neural network in a reinforcement learning model based on sample data, device data, a target action, and a reward value corresponding to the target action, where the network architecture of the target neural network is the same as that of a policy neural network.
[0096] A policy network assignment module 240, configured to assign model parameters of the target neural network to the policy neural network to obtain a policy network model.
[0097] A communication policy control module 250, configured to control the communication policy of airport equipment using the policy network model, where the communication policy is to select a target action to be executed from an action space, and the action space includes multiple actions.
[0098] As an optional implementation manner of the above device, the airport environment communication concurrency control device further includes:
[0099] A policy network initialization module, configured to initialize the policy neural network in the reinforcement learning model if the policy neural network has not been initialized, and randomly select a target action of airport equipment from the action space.
[0100] As an optional implementation manner of the above device, the action selection and execution module includes:
[0101] A target action selection sub-module, configured to input the sample data into the policy neural network of the reinforcement learning model if the policy neural network in the reinforcement learning model has been initialized, so that the policy neural network selects a target action of airport equipment from the action space.
[0102] A target action execution sub-module, configured to send the target action to airport equipment so that the airport equipment executes the target action.
[0103] As an optional implementation manner of the above device, the action reward calculation module includes:
[0104] A reward score increase sub-module, configured to increase a preset score to the original reward value if it is determined according to the sample data and device data that the target action is completed on time, the task execution efficiency of the target action is greater than a preset efficiency threshold, and the resource utilization rate of the airport equipment is greater than a preset utilization rate threshold, to obtain a reward value corresponding to the target action.
[0105] As an optional implementation manner of the above device, the action reward calculation module includes:
[0106] The reward score reduction sub-module is used to reduce a preset score from the original reward value if it is determined that the target action is completed with a delay according to the sample data and the device data, or the task execution efficiency of the target action is less than the preset efficiency threshold, or the resource utilization rate of the airport device is less than the preset utilization threshold, so as to obtain the reward value corresponding to the target action.
[0107] As an optional implementation manner of the above device, the neural network optimization module includes:
[0108] The scheduling event generation sub-module is used to generate a scheduling event according to the sample data, the device data, the target action, and the reward value corresponding to the target action.
[0109] The scheduling event storage sub-module is used to store the scheduling event into the replay memory in the reinforcement learning model.
[0110] The target network optimization sub-module is used to optimize the target neural network in the reinforcement learning model through the scheduling events in the replay memory.
[0111] As an optional implementation manner of the above device, the neural network optimization module includes:
[0112] The sensing data prediction sub-module is used to input the sample data and the device data into a machine learning model, so that the machine learning model predicts the multi-source sensor data of the airport device in the next communication cycle to obtain prediction data.
[0113] The scheduling event generation sub-module is used to generate a scheduling event according to the prediction data, the sample data, the device data, the target action, and the reward value corresponding to the target action.
[0114] The scheduling event storage sub-module is used to store the scheduling event into the replay memory in the reinforcement learning model.
[0115] The target network optimization sub-module is used to optimize the target neural network in the reinforcement learning model through the scheduling events in the replay memory.
[0116] As an optional implementation manner of the above device, the multiple actions include at least one of the following: disabling the communication action, enabling the communication action, adjusting the execution priority of the communication action, adjusting the communication parameters (such as signal strength, frequency band), and changing the communication mode (such as switching to a standby channel).
[0117] It should be understood that the device corresponds to the above-mentioned airport environment communication concurrent control method embodiment, and can execute the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be referred to in the above description, and the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the device.
[0118] See also Figure 4 The electronic device 300 provided in the embodiment of the present application includes: a processor 310 and a memory 320, wherein the memory 320 stores machine-readable instructions executable by the processor 310, and when the machine-readable instructions are executed by the processor 310, the above method is executed.
[0119] The embodiment of the present application also provides a computer-readable storage medium 330, on which a computer program is stored, and the computer program is executed by the processor 310 to execute the above method. The computer-readable storage medium 330 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0120] The embodiment of the present application also provides a computer program product, including: a computer program or a computer instruction, and the computer program or the computer instruction executes the method described above when executed by a processor.
[0121] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0122] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, mainly depending on the functions involved.
[0123] In addition, the various functional modules in the embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part. Moreover, in the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0124] The above description is only an alternative implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the embodiments of the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the embodiments of the present application.
Claims
1. A method for concurrent control of airport environment communications, characterized in that: include: Selecting and executing the target action of the airport equipment based on the sample data through the policy neural network in the reinforcement learning model, and obtaining the equipment data collected by the airport equipment in the current communication cycle, wherein the sample data is collected by the airport equipment in the previous communication cycle; Calculating a reward value corresponding to the target action according to the sample data and the device data; Optimizing the target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and the reward value corresponding to the target action, wherein the network architecture of the target neural network is the same as the network architecture of the policy neural network; Assigning the model parameters of the target neural network to the strategy neural network to obtain a strategy network model; The strategy network model is used to control the communication strategy of the airport equipment, wherein the communication strategy is to select a target action to be executed from an action space, and the action space includes multiple actions.
2. The method according to claim 1, characterized in that: Before selecting and executing the target action of the airport equipment based on the sample data through the policy neural network in the reinforcement learning model, it also includes: If the policy neural network in the reinforcement learning model has not been initialized, the policy neural network is initialized and the target action of the airport equipment is randomly selected from the action space.
3. The method according to claim 1, characterized in that The strategy neural network in the reinforcement learning model selects and executes the target action of the airport equipment based on the sample data, including: If the policy neural network in the reinforcement learning model has been initialized, inputting the sample data into the policy neural network of the reinforcement learning model so that the policy neural network selects the target action of the airport equipment from the action space; The target action is sent to the airport equipment so that the airport equipment performs the target action.
4. The method according to claim 1, characterized in that: The calculating the reward value corresponding to the target action according to the sample data and the device data includes: If it is determined based on the sample data and the equipment data that the target action is completed on time, and the task execution efficiency of the target action is greater than the preset efficiency threshold, and the resource utilization of the airport equipment is greater than the preset utilization threshold, then the preset score is added to the original reward value to obtain the reward value corresponding to the target action.
5. The method according to claim 1, characterized in that The calculating the reward value corresponding to the target action according to the sample data and the device data includes: If it is determined based on the sample data and the equipment data that the target action is delayed in completion, or the task execution efficiency of the target action is less than a preset efficiency threshold, or the resource utilization of the airport equipment is less than a preset utilization threshold, then the preset points are reduced from the original reward value to obtain the reward value corresponding to the target action.
6. The method according to claim 1, characterized in that The optimizing the target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and the reward value corresponding to the target action includes: Generate a scheduling event according to the sample data, the device data, the target action and the reward value corresponding to the target action; Storing the scheduling event in a replay memory in the reinforcement learning model; The target neural network in the reinforcement learning model is optimized by scheduling events in the replay memory.
7. The method according to claim 1, characterized in that The optimizing the target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and the reward value corresponding to the target action includes: Inputting the sample data and the device data into a machine learning model so that the machine learning model predicts the multi-source sensor data of the airport equipment in the next communication cycle to obtain prediction data; Generate a scheduling event according to the prediction data, the sample data, the device data, the target action, and the reward value corresponding to the target action; Storing the scheduling event in a replay memory in the reinforcement learning model; The target neural network in the reinforcement learning model is optimized by scheduling events in the replay memory.
8. An airport environment communication concurrent control device, characterized in that: include: An action selection and execution module is used to select and execute a target action of the airport equipment based on sample data through a policy neural network in a reinforcement learning model, and obtain equipment data collected by the airport equipment in a current communication cycle, wherein the sample data is collected by the airport equipment in a previous communication cycle; An action reward calculation module, used to calculate a reward value corresponding to the target action according to the sample data and the device data; A neural network optimization module, configured to optimize a target neural network in the reinforcement learning model based on the sample data, the device data, the target action, and a reward value corresponding to the target action, wherein the network architecture of the target neural network is the same as the network architecture of the policy neural network; A strategy network assignment module, used to assign the model parameters of the target neural network to the strategy neural network to obtain a strategy network model; The communication strategy control module is used to control the communication strategy of the airport equipment using the strategy network model, wherein the communication strategy is to select a target action to be executed from an action space, and the action space includes multiple actions.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the machine-readable instructions are executed by the processor to perform any method according to claims 1 to 7.
10. A computer program product, characterized in that include: A computer program or a computer instruction, wherein when the computer program or the computer instruction is executed by a processor, the method according to any one of claims 1 to 7 is executed.