Method for controlling a controlled device by a control server using a QoS-based wireless communication system
By employing multiple QoS flows with varying levels, the method addresses congestion and resource inefficiencies in control applications, enhancing performance and availability in QoS-based wireless communication systems.
Patent Information
- Application Number
- JP2025545448
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-11
- Filing Date
- 2023-11-02
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2043-11-02
AI Technical Summary
Conventional control applications using QoS-based wireless communication systems face congestion and rejection due to allocating excessive radio resources to ensure adequate performance under demanding circumstances, leading to inefficiencies and reduced availability for other applications.
Establishing multiple QoS flows with different QoS levels for control commands, allowing the control server to dynamically select the appropriate flow based on current needs, reducing resource usage and congestion probability while maintaining desired Quality of Experience (QoE).
This approach effectively reduces congestion and rejection rates in wireless communication systems while ensuring QoE requirements are met, optimizing radio resource allocation.
Smart Images

Figure 2025536447000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to automatic control of one or more controlled devices by a control server, and more particularly to a control method and system that uses a Quality of Service (QoS)-based wireless communication system to exchange control data. Priority is claimed from European Patent Application No. 23305537.5, filed April 11, 2023, the contents of which are incorporated herein by reference. [Background technology]
[0002] Automatic control theory is the science that deals with methods for dealing with laws for controlling dynamic systems that can be realized by automatic devices, i.e., without human intervention.
[0003] At any given time, the controlled system is in an appropriate state space
number
number
[0004] Here, we focus on control applications, where wireless communication systems are used to connect plants and controllers, send control commands to actuators, and, in the case of closed-loop control systems, capture state measurements from sensors. Control applications can rely on any wireless communication technology, such as Wi-Fi, Bluetooth, or LTE. More recently, the 3GPP standardization organization specified 5G wireless communication systems, more specifically, Ultra-Reliable Low-Latency Communications (URLLC), which incorporate dedicated means to address control applications. Wireless communication systems are suitable for mobile or portable controlled devices, such as automated guided vehicles (AGVs) and autonomous mobile robots (AMRs), and any type of robot or vehicle. However, wireless communication systems are prone to packet loss due to radio wave propagation effects, and are also prone to congestion phenomena when timely packet transmission requires more radio resources than are available. As a result, the design of control applications must take into account the possibility of packet loss and / or transmission delays.
[0005] It is common to define a minimum performance level to be achieved by a control application, measured using application-related performance metrics. This is often described as the expected application-level Quality of Experience (QoE). Due to possible variations in the communication channel, the performance of a wireless link may change over time for a given communication parameter configuration. It is not the purpose of the control application to define the communication parameters of the wireless communication system. From the perspective of the control application, the wireless communication system often appears as a black box, randomly losing packets and / or delaying packet delivery. A common solution is to define a minimum communication-related performance level, called Quality of Service (QoS), to be achieved by the wireless communication system to reach the desired QoE. The role of the wireless communication system is then to adjust its configuration to meet the objectives regardless of changes in the communication channel. In such a wireless communication system, hereafter referred to as a QoS-based wireless communication system, packets are tagged with a QoS indicator or identifier that is used by the wireless communication system to achieve the desired QoS level.
[0006] Most control applications do not support QoE or QoS indicators that can be used to tag packets being transmitted. To address this type of situation, we focus here on wireless communication systems that rely on a QoS flow mechanism, where the wireless communication system establishes a logical channel within which all packets are transmitted according to the same QoS profile. Essentially, the control application pre-negotiates the establishment of a QoS flow, and then the wireless communication system automatically tags packets with the corresponding QoS indicators, using a filtering mechanism based, for example, on input and output IP or MAC addresses, application type, etc.
[0007] The negotiation phase, also known as admission control, checks whether the wireless communication system can implement the requested QoS profile with respect to its available resources and the current load of the wireless communication system. Summary of the Invention [Problem to be solved by the invention]
[0008] Prior to transmission, the control application must select a QoS profile (i.e., a QoS level) necessary to achieve the desired QoE. Conventional approaches select a QoS profile suitable for the most demanding circumstances because establishing a QoS flow takes time and therefore is performed only once before the transmission of control data. By doing so, the wireless communication system can enable the control application to achieve the desired QoE under all circumstances. However, applying a QoS profile suitable for such a demanding situation to ensure adequate performance for the control application at all times would result in the wireless communication system allocating more radio resources than necessary most of the time, causing congestion among the control applications sharing the same radio resources. Such congestion should be avoided because it may lead to a temporary reduction in the QoS level actually achieved. Applying such a QoS profile (suitable for the most demanding circumstances) at all times would also reduce the amount of radio resources available to other control applications applying for admission to the wireless communication system, potentially resulting in the admission control rejecting many control applications.
[0009] The present disclosure aims to improve that situation, in particular to at least partially address some or all of the limitations of the prior art mentioned above by proposing a solution that makes it possible, at least in some cases, to reduce the congestion probability and / or the refusal rate of admission control, while still making it possible to meet the QoE requirements of control applications in most cases. [Means for solving the problem]
[0010] To this end, according to a first aspect, the present disclosure relates to a method for controlling a controlled device, executed by a control server, for causing the controlled device to perform a task by sending control commands to the controlled device using a wireless communication system, the wireless communication system supporting the establishment of different Quality of Service QoS flows associated with respective different QoS levels of the wireless communication system, the control method comprising: establishing at least two QoS flows for transmitting control commands to a controlled device, the at least two QoS flows including a first QoS flow and a second QoS flow, the first QoS flow being associated with a QoS level greater than a QoS level associated with the second QoS flow; - determining a control command to be sent to the controlled device; selecting, from among the at least two established QoS flows, a QoS flow to be used for transmitting the control command; - transmitting a control command to at least one controlled device via a wireless communication system using the selected QoS flow.
[0011] Thus, a control server executing a control application establishes two or more QoS flows for the same control application (as opposed to the prior art, in which only one QoS flow is associated with a control application). Thus, from the perspective of the wireless communication system, the control server is seen as executing two or more different control applications having different QoS level requirements (i.e., different QoS profiles). The wireless communication system is tricked into thinking that the control server is executing two (or more) different control applications. Thus, no modifications to the wireless communication system are required, and the control application can associate different ports with the established different QoS flows, for example, to enable automatic filtering of traffic for the different QoS flows.
[0012] The established at least two QoS flows are associated with different QoS levels, and the QoS level provided by the first QoS flow is greater than the QoS level provided by the second QoS flow. A "greater QoS level" means that the QoS profile of the first QoS flow enables achieving better control command delivery performance than the QoS profile of the second QoS flow. The "control command delivery performance" relates to the performance of control command delivery and represents how reliably and / or quickly a control command can be delivered to a controlled device. In other words, the first QoS flow can be expected to deliver a control command to a controlled device more reliably (e.g., with a lower packet error rate due to propagation loss and / or lack of wireless resources) and / or faster (i.e., with a lower delay) than the second QoS flow. Typically, the QoS level of the first QoS flow corresponds to a QoS level suitable for the most demanding situation, i.e., a QoS level that is considered to enable a control application to achieve a desired QoE under all circumstances. On the other hand, the QoS level of the second QoS flow provides a lower control command delivery performance and cannot be considered to enable the control application to achieve the desired QoE under all circumstances.
[0013] Because the at least two established QoS flows are associated with the same control application, traffic generated by the control application is divided among the at least two QoS flows, and as a result, each established QoS flow handles less traffic than if only a single QoS flow were established. However, when using the second QoS flow, the wireless communication system may potentially allocate fewer radio resources for transmitting control commands than when using the first QoS flow. Thus, using the second QoS flow at least occasionally reduces the amount of radio resources used by the control application compared to always using the first QoS flow, and therefore reduces congestion probability. For example, if the first QoS flow is used by default and the QoS level of the second QoS flow enables the desired QoE to be achieved at the current stage of the task, the second QoS flow may be used only if a predetermined condition is confirmed (e.g., if congestion is likely to occur, i.e., a congestion-related condition). In another example, the second QoS flow may be used by default, and the first QoS flow may be used only if a predetermined condition is confirmed (e.g., if a task stage requires a better QoS level to achieve the desired QoE, i.e., a task stage-related condition).
[0014] As described above, in the prior art, the QoS level of a single QoS flow to be established is determined by considering the most demanding situation. However, depending on the nature of the control application, the QoS level required to achieve a desired QoE (in terms of control command delivery performance) may change over time. For example, the packet error rate (PER) required to control a moving robot is not the same when the robot is entering a straight line at a slow speed as when the robot is making a U-turn at a high speed. While on a straight line, the control commands do not change significantly over time, so the robot can tolerate more packet losses (if a control command is lost, the previous control command can be used). Therefore, by establishing at least two QoS flows each providing a different QoS level, it is possible to use the second QoS flow (and therefore use fewer radio resources than when using the first QoS flow) when the desired QoE can be achieved even with less stringent QoS level requirements.
[0015] Thus, the proposed solution makes it possible, at least in some cases, to reduce the amount of radio resources used by the control application of the control server, and therefore reduces the congestion probability due to lack of radio resources, even though the wireless communication system is a black box for the control server. The wireless communication system provides at least two different QoS flows to the control server, as if several different control applications were running.
[0016] In certain embodiments, the control method may further comprise one or more of the following optional features, considered alone or in any technically possible combination:
[0017] In certain embodiments, the task performed by the controlled device includes multiple successive stages with respective QoS level requirements associated therewith, and selecting a QoS flow includes identifying which of at least two established QoS flows is compatible with the current stage of the task in that it provides a QoS level that is compatible with the QoS level required for the current stage of the task, and the selected QoS flow matches the identified QoS flow that is compatible with the current stage of the task. For example, the selected QoS flow matches the established compatible QoS flow associated with the lowest control command delivery performance.
[0018] In certain embodiments, selecting a QoS flow comprises: assessing whether a wireless communication system is congested or on the verge of congestion; - In response to detecting that the wireless communication system is congested or on the verge of congestion, not considering the first QoS flow during QoS flow selection (at least if the QoS level of the first QoS flow is not required to achieve the desired QoE in the current or next stage of the task). In certain embodiments, when it is detected that the wireless communication system is congested or on the verge of congestion, the first QoS flow is not considered during the selection of a QoS flow only if the QoS level required for the current stage of the task performed by the controlled device is also provided by a QoS flow other than the first QoS flow of the at least two established QoS flows.
[0019] In certain embodiments, the first QoS flow is used only if it is determined that the QoS level required for the current stage of the task is not provided by another QoS flow of the at least two established QoS flows.
[0020] In certain embodiments, the control method further includes receiving control feedback from the controlled device, and the QoS flow is selected based on the received control feedback. In some non-limiting examples, the control feedback includes a measurement of a state of the controlled device related to the task, and the QoS flow is selected based on the received state measurement. In some non-limiting examples, the control command is determined based on the received state measurement.
[0021] In certain embodiments, the QoS flow selection uses a selection policy that includes a machine learning model trained using a reinforcement learning algorithm.
[0022] In particular embodiments, the machine learning model includes a neural network trained using a deep Q-learning reinforcement learning algorithm.
[0023] In certain embodiments, the reinforcement learning algorithm uses a reward function that combines a first term and a second term; The first term represents the control performance of the control server in controlling the controlled device, and the reward is reduced when the control performance is reduced. The second term decreases the reward when the first QoS flow is used compared to when the second QoS flow is used.
[0024] In particular embodiments, the reinforcement learning algorithm uses a reward function that returns a smaller reward when a first QoS flow is used compared to when a second QoS flow is used, and the training of the machine learning model is performed under the constraint that the resulting selection policy satisfies a control performance criterion.
[0025] In certain embodiments, the control server controls multiple controlled devices, and the machine learning model of the selection policy is trained using a multi-agent reinforcement learning algorithm.
[0026] According to a second aspect, the present disclosure relates to a computer program product comprising instructions that, when executed by at least one processor, configure the at least one processor to perform a control method according to any one of the embodiments of the present disclosure.
[0027] According to a third aspect, the present disclosure relates to a computer-readable storage medium comprising instructions that, when executed by at least one processor, configure the at least one processor to perform a control method according to any one of the embodiments of the present disclosure.
[0028] According to a fourth aspect, the present disclosure relates to a control server for controlling a controlled device, the control being aimed at causing the controlled device to perform a task, the control server comprising a processing circuit and a communication unit configured to perform a control method according to any one of the embodiments of the present disclosure.
[0029] According to a fifth aspect, the present disclosure relates to a control system comprising a control server according to any one of the embodiments of the present disclosure and at least one controlled device controlled by the control server via a wireless communication system.
[0030] In certain embodiments, the wireless communication system is a 5G wireless communication system. [Effects of the Invention]
[0031] According to aspects of the present disclosure, it is possible to reduce the congestion probability and / or reduce the rejection rate of admission control while still meeting the QoE requirements of the control application in most cases. The invention will be better understood on reading the following description, given by way of example and in no way limiting, and made with reference to the following figures: [Brief explanation of the drawings]
[0032] [Figure 1]FIG. 1 is a diagram of a control system and a wireless communication system. [Figure 2] FIG. 2 is a diagram of an exemplary embodiment of a controlled device of a control system. [Figure 3] FIG. 3 is a diagram of an exemplary embodiment of a control server of the control system. [Figure 4] FIG. 4 is a diagram illustrating the main steps of an exemplary embodiment of the control method. [Figure 5] FIG. 5 is a diagram illustrating the main steps of another exemplary embodiment of the control method. [Figure 6] FIG. 6 is a plot illustrating simulation results showing the control performance of the control method. DETAILED DESCRIPTION OF THE INVENTION
[0033] In these figures, the same reference numbers from one figure to another indicate the same or similar elements. For clarity, elements shown are not to scale unless otherwise noted.
[0034] Additionally, the order of steps depicted in these figures is provided for illustrative purposes only and is not intended to limit the present disclosure, which may also apply to the same steps performed in different orders.
[0035] Figure 1 schematically illustrates an exemplary embodiment of a control system 10. As illustrated by Figure 1, the control system 10 includes a control server 30 that controls a controlled system consisting of one or more controlled devices 20. The control aims to have the controlled system perform a given task, and the task is a representation of the state space of the controlled system.
number
[0036] The wireless communication system 70 is used to establish a wireless communication link with each controlled device 20. Control commands received from the control server 30 are forwarded to the controlled device 20 via the wireless communication link. Control feedback, if any, received from the controlled device 20 over the wireless communication link is forwarded to the control server 30. The wireless communication system 70 can use any wireless communication technology, and the selection of a particular wireless communication technology represents a non-limiting specific embodiment of the present disclosure. For example, the wireless communication system can use at least one of the following wireless communication technologies: Wi-Fi, Bluetooth, LTE, and 5G.
[0037] The wireless communication system 70 comprises one or more wireless nodes (e.g., base stations, access points, etc.) that establish wireless communication links with the controlled devices 20. The one or more wireless nodes form a radio access network (RAN) 72 of the wireless communication system 70. The wireless communication system 70 may comprise a core network (CN) 71 via which it can exchange data with, for example, the control server 30, e.g., control commands sent to the controlled devices 20 over the wireless communication link and control feedback received from the controlled devices 20 over the wireless communication link.
[0038] The wireless communication system 70 supports the establishment of different QoS flows associated with different QoS levels. As described above, a QoS flow corresponds to a logical channel associated with a respective QoS profile. Such QoS flows are used to exchange application-level traffic, and the QoS profile describes the QoS level that the wireless communication system 70 should (contractually) achieve for that QoS flow. For example, separate QoS flows are established for exchanging guaranteed bit rate (GBR) traffic and best-effort traffic, respectively. The QoS profile may define the QoS level to be achieved in terms of, for example, guaranteed bit rate, priority level, PER, delay (also known as packet delay budget (PDB)), etc. In principle, a user application is associated with a single QoS profile and exchanges application-level traffic using a corresponding single QoS flow. Once QoS flows are established for the desired QoS profiles / levels, the wireless communication system 70 is viewed by the control system 10 as a black box used to exchange application-level traffic. In particular, the control server 30 does not need to be concerned with the communication parameters (e.g., modulation, coding rate, etc.) and radio resources used on the wireless communication link, or with packets that may be lost. The control server 30 simply assumes that the requested QoS profile / level is implemented by the wireless communication system 70 .
[0039] 2 schematically represents an exemplary embodiment of a controlled device 20. The controlled device 20 may be any kind of device that can be remotely controlled, such as, for example, a robot in a factory, a manned or unmanned vehicle, etc.
[0040] 2, the controlled device 20 includes a wireless communication unit 22 for transmitting and receiving data to and from a RAN 72 of the wireless communication system 70. Thus, the wireless communication unit 22 supports a wireless communication technology (e.g., Wi-Fi, Bluetooth, LTE, 5G, etc.) used by the wireless communication system 70 to establish a wireless communication link with the RAN 72.
[0041] The controlled device 20 has a state space
number
[0042] The controlled device 20 also includes a processing circuit 21 connected to a wireless communication unit 22, one or more actuators 23, and, if present, one or more sensors 24. For example, the processing circuit 21 includes one or more processors and one or more memories. The one or more processors may include, for example, a central processing unit (CPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc. The one or more memories may include any type of computer-readable volatile and non-volatile memory (such as a magnetic hard disk, a solid-state disk, an optical disk, an electronic memory, etc.). The one or more memories may store a computer program product in the form of a set of program code instructions that are executed by the one or more processors to control the one or more actuators 23 based on control commands received from the control server 30 and, in some cases, to generate control feedback based on measurements provided by the one or more sensors 24.
[0043] FIG. 3 illustrates a schematic representation of an exemplary embodiment of a control server 30 .
[0044] 3 , the control server 30 includes a communication unit 32 for exchanging data with the wireless communication system 70 to exchange data with the controlled device 20 via the RAN 72 of the wireless communication system 70. Typically, the control server 30 can exchange data with the CN 71 of the wireless communication system 70, and the data is exchanged with the controlled device 20 via the RAN 72 of the wireless communication system 70. However, in some cases, the control server 30, like the controlled device 20, may be connected to the RAN 72 of the wireless communication system 70 via a wireless communication link. Thus, the communication unit 32 of the control server 30 may support any wired and / or wireless communication technology suitable for exchanging data with the CN 71 and / or RAN 72 of the wireless communication system 70.
[0045] The control server 30 also includes a processing circuit 31 connected to the communication unit 32. For example, the processing circuit 31 includes one or more processors and one or more memories. The one or more processors may include, for example, a CPU, a DSP, an FPGA, an ASIC, etc. The one or more memories may include any type of computer-readable volatile and non-volatile memory (e.g., a magnetic hard disk, a solid-state disk, an optical disk, an electronic memory, etc.). The one or more memories may store a control application in the form of a computer program product including a set of program code instructions executed by the one or more processors to remotely control one or more controlled devices 20 via the wireless communication system 70 (e.g., by generating control commands for the controlled devices 20 and, if present, by processing control feedback).
[0046] FIG. 4 illustrates in schematic form the main steps of an exemplary embodiment of a control method 40 performed by the control server 30 .
[0047] As shown in FIG. 4 , the control method 40 includes a step S40 of establishing at least two QoS flows for a control application to transmit control commands to the controlled device 20. The established at least two QoS flows are associated with different QoS levels and include a first QoS flow and a second QoS flow. The first QoS flow is an established QoS flow that provides the best QoS level among the established QoS flows, and therefore, the QoS level provided by the first QoS flow is greater than the QoS level provided by the second QoS flow. In other words, the control command is expected to be delivered to the controlled device 20 more reliably and / or faster when using the first QoS flow than when using the second QoS flow. For example, if the QoS profile / level is defined by, for example, a priority level, a PER, and a delay, - for two QoS profiles defined by the same PER and delay, the QoS profile with a higher priority level provides a higher QoS level (in terms of control command delivery performance) since control commands are prioritized over packets with a lower priority level exchanged via the wireless communication system 70 (and therefore are less likely to be rejected due to lack of radio resources, thereby increasing the reliability and speed of control command delivery); For two QoS profiles defined by the same priority level and delay, the QoS profile with a lower PER provides a greater QoS level (in terms of control command delivery performance) since control commands are delivered with less packet loss than the other QoS profile with a greater PER; For two QoS profiles defined by the same priority level and PER, the QoS profile with lower delay provides a greater QoS level (in terms of control command delivery performance), since control commands will be delivered faster than the other QoS profile with a greater delay, etc.
[0048] Thus, in this disclosure, (at least) two different QoS flows are established for the same control application, whereas in prior art solutions, a user application is theoretically associated with a single QoS profile and exchanges application-level traffic using a corresponding single QoS flow. Therefore, from the perspective of the wireless communication system, the control server 30 is considered to be running two different control applications with different QoS level requirements (i.e., different QoS profiles). For example, a first QoS flow may be established to exchange delay-critical (e.g., periodic) GBR traffic, and a second QoS flow may be established to exchange best-effort traffic. Typically, the QoS level of the first QoS flow corresponds to a QoS level suitable for the most demanding situation, i.e., a QoS level that is considered to enable the control application to achieve the desired QoE under all circumstances. On the other hand, the QoS level of the second QoS flow provides lower control command delivery performance and cannot be considered to enable the control application to achieve the desired QoE under all circumstances.
[0049] As illustrated by FIG. 4, the control method 40 includes a step S41 of determining a control command to be sent to the controlled device 20.
[0050] The control method 40 also includes a step S42 of selecting, from among the at least two established QoS flows, a QoS flow to be used for transmitting the control command.
[0051] Next, the control method 40 includes a step S43 of transmitting a control command to the controlled device 20 via the wireless communication system 70. During step S43, the control server 30 transmits the control command to the wireless communication system 70, instructing the wireless communication system 70 to transmit the control command to the controlled device using the selected QoS flow. For example, the control application may associate different ports (e.g., TCP or UDP ports) with the different established QoS flows, and the wireless communication system 70 may filter the control command based on the port on which it was received and forward the control command to the associated QoS flow.
[0052] Of course, steps S41, S42, and S43 may be repeated during a control operation until the task to be performed is completed. In some embodiments, the QoS flow selection step S42 may be performed for each control command generated. However, in other embodiments, the QoS flow selection step S42 may be performed less frequently, i.e., not each time a new control command is generated by the control application on control server 30.
[0053] In some embodiments, it is possible to establish more than two QoS flows during step S40. For example, it is possible to establish a third QoS flow having an associated QoS level that is lower than the QoS level of the first QoS flow and different (lower or higher) than the QoS level of the second QoS flow. Thus, the QoS flow selection step S42 has three different options for selecting a QoS flow for transmitting a control command, each associated with three different QoS levels (in terms of control command delivery performance). In the following, unless otherwise specified, it is non-limitingly assumed that only two QoS flows are established (i.e., the first QoS flow and the second QoS flow), or that at least the second QoS flow matches the established QoS flow with its associated lowest QoS level.
[0054] It should be noted that the first QoS flow typically uses more radio resources than the second QoS flow to exchange the same amount of control data. Therefore, the control application of the control server 30 may affect the amount of radio resources actually used to control the controlled device 20 by not always selecting the same QoS flow. If the first QoS flow corresponds to a QoS profile / level suitable for the most demanding situation, occasionally using the second QoS flow instead of the first QoS flow reduces the amount of radio resources used compared to the prior art described above.
[0055] In particular, it should be noted that the QoS level of the first QoS flow is not necessarily required to achieve the desired QoE. For example, a task performed by the controlled device 20 may include multiple consecutive stages, each with different characteristics that may enable the desired QoE to be achieved with a different QoS profile / level. For example, the QoS level (in terms of control command delivery performance) required to control a moving robot is not the same when the robot is entering a straight line at a slow speed or when the robot is making a U-turn at a high speed. Thus, moving straight at a slow speed and making a U-turn at a high speed correspond to a first stage and a second stage, respectively, that require different QoS levels to achieve the desired QoE. While the first QoS flow enables the desired QoE to be achieved in both the first and second stages, if the QoS level of the second QoS flow is compatible with the QoS level required for the first stage (i.e., if its control command delivery performance is considered sufficient to achieve the desired QoE), it may be advantageous to use the second QoS flow in the first stage to reduce the amount of radio resources used.
[0056] The QoS flow selection step S42 uses a selection policy to select the QoS flow to use. Various selection policies can be used by the control method 40, and the selection of a particular selection policy represents a non-limiting specific embodiment of the control method 40.
[0057] For example, the selection policy may consist in using the first QoS flow by default and using the second QoS flow only if one or more predetermined conditions are confirmed. Thus, such a selection policy emphasizes control performance by default using the QoS flow with the best QoS level, i.e., the first QoS flow. For example, conditions related to congestion in the wireless communication system 70 may be taken into account. In such a case, the QoS flow selection step S42 may include evaluating whether the wireless communication system 70 is congested or on the verge of congestion (i.e., likely to become congested in the near future), and if it is detected that the wireless communication system 70 is congested (or on the verge of congestion), the second QoS flow (or another established QoS flow, if any, associated with a lower QoS level) is used instead of the first QoS flow. In such an embodiment, the control server 30 can benefit from the best QoS level guaranteed by the first QoS flow, unless this can no longer be provided by the wireless communication system 70. This may further be subject to conditions regarding the QoS level required for the current phase of the task (task phase-related conditions). Thus, if congestion is detected, the second QoS flow may be used only if its QoS level matches the QoS level required to execute the current stage of the task (i.e., if its control command delivery performance allows the desired QoE to be achieved). If not, the first QoS flow continues to be used. Note that when the control server 30 controls multiple controlled devices 20, other controlled devices 20 may be in respective stages that require a lower QoS level. Thus, these other controlled devices 20 may switch to the second QoS flow to overall reduce the amount of radio resources used. Thus, the selection policy may consider one or more congestion-related and / or task stage-related conditions to determine whether to switch from the first QoS flow to another QoS flow.
[0058] Congestion in the wireless communication system 70 can be detected using any method known to those skilled in the art, and the selection of a particular congestion detection method represents a specific, but non-limiting, embodiment of the present disclosure. For example, the control server 30 can detect that the wireless communication system 70 is congested (or on the verge of congestion) based on, for example, control feedback received (or not received) from the controlled device 20. If no control feedback is received, or if the received control feedback indicates that one or more control commands have not been received by the controlled device 20, the control server 30 can consider that the wireless communication system 70 is likely congested. According to another example, the control server 30 can detect that the wireless communication system 70 is likely congested based on, for example, reports from the wireless communication system 70. In fact, the wireless communication system 70 can provide reports, for example, regarding lost packets, which can be useful in detecting congestion. Even if such reports are not received in real time, they can be useful, for example, in building a temporal congestion model that represents the congestion probability as a function of time. Indeed, in some cases, some congestion events may occur, for example, in a roughly periodic manner, and such roughly periodic behavior can be detected using received reports and a temporal congestion model can be used to predict when the next congestion event is likely to occur.
[0059] In another example, the selection policy may consist in using the second QoS flow by default and using the first QoS flow only if one or more predetermined conditions are confirmed. Thus, such a selection policy emphasizes minimizing the amount of radio resources used by the control application (and thus reducing congestion probability) by using the QoS flow requiring the least radio resources, e.g., the second QoS flow, by default. For example, conditions related to the QoS level required to execute the current phase of the task may be considered (task phase-related conditions). In such a case, the QoS flow selection step S42 may include determining the QoS level required for the current phase of the task and identifying, from among the established QoS flows, a QoS flow that matches the QoS level required for the current phase of the task in terms of providing control command delivery performance that enables the achievement of a desired QoE during the current phase of the task. For example, the selected QoS flow will match the established conforming QoS flow with the lowest associated QoS level, so that the first QoS flow will only be used if absolutely necessary to comply with the QoS level requirements (i.e., if the desired QoE is not considered achievable using an established QoS flow other than the first QoS flow).
[0060] For example, multiple stages of a task of each controlled device 20 may be known in advance, along with associated QoS level requirements. In such a case, the control server 30 simply detects the current stage of the task being performed by the controlled device 20 and selects a QoS flow that matches the detected current stage of the task. This is possible, for example, when a predetermined series of stages must be performed to execute the task, so the control server 30 can predict which is the current stage based, for example, on the duration of each stage or the state position of the controlled device 20 along a predetermined trajectory. In the case of an open-loop control system 10, the control server 30 can predict the current stage of the task being performed by the controlled device 20 based, for example, on a control command sent to the controlled device 20. In the case of a closed-loop control system 10, the control server 30 can detect the current stage of the task being performed by the controlled device 20 based, for example, on control feedback received from the controlled device 20 (e.g., based on state measurements included in the control feedback).
[0061] 5 schematically illustrates the main steps of a particular embodiment of the control method 40. In addition to the steps discussed in connection with FIG. 4, the control method 40 also includes a step S44 of receiving control feedback from the controlled device 20. The control feedback includes, for example, measurements of the state of the controlled device 20 related to the task provided by one or more sensors 24 and / or an indication of the control command received by the controlled device 20. Such control feedback exists in the case of a closed-loop control system 10 so that the control server 30 knows the current state of the controlled device 20 and adjusts the determined control command accordingly. However, such control feedback may also exist in the case of an open-loop system so that the control server 30 knows the current state of the controlled device 20 and determines the required QoS level in light of the current state of the controlled device 20. Indeed, such state measurements can be useful to take into account unexpected situations, such as the presence of an unexpected obstacle that needs to be avoided (which can be detected, for example, if one of the sensors 24 is a camera monitoring the environment of the controlled device 20). It is therefore emphasized that the multiple successive stages of the controlled device 20 are not necessarily known in advance, at least some of them may be unexpected, and that they need to be detected and the QoS level adjusted by selecting an appropriate QoS flow from among the first QoS flow and the second QoS flow.
[0062] More generally, the selection policy selects the QoS flows to be used according to the situation of the control system 10. For example, the selection may be performed based on a set of variables called QoS flow selection variables, which may include the current state of the controlled system (which can be measured and / or predicted), the position of the controlled device 20, the target trajectory, the distance to the nearest obstacle, etc. The domain described by these QoS flow selection variables can be divided into two regions (for the first QoS flow and the second QoS flow, respectively), typically from the calculation of a norm function along the QoS flow selection variables and comparison with a predetermined decision threshold. Finally, for each control command to be sent, the control application selects a QoS flow according to the region into which the value of the QoS flow selection variable falls.
[0063] It is emphasized that, although used to configure the parameters of the wireless communication system 70 (i.e., the first QoS flow or the second QoS flow), these QoS flow selection variables are primarily related to the state of the control system 10 and not the state of the wireless communication system 70 (except when attempting to detect and / or avoid congestion). Thus, the selection policy is primarily (or in some cases exclusively) driven by the state of the controlled device 20, which is independent of the state of the wireless communication system 70.
[0064] In particular embodiments, the selection policy includes a machine learning model (e.g., a neural network) that receives values of the QoS selection variables as input and outputs a selected QoS policy, where the machine learning model is pre-trained using reinforcement learning. In other words, the selection policy is implemented by the machine learning model, and the selection policy is optimized using a reinforcement learning algorithm.
[0065] First, we will note that the problem we are trying to solve can be modeled as a Markov Decision Process (MDP), an agent-environment framework in which an agent sequentially interacts with its environment by performing actions and receiving from the environment observations of the next state and rewards for these actions. At each iteration, the agent must choose an action among a finite set of predefined actions. With each action, the agent moves to another state with a given probability that reflects the influence of the environment. The goal in an MDP is to find a good policy for the decision maker:
number
number
number
number
[0066] More formally, an MDP is a set of tuples
number
number
number
number
number
number
number
[0067] Reward is state
number
number
number
number
[0068] Solving an MDP amounts to finding a policy that maximizes the expected return from the current state.
number
number
number
number
number
[0069] In this case, an agent (i.e., a controlled device 20, such as a robot) resides in an environment (e.g., a factory) incorporating a wireless communication system 70. At each iteration, a decision maker (i.e., a control server 30) must decide whether to send control commands using a first QoS flow or a second QoS flow. A key requirement is to define a reward function given by the environment for each action. This will be discussed later. Another important aspect of MDPs is the state transition function, which characterizes what the environment will do next. In our case, the wireless communication system 70 is typically used as a black box, and therefore, we have no prior knowledge of its statistical characterization. Therefore, a solution in such situations is to resort to the reinforcement learning (RL) paradigm, a branch of machine learning that addresses how to take actions in an environment to maximize a cumulative reward.
[0070] Reinforcement learning can be combined with function approximation to deal with problems with a very large number of states. The most common RL algorithm is the so-called Q-learning algorithm, a model-free algorithm for learning the value of an action in a given state.
number
[0071] Q-learning algorithms do not require a model of the environment (hence "model-free") and can handle problems with stochastic transitions and rewards without the need for adaptation. At its simplest, Q-learning stores state value data in a table. This approach becomes unreliable as the number of states / actions increases, as the agent becomes less and less likely to visit a particular state and perform a particular action. In such situations, Q-learning can be combined with function approximation, which allows the algorithm to be applied to larger problems, even if the state space is continuous. One solution is to use a neural network (NN) as a function approximator, leading to the so-called deep Q-learning reinforcement learning.
[0072] In a preferred embodiment, a deep Q-learning RL algorithm is applied to choose an optimal selection policy for selecting between at least two established QoS flows. Using this approach, the control server 30 divides the learning state space into at least two regions in an optimal manner with respect to a long-term cumulative reward function.
number
[0073] As mentioned above, the setting of the reward function is important in that it influences the optimization of the selection policy.
[0074] According to the first example, the reward function can be set to combine two objectives: keeping the control performance within given limits and, whenever possible, prioritizing the selection of the second QoS flow over the selection of the first QoS flow, at least in case of congestion. This corresponds to a multi-objective MDP problem that can be solved using existing well-defined but complex techniques. It is also possible to resort to ad-hoc techniques where the reward function combines the first and second terms. The first term represents the control performance of the control server 30 in controlling the controlled device 20, and the reward is reduced when the control performance is reduced. The second term decreases the reward when the first QoS flow is used compared to when the second QoS flow is used.
[0075] Thus, the first term allows increasing the reward if the control performance is improved, but a larger reward is obtained if the control performance is improved while using the second QoS flow.
[0076] By way of non-limiting example, the reward function can be calculated as follows:
number
number
number
number
number
[0077] In the above, the control state
number
number
number
number
number
number
number
[0078] for example,
number
number
number
number
number
number
number
number
number
number
[0079] In this example, the learning state
number
number
number
[0080] By way of another non-limiting example, the reward function can be calculated as follows:
number
number
number
number
number
number
number
number
number
number
number
number
number
[0081] In the above example, the learning state
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0082] Each of the reward functions defined above combines the dual objective of maximizing control performance and minimizing congestion levels by prioritizing secondary QoS flows. However, it does not check whether the desired QoE is achieved. Another example proposes relying on the constrained multi-layer programming (COMDP) method. In COMDP, the optimal selection policy is searched for over possible selection policies, defined as follows:
number
number
number
number
[0083] For example, the reward and cost functions can be defined as follows:
number
number
number
number
number
number
number
[0084] Figure 6 shows a schematic representation of simulation results that demonstrate how a selection policy determined by reinforcement learning can improve control performance. In Figure 6, a basic proportional controller is used to track a sinusoidal reference signal. More specifically, Part a) of Figure 6 represents the case where the first QoS flow is always used, which basically corresponds to the prior art described above, Part b) of Figure 6 represents the case where two QoS flows are established, the first one being used by default and the second one being used when congestion is detected; Part c) of Figure 6 shows the case where two QoS flows are established, the first one being used by default and the second one being used when congestion is detected and the task stage is compatible with the use of the second QoS flow, with the selection policy determined by reinforcement learning.
[0085] In each of parts a), b), and c), the wireless communication system 70 is initially uncongested and at some point becomes congested, and the congestion is detected by the control server 30. When the wireless communication system 70 is congested, control commands are sent with errors and the PER increases from 0.1 to 0.8. If a packet is lost, the last control command received by the controlled device 20 is used.
[0086] As shown by part a) of Figure 6, if the first QoS flow is always used, the controlled device 20 will track the sinusoidal reference signal, but at the cost of using more radio resources of the wireless communication system 70, which may affect the communication performance of other devices and may cause other devices to be denied access to the wireless communication system 70.
[0087] In part b) of Figure 6, the second QoS flow is used when congestion is detected. This reduces the amount of radio resources used, which may prevent other devices from being denied access to the wireless communication system 70. However, while the control performance does not degrade in the linear portion of the sinusoidal reference signal, this is not the case in the curved portion of the sinusoidal reference signal, and the desired QoE is no longer achieved. Thus, part b) of Figure 6 highlights different stages (i.e., the linear portion and the curved portion) of the task to be performed (i.e., tracking the sinusoidal reference signal) that require different QoS levels (in terms of control command delivery performance).
[0088] In part c) of Figure 6, the selection policy determined by reinforcement learning is applied. As seen in part c) of Figure 6, when congestion is detected, the selection policy switches between the first QoS flow and the second QoS flow. More specifically, the selection policy selects the first QoS flow for the curved portion of the sinusoidal reference signal and selects the second QoS flow for the linear portion of the sinusoidal reference signal, thereby achieving a desired QoE over the entire control period while using fewer radio resources when the wireless communication system 70 is congested (during the linear portion).
[0089] It is emphasized that the present disclosure is not limited to the above exemplary embodiments, and variations of the above exemplary embodiments are also within the scope of the present invention.
[0090] For example, the above exemplary embodiments are provided primarily considering a first QoS flow and a second QoS flow. However, as mentioned above, it is also possible to establish more than two QoS flows for a control application of the control server 30. For example, it is possible to establish a third QoS flow having an associated QoS level that is lower than the QoS level of the first QoS flow and different (lower or higher) than the QoS level of the second QoS flow. Thus, the QoS flow selection step S42 has three different options for selecting a QoS flow for transmitting a control command, each associated with three different QoS levels.
[0091] Furthermore, the above examples involving reinforcement learning have been provided primarily considering the case where the control server 30 controls a single controlled device 20 (possibly sharing a wireless communication system 70 with other devices). A key assumption for finding an optimal solution in the context of MDP is the stationarity of the environment. Essentially, the state transition matrix is assumed to be constant over time. If the controlled device 20 and other devices sharing the environment (especially the wireless communication system 70) are homogeneous, the stationarity principle no longer holds, and Markov property no longer holds. A first solution consists in combining all controlled devices 20 into a single control system 10, but this increases computational complexity with the number of controlled devices 20. This also requires training the control system 10 for every possible combination of the individual controlled devices 20. Another solution is to resort to multi-agent reinforcement learning (MARL), a subfield of reinforcement learning. MARL focuses on dealing with the behavior of multiple learning agents coexisting in a shared environment. [Industrial Applicability]
[0092] The present disclosure is applicable to control systems that include a control server and one or more mobile or portable controlled devices, such as any type of robot or vehicle, such as an automated guided vehicle (AGV) or an autonomous mobile robot (AMR).
Claims
1. 1. A method (40) for controlling a controlled device (20) performed by a control server (30), wherein the control is aimed at causing the controlled device to perform a task by sending control commands to the controlled device using a wireless communication system (70), the wireless communication system supporting the establishment of different Quality of Service QoS flows associated with respective different QoS levels, the control method (40) comprising: establishing (S40) at least two QoS flows for transmitting control commands to said controlled device, said at least two QoS flows comprising a first QoS flow and a second QoS flow, said first QoS flow being associated with a QoS level greater than a QoS level associated with said second QoS flow; - determining the control command to be sent to said controlled device (S41); - selecting (S42) from among said at least two established QoS flows a QoS flow to be used for transmitting said control command; - transmitting (S43) said control command to at least one of said controlled devices via said wireless communication system using said selected QoS flow, The at least two QoS flows include a first QoS flow and a second QoS flow, the first QoS flow being associated with a QoS level that is greater than a QoS level associated with the second QoS flow. Method (40).
2. 2. The control method (40) of claim 1, wherein the task performed by the controlled device includes a plurality of successive stages each having associated QoS level requirements, and wherein selecting the QoS flow includes identifying which of the at least two established QoS flows is compatible with the current stage of the task in that it provides a QoS level that is compatible with the QoS level required for the current stage of the task, and the selected QoS flow corresponds to the QoS flow identified as compatible with the current stage of the task.
3. selecting the QoS flow assessing whether said wireless communication system is congested or on the verge of congestion; - not considering the first QoS flow during the selection of the QoS flows in response to detecting that the wireless communication system is congested or about to be congested.
4. 4. The control method (40) of claim 3, wherein in response to detecting that the wireless communication system is congested or on the verge of congestion, the first QoS flow is not considered during the selection of the QoS flow only if a QoS level required for a current stage of the task performed by the controlled device is also provided by a QoS flow other than the first QoS flow of the at least two established QoS flows.
5. 3. The control method (40) of claim 2, wherein the first QoS flow is used only if it is determined that the QoS level required for the current stage of the task is not provided by another QoS flow of the at least two established QoS flows.
6. 6. The control method (40) of claim 1, further comprising receiving control feedback (S44) from the controlled device, wherein the QoS flow is selected based on the received control feedback.
7. The control method (40) of claim 6, wherein the selection of the QoS flow uses a selection policy that includes a machine learning model trained using a reinforcement learning algorithm.
8. The control method (40) of claim 7, wherein the machine learning model comprises a neural network trained using a deep Q-learning reinforcement learning algorithm.
9. the reinforcement learning algorithm uses a reward function that combines a first term and a second term; the first term represents a control performance of the control of the controlled device by the control server, and the reward is reduced when the control performance is reduced; The control method (40) according to claim 7 or 8, wherein the second term reduces the reward when the first QoS flow is used compared to when the second QoS flow is used.
10. 9. The control method (40) of claim 7 or 8, wherein the reinforcement learning algorithm uses a reward function that returns a smaller reward when the first QoS flow is used compared to when the second QoS flow is used, and the training of the machine learning model is performed under the constraint that the resulting selection policy satisfies a control performance criterion.
11. 9. The control method (40) of claim 7 or 8, wherein the control server controls a plurality of controlled devices, and the machine learning model of the selection policy is trained using a multi-agent reinforcement learning algorithm.
12. A computer program product comprising instructions that, when executed by at least one processor, configure the at least one processor to perform the control method (40) of any one of claims 1 to 11.
13. A computer-readable storage medium comprising instructions that, when executed by at least one processor, configure the at least one processor to perform the control method (40) of any one of claims 1 to 11.
14. A control server (30) for controlling a controlled device (20), the control being aimed at causing the controlled device to perform a task, the control server (30) comprising a processing circuit (31) and a communication unit (32) configured to execute a control method (40) according to any one of claims 1 to 11.
15. A control system (10) comprising a control server (30) according to claim 14 and at least one controlled device (20) controlled by said control server via a wireless communication system (70).
Citation Information
Patent Citations
Machine learning based adaption of qoe control policy
WO2021013368A1
TECHNIQUE FOR PERFORMING QoS CONTROL IN A CLOUD ROBOTICS SYSTEM
WO2022105992A1