CONTROL DEVICE, WIRELESS COMMUNICATION TERMINAL, CONTROL METHOD, AND CONTROL PROGRAM
The control device and method enhance wireless communication quality for terminals by using a control model to optimize operation and communication settings, addressing challenges related to environmental and operational changes.
Patent Information
- Application Number
- JP2022016270
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-04
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2042-02-04
AI Technical Summary
Existing wireless communication systems face challenges in maintaining communication quality due to changes in directional antenna orientation, propagation environments, number of terminals, and surrounding objects, which affect traffic and terminal movement.
A control device and method that collect terminal information, input it into a control model for operation and communication setting control, and evaluate the control information to update the model, thereby improving wireless communication quality while maintaining task efficiency.
The solution effectively improves wireless communication quality for terminals while maintaining task efficiency, addressing issues related to communication quality and terminal operation.
Smart Images

Figure 0007672108000001 
Figure 0007672108000002 
Figure 0007672108000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a control device, a wireless communication terminal, a control method, and a control program. [Background technology]
[0002] The realization of the Internet of Things (IoT), in which various devices are connected to the Internet, is progressing, and various devices such as automobiles, drones, and construction machinery vehicles are being connected wirelessly. Wireless communication standards such as the wireless LAN (Local Area Network) defined by the standard IEEE 802.11, Bluetooth (registered trademark), cellular communication using LTE and 5G, LPWA (Low Power Wide Area) communication for IoT, ETC (Electronic Toll Collection System) used for vehicle communication, VICS (Vehicle Information and Communication System), and ARIB-STD-T109 are also developing and are expected to become more widespread in the future. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] CJ Lowrance, AP Lauf, “An active and incremental learning framework for the online prediction of link quality in robot networks,” Engineering Applications of Artificial Intelligence, 77, pp.197-211, 2018. Summary of the Invention [Problem to be solved by the invention]
[0004] While wireless communication is used for various purposes, it is problematic that wireless communication does not necessarily satisfy the communication quality requirements for some services. In particular, the movement of terminals and surrounding objects changes the direction of antenna directivity, the propagation environment, the number of terminals communicating with the base station, and the amount of traffic, which has until now been unavoidable, affecting communication quality.
[0005] In Non-Patent Document 1, communication quality is predicted using distance information between a robot and a base station. When there are multiple controllable terminals such as robots, the positions and communication settings of each terminal greatly affect communication quality, so an efficient control method is required.
[0006] When there are multiple terminals equipped with wireless communication functions, and there are requirements regarding communication quality such as the capacity, data rate, delay time, and packet loss rate of the wireless communication system, there are problems such as the fact that the communication quality does not meet the expected performance or that conditions exist that result in poor performance depending on the location, distribution, and operation of the terminals.
[0007] The present invention has been made in consideration of the above circumstances, and an object of the present invention is to improve the quality of wireless communication of a terminal while maintaining task efficiency of work by the terminal. [Means for solving the problem]
[0008] In order to achieve the above-mentioned object, one aspect of the present invention is a control device that controls a terminal that performs work involving physical movements, the terminal communicating wirelessly with a wireless communication device, the control device comprising: a collection unit that collects terminal information including status information and communication information of the terminal; a control unit that inputs the terminal information into a control model and controls the operation and communication settings of the terminal based on control information regarding the operation and communication settings of the terminal output by the control model; and an evaluation unit that calculates an evaluation value of the control information using the terminal information after control by the control unit and updates the control model so that the evaluation value of the control information output by the control model is higher.
[0009] One aspect of the present invention is a wireless communication terminal that performs a task involving physical operations, and includes an acquisition unit that acquires terminal information including status information and communication information of the wireless communication terminal, a control unit that inputs the terminal information into a control model and controls the operation and communication settings of the wireless communication terminal based on control information regarding the operation and communication settings of the wireless communication terminal output by the control model, and an evaluation unit that calculates an evaluation value of the control information using the terminal information after being controlled based on the control information, and updates the control model so that the evaluation value of the control information output by the control model is higher.
[0010] One aspect of the present invention is a control method performed by a control device to control a terminal that performs a task involving physical movement, the terminal communicating wirelessly with a wireless communication device, and the control device performing the following steps: a collection step of collecting terminal information including status information and communication information of the terminal; a control step of inputting the terminal information into a control model and controlling operation and communication settings of the terminal based on control information regarding the operation and communication settings of the terminal output by the control model; and an evaluation step of calculating an evaluation value of the control information using the terminal information after control in the control step, and updating the control model so that the evaluation value of the control information output by the control model is increased.
[0011] One aspect of the present invention is a control program that causes a computer to function as the control device. Effect of the Invention
[0012] According to the present invention, it is possible to improve the quality of wireless communication with a terminal while maintaining the task efficiency of work performed by the terminal. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 illustrates an example of a configuration of a terminal control system according to an embodiment. [Diagram 2] 4 is a flowchart showing the operation of the control device. [Diagram 3] FIG. 13 is a schematic diagram showing a communication area in a simulation of luggage transportation. [Figure 4] FIG. 4 is a diagram showing coordinates of a two-dimensional plane in the simulation of FIG. 3. [Diagram 5] FIG. 2 is a diagram illustrating an example of a neural network structure of a control model. [Figure 6] FIG. 13 is a scatter plot of the number of steps required to transport luggage in a simulation. [Figure 7] FIG. 13 shows the cumulative distribution function of the average throughput of a robot in a simulation. [Figure 8] FIG. 13 is a configuration diagram of a terminal control system according to a modified example. [Figure 9] 2 is an example of a hardware configuration. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] The following embodiments of the present invention will be described with reference to the accompanying drawings. The embodiments described below are examples of the present invention, and the present invention is not limited to the following embodiments. In this specification and drawings, the same reference numerals indicate the same or corresponding parts.
[0015] Fig. 1 is a configuration diagram of a terminal control system according to this embodiment. The illustrated terminal control system includes a base station NW1 (base station network) and at least one terminal 21-2L. L is the number of terminals and is an integer equal to or greater than 1. The terminals 21-2L may also be referred to as terminal 2i. The base station NW1 and each terminal 2i are connected via wireless communication.
[0016] Each terminal 2i includes a plurality of wireless communication units 2i1-1 to 2i1-M, an acquisition unit 2i2, and a control unit 2i3, which are connected via a NW unit 2i0. M is the number of wireless communication units and is an integer equal to or greater than 1. The wireless communication units 2i1-1 to 2i1-M may also be referred to as wireless communication unit 2i1.
[0017] The wireless communication unit 2i1 wirelessly communicates with any of the wireless communication devices 101-1 to 101-N of the base station NW1. The acquisition unit 2i2 acquires terminal information including state information and communication information of the terminal 2i, and transmits the acquired terminal information to the collection unit 102 of the control device 110. The acquisition unit 2i2 includes various devices such as a camera and a sensor (radar, LIDAR, etc.). The acquisition unit 2i2 acquires, as terminal information, GPS signals acquired from these devices or the wireless communication unit 211, radio wave propagation channel information between the terminal 2i and the surrounding wireless communication device 101, arrival times of wireless signals, camera images, sensor values, tire / leg movements, and the like. The acquisition unit 2i2 may also acquire values calculated from these data as terminal information. The control unit 2i3 controls its own terminal 2i according to control information transmitted from the control device 110.
[0018] Base station NW1 includes at least one wireless communication device (base station) 101-1 to 101-N and a control device 110, which are connected via NW unit 100. N is the number of wireless communication devices and is an integer equal to or greater than 1. Wireless communication devices 101-1 to 101-N may also be referred to as wireless communication device 101.
[0019] Each wireless communication device 101 wirelessly communicates with at least one terminal 2i. Each wireless communication device 101 may have a different frequency, a different communication method, and a different bandwidth.
[0020] The control device 110 controls at least one terminal 2i that performs a task involving a physical action. The illustrated control device 110 includes a collection unit 102, a control unit 103, and an evaluation unit 104.
[0021] The collection unit 102 collects terminal information including state information and communication information of the terminal 2i via the wireless communication device 101. The collection unit 102 can collect the terminal information acquired by the acquisition unit 2i2 of each terminal 2i together with the terminal ID (identification signal) of the terminal 2i.
[0022] The terminal information is information about the terminal 2i. The state information is information about the state of the terminal 2i, and includes, for example, at least one of the terminal position, the terminal orientation, the terminal speed, the terminal operation, the terminal operation plan, and the terminal work state.
[0023] The communication information is information related to wireless communication of the terminal 2i, and includes at least one of, for example, communication quality, communication settings, traffic, communication required quality, and service grade of the terminal 2i or the owner of the terminal 2i. The communication information is, for example, a signal-to-noise power ratio, a signal-to-interference plus noise power ratio, a received signal strength indication (RSSI), a received signal reference quality (RSRQ), a packet error rate, a number of arriving bits, a number of arriving bits per unit time, a Modular Code Scheme (MCS) Index, a number of retransmissions, a delay time, and settings of an error correction technique. The RSSI (signal power) is a numerical value indicating the strength of a received signal. The communication information may also be differential information of these values, or a value calculated from these values using a predetermined formula. The communication information may also be a setting item of the wireless communication device 101, such as a frequency of the wireless communication device 101, a resource bandwidth, a transmission power, and a Quality Of Service (QoS) setting.
[0024] The collection unit 102 may collect or estimate terminal information (state information) of each terminal 2i using an acquisition device (not shown) such as a camera or a sensor connected to the control device 110.
[0025] The control unit 103 includes a control model 105 generated by reinforcement learning. The control model 105 may be one obtained in advance by repeating trial and error in a simulation space, or a control model updated by repeating trial and error while operating in an actual terminal control system, or a control model obtained in the simulation space that reflects trial and error in an actual terminal control system may be used.
[0026] The control unit 103 inputs the terminal information collected by the collection unit 102 to the control model 105, and controls the operation and communication settings of the terminal 2i based on control information related to the operation and communication settings of the terminal 2i output by the control model 105.
[0027] Specifically, the control unit 103 transmits to each terminal 2i the control information output from the control model 105. The control unit 2i3 of each terminal 2i controls its own terminal 2i in accordance with the transmitted control information.
[0028] The control unit 103 may input terminal information of one terminal 2i to the control model 105 and output control information for that terminal 2i, or may input terminal information for each of multiple terminals 2i to the control model 105 and output control information for each of the multiple terminals 2i.
[0029] The control information regarding the operation of the terminal 2i may include, for example, at least one of a long-term operation command regarding a long-term path or position, and a short-term operation command regarding short-term forward movement, backward movement, direction change, upward movement, downward movement, acceleration, or deceleration. The control information regarding the operation may also include movement to an arbitrary point, turning, etc. The control information regarding the operation may also include operation-related information including at least one of the speed, acceleration, allowable range, and operation timing of the above-mentioned operation, and operation rules such as maximum speed, minimum speed, etc.
[0030] Furthermore, the control information regarding the operation may include control information for controlling the operation of a component (moving object) included in the terminal 2i. The control information may be information for controlling the operation parameters of the terminal 2i. For example, when the terminal 2i transports an item, the operation parameters include the work efficiency, the work amount, the power consumption, the work risk, and the like.
[0031] The control information regarding the communication settings may include at least one of the wireless communication destination (wireless communication device 101), frequency, frequency band, required quality, priority, data rate, transmission frequency, and number of retransmissions. The control information regarding the communication settings may include a base station to communicate with, the number of antennas to be used, transmission power, a wireless system to be used, transmission frequency, bit rate, packet size, and transmission mode.
[0032] The evaluation unit 104 calculates an evaluation value for the control information using the terminal information after the control by the control unit 103, and updates the control model 105 so that the evaluation value of the control information output by the control model 105 becomes higher. That is, the evaluation unit 104 updates the control model 105 so as to increase the evaluation value of the control information, and causes the control model 105 to learn so as to perform better control. Specifically, the evaluation unit 104 calculates an evaluation value for the operation of the terminal 2i based on the output control information.
[0033] The evaluation unit 104 may hold a first reward value corresponding to at least one first condition related to the task and a second reward value corresponding to at least one second condition related to the communication quality, quantify the task and communication quality of the terminal 2i after the control by the control unit 103, and calculate an evaluation value using the first reward value and the second reward value. The first reward value is set depending on whether the first condition contributes to the completion of the task, and the second reward value is set depending on the communication quality of the second condition.
[0034] The first condition (required condition) is a parameter set for the work performed by the terminal 2i, and for example, in the case of a terminal 2i performing transportation, a high first reward value is generated for the work that satisfies the desirable conditions such as the success, execution, and efficiency of transportation. Conversely, a negative first reward value may be set as a penalty for the work that satisfies an undesirable condition.
[0035] For a surveillance terminal, the first condition may be monitoring, a surveillance area, a specific number of surveillance targets, etc., and a high first reward value is set for a task that satisfies the desired conditions. For a drone, the first condition may be a small deviation from a planned route, reaching a destination, etc., and a high first reward value is generated for a task that satisfies the desired conditions.
[0036] The second condition (required condition) is a parameter related to the wireless communication of the terminal 2i, and various conditions are set for the above-mentioned communication information, and a high second reward value is set for a desirable communication quality. For example, a high second reward value may be set when the state in which the number of arriving bits per unit time is equal to or greater than a specified value continues for a longer period of time, when the minimum communication quality for performing the work is large, when the number of retransmissions is small, or when the arrival delay time of the communication packet is small, and conversely, a negative second reward value may be set as a penalty when the predetermined communication quality is not satisfied. For example, a penalty or a negative reward can be set when the number of arriving bits is equal to or less than a specified value or when the state in which the number of arriving bits is equal to or less than a specified value continues for a certain period of time, when the arrival delay time of the communication packet is large, or when the fluctuation of the communication quality is large.
[0037] FIG. 2 is a flowchart showing the operation of the control device 110 of this embodiment.
[0038] The collection unit 102 collects terminal information from at least one terminal 2i (step S101). The control unit 103 inputs the terminal information of each terminal 2i to the control model 105, and outputs control information of each terminal 2i (step S102).
[0039] In reinforcement learning, a method of outputting control information that is not an optimal solution, such as the ε-greedy method, may be used to perform better actions by performing control that is not necessarily optimal. In the ε-greedy method, the control model 105 is allowed to output control information different from the control information selected by itself with a certain probability. In this way, better control may be found.
[0040] The control unit 103 controls each terminal 2i based on the control information output by the control model 105. That is, the control unit 103 transmits control information corresponding to each terminal 2i and controls the operation and communication settings of each terminal 2i (step S103).
[0041] When the terminal 2i is controlled, the evaluation unit 104 calculates an evaluation value for the control result (step S104). For example, the evaluation unit 104 evaluates the quality of the work and wireless communication performed by the terminal 2i, and obtains a preset reward value. For example, when the terminal 2i is a transport device, the following conditions and reward values are considered.
[0042] Completion of Transport: +100 ·Continued good radio communication: +10 Poor quality wireless communication: -10 In this way, a high reward value is set for the completion of a task, and a reward for wireless communication is preset to a positive value or a negative value as a penalty according to the required conditions for wireless communication quality that terminal 2i must satisfy.
[0043] The evaluation unit 104 uses the terminal information of each terminal 2i after being controlled by the control information collected by the collection unit 102 to evaluate the control result of the control information, calculates an evaluation value using a reward value, and updates the control model 105 by feeding it back to the control model 105 (step S106). The control device 110 can control the terminal 2i while improving the control model 105 by repeatedly performing the process shown in FIG. 2.
[0044] The evaluation unit 104 may calculate the sum of reward values that satisfy the above conditions (parameters) as the evaluation value. The evaluation unit 104 may also set a predetermined weighting for each condition and calculate the sum of the reward values that take the weighting into account as the evaluation value.
[0045] The evaluation unit 104 may calculate the evaluation value only when at least one of the work, communication, date and time, application, position, speed, and acceleration of the terminal 2i satisfies a predetermined condition. This is because it may be difficult to generate a control model 105 that always outputs good control information for all conditions when the range of motion of the terminal 2i is wide. Or, it may be necessary to always control in consideration of both communication quality and work, or it may be desired to perform efficient control only under specific conditions. In actual operation, various complex factors such as the environment, communication, date and time, application, interference, and weather have an effect, so it is expected that the accuracy of the control model 105 can be improved by limiting the conditions under which the control model 105 operates. For this reason, the evaluation unit 104 may digitize and output the reward only when the work, communication, date and time, application, position, speed, acceleration, etc. of the terminal 2i satisfy a predetermined condition. For example, the control method of this embodiment may be used only when a specific application, particularly an application using wireless communication, is used, or only when the application is in a specific mode.
[0046] 3 is a schematic diagram showing a communication area in a simulation (experimental example) of luggage transportation to verify the effect of the terminal control system of this embodiment. Three robots (terminals) 41, 42, 43 move while communicating on a two-dimensional plane 301 such as a warehouse, and transport luggage 305 at a loading point 303 to one of goals A to D 304 while sorting the luggage. In addition to transporting the luggage 305, the robots 41, 42, 43 must satisfy requirements for wireless communication quality with wireless base stations 51 and 52 located on the left and right sides of the warehouse 301. The throughput between the wireless base stations 51 and 52 and the robots 41, 42, 43 is defined as follows:
[0047] C i = log2(1+S i,j ) / L i Here, C iis the throughput when a wireless base station 5i (i is 1 or 2) and a robot 4j (j is any one of 1 to 3) wirelessly communicate with each other. i,j is the signal-to-interference-plus-noise power ratio in the uplink or downlink between the wireless base station 5i and the terminal 4j when a packet transmitted by the terminal 4j is received. i represents the number of terminals 4j connected to the wireless base station 5i.
[0048] Here, the calculation is performed for a scenario in which the throughput is divided according to the number of terminals 4j communicating with the wireless base station 5i, but this may be replaced with any parameters related to throughput or line quality in an actual system, such as Wi-Fi, LTE, or 5G.
[0049] FIG. 4 shows the coordinates of two-dimensional plane 301 in the simulation of FIG. 3. The assumed two-dimensional plane 301 is 13 m wide and 7 m long. Areas 41, 42, and 43 in FIG. 4 indicate the position of robot 4j. It was assumed that the loading point is located at the point (area 42) on vertical axis 5 and horizontal axis 6 in FIG. 4. Area 304 in FIG. 4 indicates Goals A, B, C, and D from left to right. Areas 51 and 52 in FIG. 4 indicate wireless base station 5j. The range in which the robot can move is the white area in FIG. 4, Goal area 304, and Loading Point area 42.
[0050] Here, the Goal is specified by the luggage. Each robot sorts the luggage with the specified Goal into the specified Goal. At Loading Point 42, a buffer containing 20 luggage is installed. When a robot reaches Loading Point 42 without carrying any luggage, it picks up the luggage stored in the buffer, and when the robot reaches the specified Goal while carrying any luggage, it unloads the luggage and repeats this process.
[0051] In this simulation, a deep reinforcement learning algorithm is used to estimate appropriate robot movement control. Appropriate robot movement control is learned by repeating three steps: estimating appropriate robot movement control using the robot's own position and the robot's destination as input, reflecting the estimated robot operation in the simulator environment, and calculating an evaluation value (reward / penalty) based on the robot's task execution efficiency and network performance after the operation.
[0052] Fig. 5 shows the neural network structure of the control model 105 used in the simulations of Fig. 3 and Fig. 4. In this method, the vertical and horizontal coordinates of the self-position of each robot and the horizontal and vertical coordinates of the destination are given as inputs to estimate the behavior of each robot. That is, a total of four pieces of information, the vertical and horizontal coordinates of the self-position and the vertical and horizontal coordinates of the destination, are input. The selection of the wireless base station 5j imposes a condition for connection to the one with the higher received power.
[0053] In the illustrated example, features related to the robot's own position and destination are extracted from the input robot's own position and destination using a fully connected layer (FC) 501 and a rectified linear unit (ReLU) 502 twice. FC is a fully connected layer. ReLU 502 is an activation function f(x)=max(0,x) that sets negative values to 0. The subscript numbers on the right side of the figure indicate the number of neurons.
[0054] The one-dimensional tensor output by FC501 is split in half and each half is used to learn the state value function and action value function. In Figure 5, the output of FC501 with the subscript "64" is split into two parts of 32 each, and the upper FC501a learns the state value function, while the lower FC501b learns the action value function. Finally, an appropriate robot action is estimated and output from the state-action value function that combines the output obtained from the state value function of FC501a and the output of the action value function of FC501b. Five types of actions are output: moving 1m in any of the four directions (up, down, right, or left), or stopping.
[0055] In order to learn robot movement control that takes network performance into account, the conditions and reward values for calculating the evaluation value are set as follows:
[0056] 1 Action: -1 -Throughput measured below 0.2Gbps: -3 Attempting to enter an inaccessible area: -5 ·Stationary: -5 Approaching destination: +2 Completion of Transport: +100 The deduction of reward points for "one action" is intended to complete a task in as few steps as possible. As a result of selecting an action, it is determined whether any of the following is met: "Is the throughput below 0.2Gbps?", "Has an attempt been made to enter an intrusion load area?", "Has the choice been made to stop?", "Has the distance to the destination become smaller?", or "Has the transport been completed?". If any of these are met, the robot receives the corresponding reward or penalty. If the throughput falls below 0.2Gbps and the robot approaches the destination, a penalty of -3 + 2 = -1 is imposed. Note that if an attempt is made to enter an intrusion-prohibited area, the robot's position is not updated. If the robot is not transporting any luggage and there is no luggage in the entrance buffer, the reward is set to 0 regardless of which action is taken.
[0057] The effects of this simulation are shown in Figures 6 and 7. To show the effects of this simulation, we evaluated task efficiency and communication performance in two cases: one where throughput is not considered and the above evaluation item "Throughput measured below 0.2 Gbps: -3" is not included, and the other where throughput is considered and throughput is included and "Throughput measured below 0.2 Gbps: -3" is included.
[0058] Figure 6 shows a scatter plot of the number of steps required to transport 20 pieces of luggage for each number of episodes. From Figure 6, we can see that the system eventually converges to 100 steps both when throughput is taken into account and when it is not. From the perspective of task efficiency, we can see that the system operates without any loss of efficiency by taking throughput into account. In other words, when throughput is not taken into account, the system converges to 100 steps in approximately 2,300 episodes, but when throughput is taken into account, the system converges to 100 steps in approximately 1,200 episodes.
[0059] Figure 7 shows the cumulative distribution function (CDF) of the average throughput of the three robots for each episode. The CDF makes it possible to confirm the overall distribution of throughput, with 0 to 100% corresponding to 0 to 1 on the vertical axis. When throughput is not taken into account, it can be seen that 0.31 Gbps is frequently measured. When throughput is taken into account, it can be seen that 0.325 Gbps is frequently measured.
[0060] As a result, it can be said that the terminal control system of this embodiment achieves the improvement of communication quality while maintaining task efficiency by controlling the robot while taking into account the throughput. This is an effect obtained by using reinforcement learning that takes into account both communication quality and work efficiency, as opposed to conventional methods.
[0061] The control device 110 of the present embodiment described above is a control device that controls a terminal 2i that performs work involving physical operations, and the terminal 2i communicates wirelessly with the wireless communication device 101. The control device 110 includes a collection unit 102 that collects terminal information including status information and communication information of the terminal 2i, a control unit 103 that inputs the terminal information to a control model 105 and controls the operation and communication settings of the terminal 2i based on control information regarding the operation and communication settings of the terminal 2i output by the control model 105, and an evaluation unit 104 that calculates an evaluation value of the control information using the terminal information after control by the control unit 103 and updates the control model 105 so that the evaluation value of the control information output by the control model 105 is increased.
[0062] This makes it possible to improve the quality of wireless communication of the terminal 2i while maintaining the task efficiency of the work performed by the terminal 2i.
[0063] <Modification> Next, a modified example of this embodiment will be described. In this embodiment shown in Fig. 1, the base station NW1 includes a control device 110, which collects terminal information from each terminal 2i, controls each terminal 2i according to the control information, and evaluates the control results. In this modified example, each terminal 2i is provided with the function of the control device 110.
[0064] 8 is a configuration diagram of a terminal control system according to a modified example. The terminal control system according to the modified example includes a base station NW1A (base station network) and at least one terminal 21A to 2LA. L is the number of terminals and is an integer equal to or greater than 1. The terminals 21A to 2LA may also be referred to as terminal 2iA. The base station NW1 and each terminal 2iA are connected by wireless communication.
[0065] 1 in that it does not include a control device 110, but is otherwise similar to the base station NW1 in Fig. 1. That is, the base station NW1A includes at least one of wireless communication devices 101-1 to 101-N, which are connected via a NW unit 100. The wireless communication devices 101-1 to 101-N are similar to the wireless communication device 101 in Fig. 1.
[0066] Terminal 2iA is a wireless communication terminal that performs work involving physical actions, and wirelessly communicates with wireless communication device 101 of base station NW1A. Each terminal 2iA includes a plurality of wireless communication units 2i1-1 to 2i1-M, an acquisition unit 2i2, a control unit 2i3, and an evaluation unit 2i4, which are connected via NW unit 2i0. M is the number of wireless communication units and is an integer equal to or greater than 1. Wireless communication units 2i1-1 to 2i1-M may be referred to as wireless communication unit 2i1. Wireless communication unit 2i1 in the modified example is similar to wireless communication unit 2i1 in FIG. 1.
[0067] The acquisition unit 2i2 acquires terminal information including the status information and communication information of the terminal 2iA, and transmits the acquired terminal information to 2i3. The acquisition unit 2i2 differs from the acquisition unit 2i2 in Fig. 1 in that the acquisition unit 2i2 transmits the acquired terminal information to the control unit 2i3, but is otherwise similar to the acquisition unit 2i2 in Fig. 1.
[0068] The control unit 2i3 has the same function as the control unit 103 in Fig. 1. That is, the control unit 2i3 includes a control model 2i5 generated by reinforcement learning. The control unit 2i3 inputs terminal information to the control model 2i5, and controls the operation and communication settings of the terminal 2iA based on control information related to the operation and communication settings of the terminal 2iA output by the control model 2i5.
[0069] The evaluation unit 2i4 has the same function as the evaluation unit 104 in Fig. 1. That is, the evaluation unit 2i4 calculates an evaluation value of the control information by using the terminal information after being controlled based on the control information, and updates the control model 2i5 so that the evaluation value of the control information output by the control model 2i5 becomes higher.
[0070] In this manner, in this modification, each terminal 2iA inputs its own terminal information to the control model 2i5, and autonomously controls and evaluates its own terminal 2iA according to the output control information. As a result, in this modification, as in the above embodiment, it is possible to improve the quality of wireless communication of the terminal while maintaining the task efficiency of the work by the terminal.
[0071] <Hardware> The control device and the terminals 2i, 2iA described above may be realized by a computer. That is, for example, a general-purpose computer system as shown in FIG. 9 can be used for the control device 110 and the terminals 2i, 2iA. The illustrated computer system includes a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906. The memory 902 and the storage 903 are storage devices. In this computer system, the CPU 901 executes the program of the control device 110 or the terminals 2i, 2iA loaded on the memory 902, thereby realizing each function of the control device 110 or the terminals 2i, 2iA.
[0072] The control device 110 and the terminals 2i, 2iA may be implemented in one computer, or in multiple computers. The control device 110 and the terminals 2i, 2iA may be virtual machines implemented in a computer. The programs for the control device 110 and the terminals 2i, 2iA may be stored in a computer-readable recording medium such as an HDD, SSD, a Universal Serial Bus (USB) memory, a Compact Disc (CD), or a Digital Versatile Disc (DVD), or may be distributed via a network.
[0073] The computer-readable recording medium may include a recording medium that dynamically holds a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, or a recording medium that holds a program for a certain period of time, such as a volatile memory in a computer system that serves as a server or client in that case. The program may be for implementing some of the aforementioned components. The aforementioned components may be realized in combination with a program already recorded in the computer system, or may be realized using hardware such as a PLD (Programmable Logic Device) or an FPGA (Field Programmable Gate Array).
[0074] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the present invention. [Explanation of symbols]
[0075] 1: Base station network 100: Network section 101: Wireless communication device (base station) 102: Collection Department 103: Control unit 104: Evaluation section 105: Control model 110: Control device 21: Terminal 210: Network Department 211: Wireless communication department 212: Acquisition Department 213: Control unit 214: Evaluation section 215: Control model
Claims
1. A control device for controlling a terminal that performs a task involving a physical action, the terminal wirelessly communicates with a wireless communication device, The control device includes: A collection unit that collects terminal information including status information and communication information of the terminal; a control unit that inputs the terminal information into a control model and controls the operation and communication settings of the terminal based on control information related to the operation and communication settings of the terminal output from the control model; an evaluation unit that calculates an evaluation value of the control information by using terminal information after the control by the control unit, and updates the control model so that the evaluation value of the control information output by the control model is increased. Control device.
2. the evaluation unit holds a first reward value corresponding to at least one first condition related to an operation and a second reward value corresponding to at least one second condition related to communication quality, quantifies the operation and the communication quality of the terminal after control by the control unit, and calculates the evaluation value using the first reward value and the second reward value; The first reward value is set depending on whether the first condition contributes to the completion of the task, and the second reward value is set depending on the communication quality of the second condition. The control device according to claim 1 .
3. The evaluation unit calculates the evaluation value only when at least one of the operation, communication, date and time, application, position, speed, and acceleration of the terminal satisfies a predetermined condition. The control device according to claim 1 or 2.
4. The control information for motion includes at least one of long-term motion commands for a long-term path or position, and short-term motion commands for moving forward, backward, turning, ascending, descending, accelerating, or decelerating; The control information regarding communication settings includes at least one of a wireless communication destination, a frequency, a frequency band, a required quality, a priority, a data rate, a transmission frequency, and a number of retransmissions. The control device according to any one of claims 1 to 3.
5. A wireless communication terminal for performing a task involving a physical action, an acquisition unit that acquires terminal information including status information and communication information of the wireless communication terminal; a control unit that inputs the terminal information into a control model and controls an operation and communication setting of the wireless communication terminal based on control information related to the operation and communication setting of the wireless communication terminal output from the control model; an evaluation unit that calculates an evaluation value of the control information by using terminal information after being controlled based on the control information, and updates the control model so that the evaluation value of the control information output by the control model is increased. Wireless communication terminal.
6. A control method for controlling a terminal that performs a task involving a physical operation, the method comprising: The control device includes: a collection step of collecting terminal information including status information and communication information of the terminal; a control step of inputting the terminal information into a control model and controlling an operation and communication setting of the terminal based on control information related to the operation and communication setting of the terminal outputted from the control model; an evaluation step of calculating an evaluation value of the control information using terminal information after the control step, and updating the control model so that the evaluation value of the control information output by the control model is increased. Control methods.
7. A control program that causes a computer to function as the control device according to any one of claims 1 to 4.
8. A control program for causing a computer to function as the wireless communication terminal according to claim 5.
Citation Information
Patent Citations
Testing system and testing method for onboard electrical components
JP2005181113A
Testing system for on-vehicle electrical component, and and testing method
JP2006329787A
Learning control system and learning control method
JP2019046422A
Systems and methods for operator skill reduction
JP2021517534A
Optimization method of wireless communication system, wireless communication system, and program for wireless communication system
JP2022018901A