Control device, control method, and control system

The control device uses learning models to optimize the path of IoT devices to power supply units, addressing inefficient power supply by determining optimal routes based on past experiences and environmental data.

JP2026089855AActive Publication Date: 2026-06-02INTERNET INITIATIVE JAPAN INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERNET INITIATIVE JAPAN INC
Filing Date
2024-11-21
Publication Date
2026-06-02

Smart Images

  • Figure 2026089855000001_ABST
    Figure 2026089855000001_ABST
Patent Text Reader

Abstract

The aim is to find the optimal travel route with a simpler configuration and support efficient power supply to IoT devices. [Solution] The control device 1 includes a first learning unit 13 that learns a strategy for the path the first device should take sequentially from its current position using a reinforcement learning model, and a second learning unit 14 that learns the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach its destination point, obtained through learning by the first learning unit 13, using a supervised learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a control device, a control method, and a control system. [Background technology]

[0002] Conventionally, communication standards such as LPWA (Low Power Wide Area, registered trademark) and eDRX (extended Discontinuous Reception, registered trademark) have been used to reduce power consumption in IoT terminals. While these communication standards allow IoT terminals to use power for longer periods, if the IoT terminal's power runs out, it becomes inoperable. Patent Document 1 discloses an autonomous mobile IoT terminal that decides whether to move to a charging facility based on the remaining operating time based on the remaining power and the time it will take to reach the charging facility.

[0003] However, in Patent Document 1, the route used to move the IoT terminal to the charging equipment is a pre-set, predetermined route. Therefore, the route from the IoT terminal's location when its remaining battery level falls below a threshold to the charging equipment is not necessarily the optimal route. Consequently, it was sometimes difficult to support efficient power supply with a simpler configuration when the IoT terminal's remaining battery level was limited. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2023-111344 [Overview of the project] [Problems that the invention aims to solve]

[0005] With conventional technology, it was difficult to efficiently supply power to IoT devices by finding the optimal travel route with a simpler configuration when an anomaly occurred regarding the remaining power of the IoT device.

[0006] This invention was made to solve the above-mentioned problems and aims to support efficient power supply to IoT terminals by finding the optimal travel route with a simpler configuration. [Means for solving the problem]

[0007] To solve the above-mentioned problems, the control device according to the present invention is a control device that controls the path of a first device moving to a destination point set in a moving space, comprising: a first acquisition unit configured to acquire abnormality information regarding the remaining power of the first device or a second device arranged in the moving space; a second acquisition unit configured to acquire the position of the second device as the position of the destination point; a third acquisition unit configured to acquire the current position of the first device attached to a position registration request signal transmitted by the first device; and a control device that calculates the path that the first device should sequentially take from the initial position of the first device to the position of the destination point. The system comprises: a first learning unit configured to apply a reward function to the calculated estimation result to update it so that the reward for the first device to reach the destination point is maximized, and to learn a strategy for the path the first device should take sequentially from its current position using a reinforcement learning model; a second learning unit configured to learn, using a supervised learning model, the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach the destination point, obtained through learning by the first learning unit; and a storage unit configured to store the learned supervised learning model constructed by the second learning unit.

[0008] Furthermore, in the control device according to the present invention, the second acquisition unit may acquire the position of the second device, which is added to the position registration request signal transmitted by the second device, as the position of the destination point.

[0009] Further, in the control device according to the present invention, when the remaining power amount of the first device becomes less than a threshold value, the first device adds the abnormality occurrence information to the position registration request signal and transmits it. The second device is a power supply device that supplies power to the first device, and the third acquisition unit may acquire the current position of the first device契机 with the acquisition of the abnormality occurrence information by the first acquisition unit.

[0010] Further, in the control device according to the present invention, the first device is a power supply device that supplies power to the second device. When the remaining power amount of the second device becomes less than a threshold value, the second device adds the abnormality occurrence information to the position registration request signal and transmits it. The third acquisition unit may acquire the current position of the first device契机 with the acquisition of the abnormality occurrence information by the first acquisition unit.

[0011] Further, in the control device according to the present invention, it may further include a setting unit configured to set the learned supervised learning model in the first device as control information for controlling a path from the position of the initial point to the position of the destination point by the first device.

[0012] To solve the above-mentioned problems, the control method according to the present invention is a control method for controlling the path of a first device moving to the position of a destination point set in a moving space, comprising: a first acquisition step of acquiring abnormality occurrence information regarding the remaining power of the first device or a second device arranged in the moving space; a second acquisition step of acquiring the position of the second device as the position of the destination point; a third acquisition step of acquiring the current position of the first device attached to a position registration request signal transmitted by the first device; and calculating the path that the first device should sequentially take from the initial position of the first device to the position of the destination point. The system includes: a first learning step in which a reward function is applied to the estimation result to update it so that the reward for the first device to reach the destination point is maximized, and a strategy for the path the first device should take sequentially from its current position is learned using a reinforcement learning model; a second learning step in which the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach the destination point, obtained through learning in the first learning step, is learned using a supervised learning model; and a storage step in which the learned supervised learning model constructed in the second learning step is stored in a memory unit.

[0013] Furthermore, in the control method according to the present invention, the second acquisition step may acquire the position of the second device, which is attached to the position registration request signal transmitted by the second device, as the position of the destination point.

[0014] Furthermore, in the control method according to the present invention, the first device transmits the abnormality information attached to the position registration request signal when the remaining power of the device falls below a threshold, the second device is a power supply device that supplies power to the first device, and the third acquisition step may acquire the current position of the first device triggered by the acquisition of the abnormality information in the first acquisition step.

[0015] Furthermore, in the control method according to the present invention, the first device is a power supply device that supplies power to the second device, and the second device transmits the abnormality information attached to the position registration request signal when the remaining power of its own device falls below a threshold, and the third acquisition step may acquire the current position of the first device triggered by the acquisition of the abnormality information in the first acquisition step.

[0016] Furthermore, the control method according to the present invention may further include a setting step in which the learned supervised learning model is set in the first device as control information that controls the path the first device takes from the initial location to the destination location.

[0017] Furthermore, in the control method according to the present invention, the first device further comprises: a fourth acquisition step of acquiring the trained supervised learning model constructed in the second learning step; a fifth acquisition step of acquiring the current position of the device itself; a calculation step of providing the current position of the device itself acquired in the third acquisition step as an unknown input to the trained supervised learning model, performing calculations on the trained supervised learning model, and outputting a strategy for the path the device should sequentially take from its current position; and a movement control step of controlling the movement of the device itself from the initial point to the destination point based on the strategy for the path the device should sequentially take from its current position output in the calculation step.

[0018] To solve the above-mentioned problems, the control system according to the present invention is a control system comprising the control device described above and the first device, wherein the first device comprises a fourth acquisition unit configured to acquire the learned supervised learning model constructed by the control device, a fifth acquisition unit configured to acquire the current position of the device itself, a calculation unit configured to provide the current position of the device itself acquired by the third acquisition unit as an unknown input to the learned supervised learning model, perform calculations on the learned supervised learning model, and output a strategy for the path the device should sequentially take from its current position, and a movement control unit configured to control the movement of the device itself from the initial point to the destination point based on the strategy for the path the device should sequentially take from its current position output by the calculation unit. [Effects of the Invention]

[0019] According to the present invention, the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach its destination, obtained through learning by the first learning unit, is learned using a supervised learning model. Therefore, it is possible to find the optimal travel route with a simpler configuration and support efficient power supply to IoT terminals. [Brief explanation of the drawing]

[0020] [Figure 1] Figure 1 is a block diagram showing the configuration of a control system equipped with a control device according to an embodiment of the present invention. [Figure 2] Figure 2 is a diagram illustrating the overview of the control system according to this embodiment. [Figure 3] Figure 3 is a diagram illustrating the learning process performed by the first learning unit of the control device according to this embodiment. [Figure 4] Figure 4 is a block diagram showing the configuration of the first learning unit included in the control device according to this embodiment. [Figure 5] Figure 5 is a diagram illustrating the learning process performed by the second learning unit of the control device according to this embodiment. [Figure 6] Figure 6 is a block diagram showing the hardware configuration of the control device according to this embodiment. [Figure 7] Figure 7 is a block diagram showing the configuration of the mobile terminal device included in the control system according to this embodiment. [Figure 8] Figure 8 is a block diagram showing the hardware configuration of the mobile terminal device included in the control system according to this embodiment. [Figure 9] Figure 9 is an operation sequence diagram of the control system according to this embodiment. [Figure 10] Figure 10 is a flowchart showing the first learning process of the control device according to this embodiment. [Figure 11] Figure 11 is a flowchart showing the first learning process of the control device according to this embodiment. [Figure 12] Figure 12 is a flowchart showing the operation of the mobile terminal device included in the control system according to this embodiment. [Figure 13] Figure 13 is a diagram illustrating an overview of a control system according to a modified example of this embodiment. [Figure 14] Figure 14 is an operation sequence diagram of a control system according to a modified example of this embodiment. [Modes for carrying out the invention]

[0021] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to Figures 1 to 14.

[0022] [Control system configuration] First, with reference to Figure 1, an overview of a control system comprising a control device 1 and a mobile terminal device (first device) 2 according to an embodiment of the present invention will be described.

[0023] The control system according to this embodiment comprises a control device 1, a mobile terminal device 2, a power supply device (second device) 2B, a base station 3, and a core network 4. The control device 1, the mobile terminal device 2, and the power supply device 2B are connected to each other so as to be able to communicate via a wireless communication network NW that conforms to a predetermined communication standard such as LTE / 4G, 5G, or 6G. The control system controls the path that the mobile terminal device 2 takes to travel to a set destination point in the mobile space. As shown in Figures 1 and 2, the mobile space A on which the mobile terminal device 2 travels is, for example, capable of communication using a 5G wireless communication method.

[0024] Mobile terminal device 2 includes flying objects capable of autonomous flight such as mobile robots, drones, and unmanned aerial vehicles, as well as autonomous vehicles and ships. When the remaining charge of the mobile terminal device 211 falls below a threshold, it determines the optimal route to the power supply device 2B and autonomously moves to the location of the power supply device 2B, which is the destination point. As shown in Figure 2, the mobile terminal device 2 is moving along an arbitrary route in mobile space A, and the position at which the remaining power falls below the threshold is considered the initial point. Then, the mobile terminal device 2 moves from the initial point to the location of the power supply device 2B, which is the destination point, and receives power from the power supply device 2B. The mobile terminal device 2 and the power supply device 2B can communicate with each other via short-range wireless communication.

[0025] In the following explanation, we will use the case where the mobile terminal device 2 is a mobile robot as an example. The mobile terminal device 2 controls autonomous movement using a controller that processes information from sensors 208 and other devices (described later) to control the rotation speed of the motor 209 and the drive mechanism 210. The mobile terminal device 2 also acquires its own GPS position using a GPS receiver 207. The mobile terminal device 2 is configured as an IoT terminal with an IP address, and each IP address can uniquely identify the mobile terminal device 2.

[0026] Furthermore, the mobile terminal device 2 according to this embodiment is equipped with a SIM and has an IMSI (International Mobile Subscriber Identity) stored in the SIM. Details of the functional blocks and hardware configuration of the mobile terminal device 2 will be described later. Note that there may be multiple mobile terminal devices 2, in which case each mobile terminal device 2 will have the same configuration.

[0027] The power supply unit 2B is fixedly positioned in the mobile space. The power supply unit 2B is implemented by a computer equipped with a processor, main memory, communication interface, auxiliary storage, and input / output I / O, and a program that controls these hardware resources. The power supply unit 2B further includes a charging device and a charging control unit for supplying power to the mobile terminal device 2. More specifically, the power supply unit 2B includes a power management unit, a charging module, a DC-DC converter, and contact-type charging terminals. As an example, the power supply unit 2B is implemented as an integral part of a utility pole positioned in the mobile space. Multiple power supply units 2B may be positioned in the mobile space. For example, they may be arranged in a grid pattern at predetermined intervals within mobile space A in Figure 2.

[0028] The power supply device 2B can uniquely identify the mobile terminal device 2 by its IP address. Furthermore, the power supply device 2B according to this embodiment is equipped with a SIM card and can also be uniquely identified by the IMSI stored in the SIM card. The power supply device 2B is also equipped with a GPS receiver and acquires the GPS location of its own device.

[0029] The mobile terminal device 2 and the power supply device 2B transmit a location registration request signal to the core network 4 via the base station 3 at regular intervals. The mobile terminal device 2 and the power supply device 2B transmit the location registration request signal with their own GPS location added to it. The mobile terminal device 2 can transmit the location registration request signal with its own location added when crossing the area of ​​base station 3 or when power is turned on. Similarly, the power supply device 2B can transmit the location registration request signal with its own location added when power is turned on. The location registration request signal includes the IMSI of the transmitting device.

[0030] The mobile terminal device 2 also transmits information regarding an abnormality in the remaining power level, added to the location registration request signal, when the remaining power of its own battery 211 falls below a threshold.

[0031] As shown in Figure 2, the mobile space A in which the mobile terminal device 2 moves is a three-dimensional matrix-like space composed of unit spaces divided into multiple spaces. Each unit space constituting the mobile space A has the same volume. Furthermore, each unit space has a node ID, and each unit space is represented by a single position (x, y, z). The position information can be three-dimensional GPS position coordinates consisting of latitude, longitude, and altitude. For example, a representative value such as the center position of the unit space can be used as the position of that unit space.

[0032] Furthermore, as shown in Figure 2, the mobile terminal device 2 moves from a unit space position corresponding to the initial point where the battery 211's remaining charge falls below a threshold, using each unit space as a waypoint, to the unit space position of the destination point where the power supply device 2B is located.

[0033] Base station 3 is a wireless base station compatible with the 5G communication standard and relays communication between the mobile terminal device 2 and power supply device 2B located within the communication area and the core network 4. Base station 3 is connected to the core network 4 via a backhaul link.

[0034] Core network 4 is connected to control device 1 via a network NW such as LAN, WAN, or the Internet. Core network 4 includes AMF (Access and Mobility Management Function) 40 and UDM (Unified Data Management) / UDR (Unified Data Repository) 41, which are nodes within the C-plane. Core network 4 also includes multiple UPF (User Plane Function) 42 within the U-plane. Functional nodes within the U-plane and C-plane that are included in core network 4 other than those mentioned above are not shown in the diagram.

[0035] The AMF40 is a node that provides mobility control functions and performs movement control such as location registration, paging, and handover. Based on location registration request signals from the mobile terminal device 2 and the power supply device 2B, the AMF40 performs authentication, service connection initiation, session management initiation, etc. Location information and abnormality information from the mobile terminal device 2 and the power supply device 2B, which are attached to the location registration request signals received by the AMF40, are transmitted to the UDM / UDR41.

[0036] The UDM / UDR41 manages subscriber profiles, performs authentication, and manages mobility. In this embodiment, the UDM / UDR41 adds fields for location information and anomaly occurrence information to the subscriber profile. The UDM / UDR41 according to this embodiment is equipped with a communication interface 41a for communicating with the control device 1. In the UDM / UDR41, the presence or absence of location information and anomaly occurrence information is stored using the IMSI included in the location registration request signal as a key. The transmission timestamp of the location registration request signal identifies when the mobile terminal device 2 and power supply device 2B, identified by the IMSI, were located at the GPS location. In addition, the UDM / UDR41 may be configured as a single device with the UDM and UDR, or it may be a device in which the UDM and UDR are arranged separately.

[0037] UPF42 is a user plane function that handles data between base station 3 and data networks (DNs) such as the internet.

[0038] The control system according to this embodiment learns the optimal path from the initial location where the battery 211 of the mobile terminal device 2 has a remaining charge below a threshold, to the destination location where the power supply device 2B is installed, using reinforcement learning. Furthermore, using the policy for the mobile terminal device 2's path obtained through reinforcement learning as training data, the system learns the relationship between the current position of the mobile terminal device 2 in a unit space and the policy for the path the mobile terminal device 2 should sequentially take, using a supervised learning model. In addition, the learned supervised learning model is set in the mobile terminal device 2 as control information to control the path of the mobile terminal device 2. Based on the set control information, the mobile terminal device 2 takes its current position in a unit space as an unknown input, performs calculations using the learned supervised learning model, and outputs the policy for the path it should sequentially take. Then, based on the output policy for the path to sequentially take, the system determines the path from the initial location to the destination location and controls its movement to the power supply device 2B at the destination location.

[0039] The mobile terminal device 2, with control information set, changes its course in one of the directions indicated by each arrow, as shown in Figure 2, according to the course determined based on the course policy obtained from calculations of a trained supervised learning model, and moves in the direction it should move in each unit space. The course can include various courses, i.e., directions of movement. In Figure 2, the movement space A is explained in a two-dimensional plane, but the course of the mobile terminal device 2 can be a three-dimensional course. With the control information set by the control device 1 to the mobile terminal device 2, the mobile terminal device 2 can reach the unit space where the power supply device 2B at the destination is located, from the unit space of the initial point.

[0040] Here, "path" refers to the direction of movement from one position in a unit space to an adjacent position in another unit space. "Route" includes the entire route from the initial point to the destination point.

[0041] [Controller Function Blocks] As shown in Figure 1, the control device 1 comprises a first acquisition unit 10, a second acquisition unit 11, a third acquisition unit 12, a first learning unit 13, a second learning unit 14, a first storage unit (storage unit) 15, a second storage unit 16, a third storage unit 17, and a setting unit 18. The control device 1 controls the path of the mobile terminal device 2 to the location of a set destination point in the moving space.

[0042] The first acquisition unit 10 acquires abnormality information regarding the remaining power of the mobile terminal device 2 (first device). In this embodiment, the first acquisition unit 10 acquires abnormality information attached to the location registration request signal transmitted by the mobile terminal device 2 from the UDM / UDR 41 of the core network 4 via the network NW. The abnormality information is transmitted by the mobile terminal device 2 when the remaining battery level of the battery 211 falls below a threshold.

[0043] The second acquisition unit 11 acquires the position of the power supply device 2B (second device) as the position of the destination. The second acquisition unit 11 acquires the position information attached to the position registration request signal transmitted by the power supply device 2B from the UDM / UDR 41 of the core network 4 via the network NW. As described above, the power supply device 2B adds its own GPS position to the position registration request signal and transmits it to the core network 4 when installed and at regular intervals. The second acquisition unit 11 also acquires the node ID of the unit space corresponding to the GPS position, which is stored in the third storage unit 17, described later. If there are multiple power supply devices 2B in the mobile space, the power supply device 2B with the shortest straight-line distance can be selected and set as the destination based on the GPS position of the mobile terminal device 2 that is the source of the position registration request signal with abnormality information attached, acquired by the first acquisition unit 10.

[0044] The third acquisition unit 12 acquires the current location of the mobile terminal device 2 (first device) attached to the location registration request signal transmitted by the mobile terminal device 2 (first device). The third acquisition unit 12 acquires the current location of the mobile terminal device 2 triggered by the acquisition of abnormality information by the first acquisition unit 10. The third acquisition unit 12 acquires location information attached to the location registration request signal transmitted by the mobile terminal device 2 from the UDM / UDR 41 of the core network 4 via the network NW. As described above, the mobile terminal device 2 adds its own GPS location to the location registration request signal and transmits it to the core network 4 at regular intervals. The third acquisition unit 12 acquires the node ID of the unit space corresponding to the GPS location from the third storage unit 17. The third acquisition unit 12 acquires the unit space in which the mobile terminal device 2 is currently located as the current location, and the current location is the location of the unit space in which the mobile terminal device 2 exists at time t.

[0045] The first learning unit 13 applies a reward function to the estimated path that the mobile terminal device 2 (first device) should sequentially take from its initial position to its destination position, and updates the result to maximize the reward for the mobile terminal device 2 to reach its destination position. The unit then learns a strategy for the path that the mobile terminal device 2 should sequentially take from its current position using a reinforcement learning model.

[0046] In this embodiment, the mobile terminal device 2 has a strategy for the path it should sequentially take from each unit space position, which involves actions a related to movement in a predetermined n (where n is an integer of 2 or more) directions relative to the direction of travel. n An example of the case where this is adopted is given. Also, the direction of travel is the direction based on the position in the unit space where the mobile terminal device 2 was immediately before.

[0047] The first learning unit 13 uses a neural network model as a reinforcement learning model, which includes an input layer s, a hidden layer h, and an output layer q as shown in Figure 3. Furthermore, the neural network model uses the state s, which represents the position of the mobile terminal device 2. t It receives all action value functions Q(s t a1), Q(s t a2), Q(s t, a3), ···, Q(s t , a n-1 ), Q(s t , a n ) outputs the Deep Q-Network (DQN), which is a neural network.

[0048] More specifically, the first learning unit 13 gives the position of the current unit space indicating the position of the current mobile terminal device 2 as an input to the neural network model, performs the operation of the neural network model, and as the route that the mobile terminal device 2 should proceed to next from the current position of the unit space, the action a related to each movement in n directions n outputs the first estimated value Q1 of the action value function representing the expected value of the cumulative value of the future reward obtained when taking.

[0049] The reward is given by the reward function r = r(s, a, s') of the state s indicating the current position of the mobile terminal device 2, the action a of the mobile terminal device 2 moving in a predetermined direction n , and the next position of the mobile terminal device 2, that is, the next state s'. In the present embodiment, the reward function includes, as a variable, the degree of reach to the position of the unit space of the power supply device 2B related to the destination point of the mobile terminal device 2. In addition, it can include, as a variable, the degree of reach to the position of the unit space corresponding to the space with obstacles. For example, when the mobile terminal device 2 approaches the destination point or reaches the power supply device 2B at the shortest distance by the action related to the movement in a predetermined direction, the reward, which is a scalar quantity, is set as a larger value.

[0050] On the other hand, when the mobile terminal device 2 moves away from the power supply device 2B, which is the destination point, or reaches the unit space where there are obstacles, it can be designed to give a negative reward value (for example, r = -1). Thus, by setting the reward of the unit space where there are obstacles as a negative value, the mobile terminal device 2 can avoid these points and reach the position of the power supply device 2B.

[0051] Furthermore, the first learning unit 13 receives the next unit space location reached by the mobile terminal device 2 as input to the neural network model, performs calculations on the neural network model, and outputs a second estimate Q2 of the action-value function. The first learning unit 13 learns the weight parameters of the neural network model so that the first estimate Q1 becomes the target value calculated from the second estimate Q2.

[0052] If we denote the weight parameters of the neural network model as θ and the action-value function as Q(s,a;θ), the learning minimization loss function is given by the following equation (1). L(θ) = 1 / 2{r + γmax} a’ Q(s',a';θ)-Q(s,a;θ)} 2 ...(1)

[0053] In equation (1) above, r is the reward (immediate reward) and γ is the discount rate. Q(s,a;θ) corresponds to the first estimate Q1, and Q(s',a';θ) corresponds to the value of the action in state s' after one step, i.e., the second estimate Q2. The target value is r + γmax a’ It can be represented by Q(s',a';θ).

[0054] The first learning unit 13 can update the weight parameters of the neural network model by backpropagating the gradient of the loss function given by equation (1) above.

[0055] More specifically, the first learning unit 13 can employ a Fixed Target Q-Network using two neural networks, main QN131 and target QN133, as shown in Figure 4. Main QN131 selects the optimal action and updates the action-value function Q. Meanwhile, target QN133 estimates and evaluates the value of the action a' to be taken in the next state s' resulting from the action. Main QN131 and target QN133 have neural networks with the same layer structure, but the parameter of main QN131 is "θ" and the parameter of target QN133 is "θ". - It is given by ".

[0056] The main QN131 receives the current position of the mobile terminal device 2 as state s from environment 130. Environment 130 is a system of the mobile space in which the mobile terminal device 2 is located. Under this environment 130, the mobile terminal device 2 moves to another unit space by taking action a related to movement in a predetermined direction, transitions to the next state s', and simultaneously receives a reward r from environment 130.

[0057] The first learning unit 13 inputs the state s related to the current position of the mobile terminal device 2 to the main QN131 and calculates the action-value function Q(s,a;θ). The first learning unit 13 calculates action a using, for example, the ε-greedy method, or the optimal action argmax at the present time. a We find Q(s,a;θ). In environment 130, the mobile terminal device 2 takes action related to the optimal path at the present time. a Perform Q(s,a;θ). Environment 130 is where mobile terminal device 2 performs action argmax. a The result of performing Q(s,a;θ) is observed as the next state s' in the unit space where the user moved, and the reward r is output. Experience data 134 stores the experience (s,a,r,s') output from environment 130.

[0058] The first learning unit 13 calculates the loss function L in the DQN loss calculation 132 and updates the weights of the main QN 131 using the gradient of the loss function L.

[0059] The first learning unit 13 periodically copies the weights of the main QN131 to the target QN133 and synchronizes them. The synchronization of the target QN133 is performed at a lower frequency than the update frequency of the weights of the main QN131. The first learning unit 13 extracts experience from the experience data 134, inputs the past state into the target QN133, and estimates the max value. a’ Q(s',a';θ - The first learning unit 13 outputs the estimated value max output by target QN133. a’ Q(s',a';θ - ) Target value r+γmax a’ Q(s',a';θ- Using this method, the weights of the main QN131 are trained using DQN loss calculation 132.

[0060] The strategy for the path that the mobile terminal device 2 should sequentially take from the initial position where the battery level of the battery 211 falls below a threshold to the destination position where the power supply device 2B is located, obtained through learning by the first learning unit 13, is stored in the first storage unit 15. Furthermore, the constructed and learned reinforcement learning model is used as training data in learning by the second learning unit 14.

[0061] The second learning unit 14 learns the relationship between the current position of the mobile terminal device 2 and the strategy for the path that the mobile terminal device 2 should take sequentially from its current position, which was obtained through learning by the first learning unit 13, using a supervised learning model.

[0062] Figure 5 shows the structure of a neural network model adopted as an example of a supervised learning model used by the second learning unit 14. The neural network model comprises an input layer x, a hidden layer h, and an output layer y. The second learning unit 14 provides the current unit space position of the mobile terminal device 2, i.e., the unit space position corresponding to the GPS position of the mobile terminal device 2 at each time t, to the input layer of the neural network model, applies an activation function to the weighted sum of the inputs, and passes the output determined by thresholding to the output layer. Each output node of the output layer outputs the model's predicted output corresponding to the n action-value functions Q of each subpath.

[0063] The second learning unit 14 learns the parameters of the neural network model so that the predicted path that the mobile terminal device 2 should take sequentially from its current position, which is the predicted value from the neural network model for the current position of the mobile terminal device 2, becomes the value of the optimized path that the mobile terminal device 2 should take sequentially from its current position, which has been reinforced by the first learning unit 13 for each subpath.

[0064]

number

[0065] In equation (2) above, y1, y2, ..., y n The predicted output values ​​for each output node are shown. Also, Y1, Y2, ..., Y n This is the training data, and here, n optimized action-value functions Q(s) for the current position are obtained by reinforcement learning by the first learning unit 13. t a1), Q(s t a2), Q(s t ,a3),...,Q(s t ,a n-1 ), Q(s t ,a n )

[0066] The value of the objective function E in equation (2) above is the output value y1, y2, ..., y for the unit space position corresponding to the GPS position of the mobile terminal device 2 at time t, which is the input value x of the supervised learning model. n-1 ,y n The target output of the training data is Y1, Y2, ..., Y n-1 ,Y n The value becomes 0 when it matches the condition. The second learning unit 14 adjusts the weight parameters of the neural network related to the supervised learning model so that the objective function E is minimized, i.e., becomes 0. The second learning unit 14 can optimize the objective function E using methods such as backpropagation.

[0067] The first memory unit 15 stores the trained reinforcement learning model constructed by the first learning unit 13 through reinforcement learning.

[0068] The second memory unit 16 stores the trained supervised learning model constructed by the second learning unit 14 through supervised learning.

[0069] The third storage unit 17 stores location information of the unit space constituting the mobile space, and identification information of the mobile terminal device 2 and the power supply device 2B. The IP addresses and IMSIs of the mobile terminal device 2 and the power supply device 2B can be used as identification information.

[0070] The configuration unit 18 sets the trained supervised learning model in the mobile terminal device 2 as control information that controls the path the mobile terminal device 2 takes from its initial location to its destination location. For example, the configuration unit 18 can transmit the control information to the mobile terminal device 2 via a network NW.

[0071] [Control device hardware configuration] Next, an example of a hardware configuration for realizing the control device 1 having the functions described above will be explained using Figure 6.

[0072] As shown in Figure 6, the control device 1 can be implemented by a computer equipped with a processor 102, main memory 103, communication interface 104, auxiliary storage 105, and input / output I / O 106 connected via a bus 101, and a program for controlling these hardware resources. Furthermore, the control device 1 may include a display device 107 connected via the bus 101.

[0073] Processor 102 is implemented using CPUs, GPUs, FPGAs, ASICs, etc.

[0074] The main memory 103 contains pre-stored programs for the processor 102 to perform various controls and calculations. The processor 102 and the main memory 103 work together to realize the various functions of the control device 1, such as the first acquisition unit 10, the second acquisition unit 11, the third acquisition unit 12, the first learning unit 13, the second learning unit 14, and the setting unit 18, as shown in Figure 1.

[0075] The communication interface 104 is an interface circuit for networking the control device 1 with various external electronic devices.

[0076] The auxiliary storage device 105 consists of a read / write storage medium and a drive device for reading and writing various information such as programs and data to the storage medium. The auxiliary storage device 105 can use semiconductor memory such as a hard disk or flash memory as the storage medium.

[0077] The auxiliary storage device 105 has a program storage area for storing the control program executed by the control device 1. It also has a program storage area for storing the reinforcement learning program executed by the control device 1. Furthermore, the auxiliary storage device 105 has an area for storing the supervised learning program. The auxiliary storage device 105 enables the realization of the first storage unit 15, the second storage unit 16, and the third storage unit 17 described in Figure 1. The auxiliary storage device 105 also has an area for storing the position coordinates of the mobile space and the position coordinates of the unit space. Furthermore, the auxiliary storage device 105 has an area for storing identification information such as the IP address and IMSI of the mobile terminal device 2 and the power supply device 2B. In addition, it may have, for example, a backup area for backing up the above-mentioned data and programs.

[0078] The I / O106 is an input / output device that accepts signals from external devices and outputs signals to external devices.

[0079] The display device 107 is composed of an organic EL display or a liquid crystal display. The display device 107 can display a map of the moving space, as well as the current location and route of the mobile terminal device 2 and the power supply device 2B.

[0080] [Functional blocks of mobile terminal devices] Next, the functional blocks of the mobile terminal device 2 will be explained with reference to Figure 7.

[0081] The mobile terminal device 2 comprises a transmission unit 20, a fourth storage unit 21, a fourth acquisition unit 22, a fifth storage unit 23, a fifth acquisition unit 24, a calculation unit 25, a determination unit 26, and a movement control unit 27. Based on the control information set by the control device 1, the mobile terminal device 2 determines the next path it should take from its current position and controls its movement to the destination point, the power supply device 2B.

[0082] The transmitting unit 20 transmits an anomaly occurrence information along with a location registration request signal when the remaining power of its own device falls below a threshold. The anomaly occurrence information indicates that the battery 211 of the mobile terminal device 2 needs to be charged and requests the control device 1 to control the route, i.e., to learn the optimal route to the power supply device 2B. The mobile terminal device 2 monitors the remaining power of the battery 211 and, when it falls below the threshold, transmits a location registration request signal with the anomaly occurrence information attached to it to the core network 4 via the base station 3. The transmitting unit 20 can transmit the location information of its own device along with the location registration request signal with the anomaly occurrence information attached. The threshold set for the remaining power is determined considering the size of the mobile space, the computational load required by the mobile terminal device 2, and the power required for movement control to the power supply device 2B.

[0083] Furthermore, the transmitting unit 20 transmits a location registration request signal with its own location information added at regular intervals. The location information is the current GPS location received by the GPS receiver 207.

[0084] The fourth memory unit 21 stores control information set by the setting unit 18 of the control device 1. The control information is a pre-trained supervised learning model in which the second learning unit 14 of the control device 1 has learned through supervised learning the strategy for the path that the mobile terminal device 2 should take sequentially from its current position until it reaches the destination point from the initial point. The control information includes the location information of the power supply device 2B, which is the destination point.

[0085] The fourth acquisition unit 22 acquires control information set by the control device 1. Specifically, the fourth acquisition unit 22 loads the control information stored in the fourth storage unit 21.

[0086] The fifth storage unit 23 stores map data including the position coordinates of the moving space, and information that associates the position coordinates of the unit spaces constituting the moving space with the node IDs of the unit spaces. The fifth storage unit 23 also stores the position information of the destination point.

[0087] The fifth acquisition unit 24 acquires the current position of the device. Based on the GPS position of the device, the fifth acquisition unit 24 acquires the position in the unit space where the device is located at each time t. The fifth acquisition unit 24 refers to the fifth storage unit 23 and acquires the position in the unit space corresponding to the current GPS position received by the GPS receiver 207 as the current position of the device.

[0088] The calculation unit 25 provides the current position of the device, acquired by the fifth acquisition unit 24, as an unknown input to the trained supervised learning model, performs calculations on the trained supervised learning model, and outputs a strategy for the path the device should sequentially move from its current position.

[0089] The determination unit 26 determines the next path to take from the current position of the device, based on the strategy for the path the device should sequentially take from its current position, which is output by the calculation unit 25. More specifically, based on the path strategy output by the calculation unit 25, the current position in the unit space is determined to state s t For each state s t The system then determines the path to take by selecting the path that maximizes the value of the action-value function Q, which is action a.

[0090] The movement control unit 27 controls the movement of the device based on the strategy for the path the device should take sequentially from its current position, which is output by the calculation unit 25. Specifically, the movement control unit 27 controls the movement of the device based on the next path to be taken, which is output by the calculation unit 25 and determined by the determination unit 26. The movement control unit 27 can calculate a control command for the next path to be taken from the current position and transmit a control command value to the motor 209. In this way, the mobile terminal device 2 controls each state s tBy selecting the action a that maximizes the value of the action-value function Q, it is possible to move to the destination point where the power supply device 2B is located via the optimal path.

[0091] [Hardware configuration of mobile terminal devices] Next, an example of a hardware configuration for realizing the mobile terminal device 2 having the functions described above will be explained using Figure 8.

[0092] As shown in Figure 8, the mobile terminal device 2 can be realized by, for example, a microcomputer equipped with a processor 202 connected via a bus 201, main memory 203, communication interface 204, auxiliary storage 205, and input / output I / O 206, a program to control these hardware resources, a GPS receiver 207, a sensor 208, a motor 209, a drive mechanism 210, a battery 211, and a power management module 212. A controller that controls the autonomous movement of the mobile terminal device 2 is realized by a computer such as a microcomputer and a program.

[0093] The main memory 203 contains pre-stored programs for the processor 202 to perform movement control and calculations. The processor 202 and the main memory 203 work together to realize the various functions of the mobile terminal device 2, such as the transmission unit 20, the fourth acquisition unit 22, the calculation unit 25, the determination unit 26, and the movement control unit 27, as shown in Figure 7.

[0094] The communication interface 204 is an interface circuit for network connection between the mobile terminal device 2 and the control device 1. The communication interface 204 also supports short-range wireless communication standards such as Bluetooth Low Energy (BLE, registered trademark) with the power supply device 2B.

[0095] The auxiliary storage device 205 consists of a read / write storage medium and a drive device for reading and writing various information such as programs and data to the storage medium. The auxiliary storage device 205 can use semiconductor memory such as a hard disk or flash memory as the storage medium.

[0096] The auxiliary storage device 205 has a program storage area for storing the mobile control program executed by the mobile terminal device 2. The auxiliary storage device 205 also has an area for storing calculation programs for performing calculations on a trained supervised learning model. Furthermore, the auxiliary storage device 205 has an area for storing a battery management program for battery management. The auxiliary storage device 205 enables the realization of the fourth storage unit 21 and the fifth storage unit 23 described in Figure 7. The auxiliary storage device 205 also has an area for storing identification information such as the IP address of the mobile terminal device 2. In addition, it may have, for example, a backup area for backing up the aforementioned data and programs.

[0097] The I / O206 is an input / output device that accepts signals from external devices and outputs signals to external devices.

[0098] The GPS receiver 207 has a built-in antenna for receiving GPS signals. The GPS receiver 207 enables the fifth acquisition unit 24 shown in Figure 7.

[0099] Sensor 208 consists of various sensors such as an altitude sensor, attitude sensor, camera, LiDAR, and RADAR. In addition to the GPS receiver 207, the altitude sensor enables the fifth acquisition unit 24 shown in Figure 6. Furthermore, the mobile controller controls the movement of the mobile terminal device 2 based on the various sensor data measured by sensor 208. Sensor 208 also includes a voltage sensor, current sensor, temperature sensor, etc., for measuring the remaining charge of the battery 211.

[0100] The motor 209 rotates due to a rotational drive, driving the drive mechanism 210 attached to the rotation shaft of the motor 209.

[0101] Battery 211 is an internal battery such as a lithium-ion battery, and supplies power to the mobile terminal device 2.

[0102] The power management module 212 includes a power input circuit, a voltage regulator, a charge management circuit, and a battery monitoring circuit. The power management module 212 monitors the remaining charge of the battery 211.

[0103] Furthermore, the mobile terminal device 2 is equipped with a SIM card and has an IMSI (International Mobile Subscriber Identity) for the SIM card.

[0104] The hardware configuration of the power supply unit 2B includes all components of the mobile terminal device 2 except for the motor 209 and the drive mechanism 210. The battery 211 of the power supply unit 2B includes an internal battery and a charging battery. The power management module 212 manages and controls the charging process.

[0105] [Control system operation] Next, the operation of the control system comprising the control device 1 and mobile terminal device 2 having the above-described configuration will be explained with reference to the sequence in Figure 9. The power supply device 2B is assumed to be installed at a fixed position within the mobile space. The mobile terminal device 2 is assumed to be performing a predetermined task while moving within the mobile space.

[0106] First, the power supply device 2B transmits a location registration request signal to the UDM / UDR41 of the core network 4 via the base station 3, adding its own GPS location (step S100). The UDM / UDR41 stores the transmission timestamp of the location registration request signal and the GPS location in association with the IMSI of the power supply device 2B. Subsequently, the first acquisition unit 10 of the control device 1 acquires the GPS location of the power supply device 2B from the UDM / UDR41 as the destination point (step S1).

[0107] Subsequently, when the power management module 212 of the mobile terminal device 2 detects that the remaining charge of the battery 211 has fallen below a threshold, the transmission unit 20 transmits a location registration request signal to the core network 4, adding abnormality information and GPS location (step S101). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal, the abnormality information, and the GPS location in association with the IMSI of the mobile terminal device 2. Subsequently, the second acquisition unit 11 of the control device 1 acquires the abnormality information from the UDM / UDR 41 (step S2). The acquisition of the abnormality information in step S2 triggers the execution of the following processes from steps S3 to S8.

[0108] Next, the mobile terminal device 2 transmits a location registration request signal to the core network 4, with its own GPS location added (step S102). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS location in association with the IMSI of the mobile terminal device 2. Subsequently, the third acquisition unit 12 of the control device 1 acquires the GPS location of the mobile terminal device 2 from the UDM / UDR 41 as its current location (step S3).

[0109] In step S3, the third acquisition unit 12 acquires the position of the mobile terminal device 2 in the unit space where it is currently located at each time t. The initial position can be the GPS position added to the position registration request signal along with the abnormality information in step S101. The third acquisition unit 12 acquires the position in the unit space corresponding to the current GPS position received by the GPS receiver 207 of the mobile terminal device 2 as the current position of the mobile terminal device 2.

[0110] Next, the first learning unit 13 performs the first learning process (step S4). In the first learning process, the first learning unit 13 applies a reward function to the estimated path that the mobile terminal device 2 should sequentially take from its initial position to its destination position, and updates the result to maximize the reward for the mobile terminal device 2 to reach its destination position. The first learning unit 13 then learns a strategy for the path that the mobile terminal device 2 should sequentially take from its current position using a reinforcement learning model. Details of the first learning process will be described later.

[0111] Subsequently, the first memory unit 15 stores the reinforcement learning model obtained in step S4 (step S5). Next, the second learning unit 14 learns the relationship between the current position of the mobile terminal device 2 and the strategy for the path that the mobile terminal device 2 should sequentially take from its current position, obtained in the first learning process in step S5, using a supervised learning model (second learning process) (step S6).

[0112] Specifically, the second learning unit 14 repeatedly adjusts and updates parameters such as weights and thresholds to determine the values ​​of these parameters, such that the error between the predicted output value of the sequential path to be taken (when the current GPS position of the mobile terminal device 2 in a unit space, i.e., the current state, is given to the supervised learning model as input) and the training data minimizes the objective function E in equation (2) above. In step S6, the second learning unit 14 can determine the parameters that minimize the objective function E by backpropagation or the like.

[0113] The training data used in step S6 is a strategy for the path to take sequentially from the current position in the unit space, obtained by the trained reinforcement learning model constructed in the first learning process of step S4.

[0114] Next, the second memory unit 16 stores the trained supervised learning model constructed in step S6 (step S7). Then, the setting unit 18 sets the trained supervised learning model as control information for the mobile terminal device 2 (step S8). In step S8, the setting unit 18 can transmit the trained supervised learning model to the mobile terminal device 2 via the network NW. Subsequently, the mobile terminal device 2 performs movement control based on the control information, as will be described later in Figure 12.

[0115] Next, the first learning process by the control device 1 (step S4 in Figure 9) will be explained using the flowcharts in Figures 10 and 11. First, step S3, which was explained in Figure 9, is executed. First, the first learning unit 13 provides the current state of the mobile terminal device 2, which is the position in the unit space where the mobile terminal device 2 is currently located, as input to the neural network model, performs calculations on the neural network model, and outputs a first estimated value Q1 of the action-value function, which represents the expected value of the cumulative value of future rewards obtained when each action related to movement in a predetermined direction relative to the direction of travel is taken as the next path that the mobile terminal device 2 should take from its current position in the unit space (step S20).

[0116] Next, the first acquisition unit 10 acquires the position of the mobile terminal device 2 in the unit space at the next time t as the next state s' (step S21). The next position in the unit space reached by the mobile terminal device 2 is determined based on the GPS position of the mobile terminal device 2 acquired by the third acquisition unit 12 at each time step. Furthermore, the first learning unit 13 provides the position in the unit space reached by the mobile terminal device 2, acquired in step S21, as input to the neural network model, performs calculations on the neural network model, and outputs the second estimated value Q2 of the action-value function (step S22).

[0117] Next, the first learning unit 13 calculates the target value from the second estimate Q2 (step S23). Subsequently, the first learning unit 13 learns the weight parameters of the neural network model so that the first estimate Q1 becomes the target value calculated from the second estimate Q2 (step S24). Specifically, the first learning unit 13 updates the weight parameters of the neural network model to minimize the loss function in equation (1) above.

[0118] Subsequently, the first memory unit 15 stores the trained reinforcement learning model obtained in step S24 (step S5).

[0119] Next, referring to Figure 11, we will explain the first learning process performed by the first learning unit 13 when a Fixed Target Q-Network is adopted, which uses two neural networks: the main QN131 and the target QN133.

[0120] The process in step S3 is the same as the steps of the first learning process described in Figure 10. Subsequently, the first learning unit 13 provides the main QN 131 with the position in the unit space where the mobile terminal device 2 is currently located, obtained in step S3, as input, performs calculations on the neural network model, outputs the action-value function Q, and calculates the next path a to take (step S120).

[0121] Next, the first learning unit 13 returns the action of the mobile terminal device 2 along the path a determined in step S130 to the environment 130, and obtains the next state s' of the mobile terminal device 2, which is the position in the unit space where the mobile terminal device 2 has moved and the reward r (step S121).

[0122] The first learning unit 13 saves the experience (s, a, r, a') obtained in step S121 to the experience data 134 (step S122). Next, in the DQN loss calculation 132, the first learning unit 13 calculates the loss function L and updates the weights of the main QN 131 using the gradient of the loss function L (step S123). The first learning unit 13 repeats the process from step S120 to step S123 a set number of times.

[0123] Subsequently, the first learning unit 13 periodically copies the weights of the main QN131 to the target QN133 and synchronizes them (step S124). The synchronization of the target QN133 is performed at a lower frequency than the update frequency of the weights of the main QN131. Next, the first learning unit 13 extracts experience from the experience data 134, inputs the past state into the target QN133, and estimates the max a’ Q(s',a';θ - Output ) (step S126).

[0124] Next, the first learning unit 13 processes the estimated value max output by the target QN133. a’Q(s',a';θ - ) Target value r+γmax a’ Q(s',a';θ - The first learning unit 13 calculates the target value (step S127). Next, the first learning unit 13 calculates the loss function L in the DQN loss calculation 132 using the target value calculated in step S127 (step S128). Next, the first learning unit 13 learns the weights of the main QN 131 to minimize the loss given by the loss function L (step S129). After that, the trained reinforcement learning model is stored in the first memory unit 15 (step S5).

[0125] Next, the operation of the mobile terminal device 2 having the above-described configuration will be explained with reference to the flowchart in Figure 12. Below, each process after the control information is set in the mobile terminal device 2 in step S8 of Figure 10 will be explained.

[0126] First, the fourth acquisition unit 22 of the mobile terminal device 2 acquires the trained supervised learning model from the fourth storage unit 21 (step S30). The fourth acquisition unit 22 reads out the control information transmitted from the control device 1 and stored in the fourth storage unit 21, i.e., the trained supervised learning model.

[0127] Next, the fifth acquisition unit 24 acquires the current position of the device in unit space as the current position (step S10). Specifically, it can acquire the position in unit space corresponding to the GPS position received by the GPS receiver 207 as the current position of the device. Next, the calculation unit 25 uses the control information acquired in step S30, provides the current position in unit space of the device acquired in step S31 as an unknown input, performs calculations on the trained supervised learning model, and outputs a strategy for the path to be taken sequentially from the current position in unit space (step S32).

[0128] For example, in the mobile space A shown in Figure 2, if the initial position of the mobile terminal device 2 is input to a pre-trained supervised learning model as its current position at time t=1, then the model will output n action-value functions Q indicating the next steps to take from the initial position at time t=1.

[0129] Next, the decision unit 26 determines the path to be taken sequentially by selecting the path that takes the action a with the maximum value of the n action value functions Q output in step S32 (step S33). Next, the movement control unit 27 controls the movement of the device based on the path to be taken next, which was determined in step S33 (step S34). More specifically, the movement control unit 27 can calculate a control command for the next path to be taken from the current position and transmit the control command value to the motor 209.

[0130] The mobile terminal device 2 repeats the processes from step S31 to step S34 until it reaches the unit space where the power supply device 2B at the destination point is located (step S35: NO). Then, when the mobile terminal device 2 reaches the unit space location of the destination point (step S35: YES), it connects to the power supply device 2B and starts charging (step S36). In this way, the mobile terminal device 2 can autonomously move from the initial point where the battery level of the battery 211 falls below a threshold to the destination point where the power supply device 2B is located and start charging by executing the processes from step S30 to step S36 using a trained supervised learning model which is the control information.

[0131] As described above, the control device 1 according to this embodiment learns the optimal route strategy for the mobile terminal device 2 from the initial point to the destination point where the power supply device 2B is located, using reinforcement learning. Using the route strategy obtained through reinforcement learning as training data, the relationship between the current position of the mobile terminal device 2 and the route strategy to be followed sequentially is learned using a supervised learning model. Furthermore, the learned supervised learning model is set in the mobile terminal device 2 as control information to control the route. Therefore, it is possible to find the optimal travel route with a simpler configuration and support efficient power supply to IoT terminals.

[0132] Furthermore, according to the control device of this embodiment, the control device learns, through reinforcement learning, the optimal route for the mobile terminal device 2 from the initial point to the destination point where the power supply device 2B is located, using the abnormality information indicating that the remaining power level has fallen below a threshold, which is added to the position registration request signal transmitted by the mobile terminal device 2, and the position information of the mobile terminal device 2. Therefore, even when the travel space is wider, the optimal travel route can be determined with a simpler configuration.

[0133] Furthermore, according to the control system of this embodiment, in order to set control information in the mobile terminal device 2, it is possible to realize a mobile terminal device 2 that can autonomously move along the optimal route from the initial point where the remaining power falls below a threshold to the destination point where the power supply device 2B is located.

[0134] [Differentiation] Next, a modified example of the embodiment of the present invention will be described. In the embodiment described above, a configuration was described in which control information indicating the optimal path from the initial point where the remaining power of the mobile terminal device 2 falls below a threshold to the position of the power supply device 2B fixedly arranged in the moving space is set in the mobile terminal device 2. That is, in the embodiment described above, the case in which the first device that performs movement is the mobile terminal device 2 and the second device that is fixedly arranged is the power supply device 2B was described.

[0135] In contrast, in this modified example, as shown in Figure 13, when the remaining power of terminal device 2' falls below a threshold, control information indicating the optimal route to terminal device 2', which is fixedly positioned within the mobile space A, is set in the mobile power supply device 2B'. Therefore, in this modified example, the first device that moves is the mobile power supply device 2B', and the second device that is fixedly positioned is terminal device 2'.

[0136] The configuration of the control device 1 and the control system in this modified example is the same as the configuration of the control device 1 and the control system described in Figure 1. Furthermore, the mobile power supply device 2B' provided in the control system according to this modified example has the same functional blocks as the mobile terminal device 2 according to this embodiment, as described in Figure 7. In this modified example, the transmitting unit 20 of the mobile power supply device 2B' transmits a location registration request signal with its own location information added at regular intervals, but abnormality information is transmitted by the terminal device 2'.

[0137] The hardware configuration of the mobile power supply device 2B' is the same as that of the mobile terminal device 2 according to this embodiment, as described in Figure 8. In this modified example, the battery 211 comprises an internal battery and a charging battery for charging the terminal device 2'. The power management module 212 manages and controls the charging process.

[0138] Terminal device 2' is fixedly positioned in the moving space and includes the transmitter 20, which is part of the functional block of mobile terminal device 2 described in Figure 7. Furthermore, the hardware configuration of terminal device 2' includes all components of mobile terminal device 2 described in Figure 8 except for the motor 209 and the drive mechanism 210.

[0139] [Control system operation] Next, the operation sequence of the control system according to this modified example will be explained with reference to the sequence diagram in Figure 14. First, terminal device 2' is fixedly positioned at a predetermined location in the moving space and monitors the remaining power supply of its own device. Mobile power supply device 2B' is movably positioned at any location within the moving space.

[0140] First, terminal device 2' transmits a location registration request signal to the UDM / UDR41 of the core network 4 via base station 3, adding its own GPS location (step S110). The UDM / UDR41 stores the transmission timestamp of the location registration request signal and the GPS location in association with the IMSI of terminal device 2'. Subsequently, the first acquisition unit 10 of control device 1 acquires the GPS location of terminal device 2' from the UDM / UDR41 as the destination point (step S1A).

[0141] Subsequently, when it is detected that the remaining power in terminal device 2' has fallen below a threshold, it sends a location registration request signal to the core network 4 with an error occurrence information and GPS location added (step S111). The UDM / UDR41 stores the transmission timestamp of the location registration request signal, the error occurrence information, and the GPS location in association with the IMSI of terminal device 2'. Then, the second acquisition unit 11 of the control device 1 acquires the error occurrence information from the UDM / UDR41 (step S2A). The acquisition of the error occurrence information in step S2A triggers the execution of the following processes from steps S3A to S8A.

[0142] Next, the mobile power supply device 2B' transmits a location registration request signal to the core network 4, with its own GPS location added (step S112). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS location in association with the IMSI of the mobile power supply device 2B'. Subsequently, the third acquisition unit 12 of the control device 1 acquires the GPS location of the mobile power supply device 2B' from the UDM / UDR 41 as its current location (step S3A).

[0143] Next, the first learning unit 13 performs the first learning process (step S4). In the first learning process, the first learning unit 13 applies a reward function to the estimated path that the mobile power supply device 2B' should sequentially take from its initial position to its destination position, and updates the result to maximize the reward for the mobile power supply device 2B' to reach its destination position. The first learning unit 13 then learns a strategy for the path that the mobile power supply device 2B' should sequentially take from its current position using a reinforcement learning model.

[0144] Subsequently, the first memory unit 15 stores the reinforcement learning model obtained in step S4 (step S5). Next, the second learning unit 14 learns the relationship between the current position of the mobile power supply device 2B' and the strategy for the path that the mobile power supply device 2B' should sequentially take from its current position, obtained in the first learning process in step S5, using a supervised learning model (second learning process) (step S6).

[0145] Next, the second memory unit 16 stores the trained supervised learning model constructed in step S6 (step S7). Then, the setting unit 18 sets the trained supervised learning model as control information for the mobile power supply device 2B' (step S8A). In step S8A, the setting unit 18 can transmit the trained supervised learning model to the mobile power supply device 2B' via the network NW. Subsequently, the mobile power supply device 2B' performs movement control based on the control information.

[0146] The movement control performed by the mobile power supply device 2B' based on the control information corresponds to the processing from step S30 to step S35 of the mobile terminal device 2 in this embodiment, as explained in Figure 12. When the mobile power supply device 2B' reaches the destination location of terminal device 2', it connects to terminal device 2' and charges terminal device 2'.

[0147] As described above, according to a modified version of this embodiment, the optimal path strategy for the mobile power supply device 2B' from the initial point, which is the position of the mobile power supply device 2B' when the remaining power of the terminal device 2' falls below a threshold, to the destination point where the terminal device 2' is located is learned by reinforcement learning. The path strategy obtained by reinforcement learning is used as training data, and the relationship between the current position of the mobile power supply device 2B' and the path strategy to be followed sequentially is learned using a supervised learning model. Furthermore, the learned supervised learning model is set in the mobile power supply device 2B' as control information to control the path. Therefore, it is possible to find the optimal travel route with a simpler configuration and support efficient power supply to IoT terminals.

[0148] In the embodiment described, the reinforcement learning model used by the first learning unit 13 is exemplified as a DQN related to a Fixed Target Q-Network composed of a multilayer neural network. However, other reinforcement learning models such as CNNs and multilayer perceptrons can be used. In addition to the DQN exemplified as a reinforcement learning model, Double DQN, Dueling DQN, Actor-Critic (AC) method, Soft Actor-Critic (SAC), Deep Deterministic Policy Gradient (DDPG), Q-learning, etc., can also be used.

[0149] Furthermore, in the embodiment described, the supervised learning model used by the second learning unit 14 was exemplified as a multilayer neural network. However, the supervised learning model can also be a multilayer perceptron, a decision tree-based model such as a random forest, or a support vector machine.

[0150] Although embodiments of the control device, control method, and control system of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications that a person skilled in the art can envision are possible within the scope of the invention described in the claims. [Explanation of symbols]

[0151] 1...Control device, 10...First acquisition unit, 11...Second acquisition unit, 12...Third acquisition unit, 13...First learning unit, 14...Second learning unit, 15...First storage unit, 16...Second storage unit, 17...Third storage unit, 18...Setting unit, 2...Mobile terminal device, 2B...Power supply device, 20...Transmission unit, 21...Fourth storage unit, 22...Fourth acquisition unit, 23...Fifth storage unit, 24...Fifth acquisition unit, 25...Calculation unit, 26...Decision unit, 27...Movement control unit, 101, 201...Bus, 102, 202...Processor S, 103, 203... Main memory, 104, 204... Communication interface, 105, 205... Auxiliary memory, 106, 206... Input / output I / O, 107... Display device, 207... GPS receiver, 208... Sensor, 209... Motor, 210... Drive mechanism, 211... Battery, 212... Power management module, 130... Environment, 131... Main QN, 132... DQN loss calculation, 133... Target QN, 134... Experience data, NW... Network.

Claims

1. A control device for controlling the path of a first device that moves to the position of a set destination point in a moving space, A first acquisition unit configured to acquire abnormality information regarding the remaining power of the first device or the second device located in the moving space, A second acquisition unit configured to acquire the position of the second device as the position of the destination point, A third acquisition unit configured to acquire the current position of the first device, which is added to the position registration request signal transmitted by the first device, A first learning unit is configured to apply a reward function to an estimated path calculated by the first device from its initial position to the destination position, update the result so that the reward for the first device to reach the destination position is maximized, and learn a strategy for the path the first device should take sequentially from its current position using a reinforcement learning model. A second learning unit is configured to learn, using a supervised learning model, the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach the destination point, which is obtained through learning by the first learning unit. A storage unit configured to store the trained supervised learning model constructed by the second learning unit, A control device equipped with the following features.

2. In the control device according to claim 1, The second acquisition unit acquires the position of the second device, which is attached to the position registration request signal transmitted by the second device, as the position of the destination point. A control device characterized by the following features.

3. In the control device according to claim 1, The first device transmits the abnormality information added to the location registration request signal when the remaining power of the device falls below a threshold value. The second device is a power supply device that supplies power to the first device, The third acquisition unit acquires the current position of the first device, triggered by the acquisition of the abnormality information by the first acquisition unit. A control device characterized by the following features.

4. In the control device according to claim 1, The first device is a power supply device that supplies power to the second device, The second device, when the remaining power of its own device falls below a threshold, adds the abnormality information to the location registration request signal and transmits it. The third acquisition unit acquires the current position of the first device, triggered by the acquisition of the abnormality information by the first acquisition unit. A control device characterized by the following features.

5. In the control device according to claim 1, Furthermore, the system includes a setting unit configured to set the previously learned supervised learning model in the first device as control information that controls the path the first device takes from the initial location to the destination location. Control device.

6. A control method for controlling the path of a first device that moves to a designated destination point in a moving space, A first acquisition step of acquiring abnormality information regarding the remaining power of the first device or the second device located in the moving space, A second acquisition step in which the position of the second device is acquired as the position of the destination point, A third acquisition step of acquiring the current position of the first device, which is added to the position registration request signal transmitted by the first device, A first learning step involves applying a reward function to an estimated path calculated by first device from its initial position to the destination position, updating the path so as to maximize the reward for first device reaching the destination position, and learning a strategy for the path first device should take sequentially from its current position using a reinforcement learning model. A second learning step involves learning, using a supervised learning model, the relationship between the current position of the first device and the strategy for the path the first device should take sequentially from its current position to reach the destination point, which was obtained through learning in the first learning step. A storage step in which the trained supervised learning model constructed in the second learning step is stored in the memory unit. A control method comprising the following features.

7. In the control method described in claim 6, The second acquisition step involves acquiring the position of the second device, which is attached to the position registration request signal transmitted by the second device, as the position of the destination point. A control method characterized by the following:

8. In the control method described in claim 6, The first device transmits the abnormality information added to the location registration request signal when the remaining power of the device falls below a threshold value. The second device is a power supply device that supplies power to the first device, The third acquisition step, triggered by the acquisition of the abnormality information in the first acquisition step, acquires the current position of the first device. A control method characterized by the following:

9. In the control method described in claim 6, The first device is a power supply device that supplies power to the second device, The second device, when the remaining power of its own device falls below a threshold, adds the abnormality information to the location registration request signal and transmits it. The third acquisition step, triggered by the acquisition of the abnormality information in the first acquisition step, acquires the current position of the first device. A control method characterized by the following:

10. In the control method described in claim 6, Furthermore, the system includes a setting step in which the trained supervised learning model is set in the first device as control information that controls the path the first device takes from the initial location to the destination location. Control method.

11. In the control method described in claim 6, Furthermore, the first device, A fourth acquisition step involves acquiring the trained supervised learning model constructed in the second learning step, A fifth acquisition step involves obtaining the current position of the device, A calculation step in which the current position of the device obtained in the third acquisition step is given as an unknown input to the trained supervised learning model, the trained supervised learning model performs calculations and outputs a strategy for the path the device should sequentially take from its current position, A movement control step that controls the movement of the device from the initial point to the destination point based on the strategy for the path the device should sequentially take from its current position, which is output in the calculation step; A control method comprising the following features.

12. A control device according to any one of claims 1 to 5, The first apparatus and A control system comprising, The first apparatus is A fourth acquisition unit configured to acquire the trained supervised learning model constructed by the control device, A fifth acquisition unit configured to acquire the current position of the device, A calculation unit is configured to provide the current position of the device acquired by the third acquisition unit as an unknown input to the trained supervised learning model, perform calculations on the trained supervised learning model, and output a strategy for the path the device should sequentially take from its current position. A movement control unit is configured to control the movement of the device from the initial point to the destination point based on a strategy for the path the device should sequentially take from its current position, which is output by the calculation unit. A control system equipped with the following features.