Control device, control method, and control system
The control device optimizes IoT terminal routes to power supply devices using reinforcement and supervised learning, addressing inefficient power supply by learning optimal paths.
Patent Information
- Application Number
- JP2024202998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Existing IoT terminals face challenges in finding an optimal movement route to a charging facility when their battery level falls below a threshold, leading to inefficient power supply due to pre-set fixed routes.
A control device utilizing reinforcement and supervised learning models to learn and optimize the movement route of IoT devices to power supply devices, incorporating reward functions and policy learning to maximize efficiency.
Enables IoT terminals to find an optimal movement route with a simpler configuration, ensuring efficient power supply by learning the relationship between current positions and destination points.
Smart Images

Figure 0007712459000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a control device, a control method, and a control system.
Background Art
[0002] Conventionally, in order to reduce the power consumption of IoT terminals, communication standards such as LPWA (Low Power Wide Area, registered trademark) and eDRX (extended Discontinuous Reception, registered trademark) have been used. By using these communication standards, IoT terminals can use their power sources for a longer time. However, when the remaining power of the IoT terminal runs out, the IoT terminal cannot operate. Therefore, Patent Document 1 discloses an IoT terminal that moves autonomously and determines whether to move to a charging facility based on the remaining operating time based on the remaining power and the time until it reaches a charging facility.
[0003] However, in Patent Document 1, the movement route for moving the IoT terminal to the charging facility is a pre-set fixed route. Therefore, the movement route from the position when the remaining power of the IoT terminal falls below the threshold value to the charging facility is not necessarily the optimal route. Therefore, it has been difficult to support efficient power supply with a simpler configuration while the remaining power of the IoT terminal is limited.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the prior art, when an abnormality occurs regarding the remaining battery level of an IoT terminal, it has been difficult to find an optimal movement route with a simpler configuration and support efficient power supply to the IoT terminal.
[0006] The present invention has been made to solve the above-described problems, and an object thereof is to find an optimal movement route with a simpler configuration and support efficient power supply to an IoT terminal.
Means for Solving the Problems
[0007] In order to solve the above-described problems, a control device according to the present invention is a control device that controls a route of a first device that moves to a position of a destination point set in a movement space, and is configured to acquire abnormality occurrence information regarding the remaining battery level of the first device or a second device arranged in the movement space, a second acquisition unit configured to acquire the position of the second device as the position of the destination point, a third acquisition unit configured to acquire the current position of the first device added to a position registration request signal transmitted by the first device, a first learning unit configured to apply a reward function to an estimation result obtained by calculating a route that the first device should sequentially follow from the position of the initial point of the first device to the position of the destination point, and update the reward so that the reward for the first device to reach the position of the destination point is maximized, and learn a policy of a route that the first device should sequentially follow from the current position using a reinforcement learning model, a second learning unit configured to learn a relationship between the current position of the first device and a policy of a route that the first device should sequentially follow from the current position until the first device reaches the destination point, obtained by learning by the first learning unit, using a supervised learning model, and a storage unit configured to store the learned supervised learning model constructed by the second learning unit.
[0008] Further, in the control device according to the present invention, the second acquisition unit may acquire the position of the second device added to a position registration request signal transmitted by the second device as the position of the destination point.
[0009] In addition, in the control device of the present invention, when the remaining power of the first device falls below a threshold value, the first device adds the abnormality occurrence information to the location registration request signal and transmits it, the second device is a power supply device that supplies power to the first device, and the third acquisition unit may acquire the current location of the first device in response to acquisition of the abnormality occurrence information by the first acquisition unit.
[0010] In addition, in the control device of the present invention, the first device is a power supply device that supplies power to the second device, and when the remaining power of the second device becomes less than a threshold value, the second device adds the abnormality occurrence information to the location registration request signal and transmits it, and the third acquisition unit acquires the current location of the first device in response to the acquisition of the abnormality occurrence information by the first acquisition unit.
[0011] In addition, the control device of the present invention may further include a setting unit configured to set the learned supervised learning model in the first device as control information for controlling a route from the position of the initial point to the position of the destination point.
[0012] In order to solve the above-mentioned problems, the control method according to the present invention is a control method for controlling a route of a first device moving to a destination position set in a moving space, the control method including a first acquisition step of acquiring abnormality occurrence information related to the remaining power supply of the first device or a second device arranged in the moving space, a second acquisition step of acquiring the position of the second device as the destination position, a third acquisition step of acquiring the current position of the first device added to a location registration request signal transmitted by the first device, and a third acquisition step of calculating a course to be taken by the first device from the initial position of the first device to the destination position. The method includes a first learning step of applying a reward function to the estimation result to update the estimation result so as to maximize the reward for the first device to reach the destination position, and learning a course plan for the first device to take from the current position using a reinforcement learning model; a second learning step of learning, using a supervised learning model, the relationship between the current position of the first device and the course plan for the first device to take from the current position until it reaches the destination position, which is obtained by learning in the first learning step; and a storage step of storing the learned supervised learning model constructed in the second learning step in a storage unit.
[0013] In the control method according to the present invention, the second acquisition step may acquire, as the location of the destination point, a location of the second device added to a location registration request signal transmitted by the second device.
[0014] Furthermore, in the control method of the present invention, when the remaining power of the first device falls below a threshold value, the first device adds the abnormality occurrence information to the location registration request signal and transmits it, the second device is a power supply device that supplies power to the first device, and the third acquisition step may acquire the current location of the first device in response to acquisition of the abnormality occurrence information in the first acquisition step.
[0015] Furthermore, in the control method of the present invention, the first device may be a power supply device that supplies power to the second device, and when the remaining power of the second device becomes less than a threshold value, the second device may add the abnormality occurrence information to the location registration request signal and transmit it, and the third acquisition step may acquire the current location of the first device in response to acquisition of the abnormality occurrence information in the first acquisition step.
[0016] In addition, the control method of the present invention may further include a setting step of setting the learned supervised learning model in the first device as control information for controlling a route from the position of the initial point to the position of the destination point.
[0017] In addition, the control method of the present invention further includes a fourth acquisition step in which the first device acquires the trained supervised learning model constructed in the second learning step, a fifth acquisition step in which the first device acquires a current position of the device, a calculation step in which the current position of the device acquired in the third acquisition step is provided as an unknown input to the trained supervised learning model, and the trained supervised learning model is calculated to output a course plan for the device to proceed sequentially from the current position of the device, and a movement control step in which the first device controls movement of the device from the initial point to the destination point based on the course plan for the device to proceed sequentially from the current position of the device output in the calculation step.
[0018] To solve the above problems, a control system according to the present invention is a control system including the above control device and the first device, wherein the first device is configured to obtain the learned supervised learning model constructed by the control device, a fifth acquisition unit configured to acquire the current position of the own device, and the current position of the own device acquired by the third acquisition unit is given as an unknown input to the learned supervised learning model, and the learned supervised learning model is calculated to output a policy for a route to be sequentially advanced from the current position of the own device, and a movement control unit configured to control the movement of the own device from the initial point to the destination point based on the policy for the route to be sequentially advanced from the current position of the own device output by the calculation unit.
Effect of the Invention
[0019] According to the present invention, the relationship between the current position of the first device and the policy for the route to be sequentially advanced from the current position until the first device reaches the destination point obtained by learning by the first learning unit is learned using a supervised learning model. Therefore, it is possible to obtain an optimal movement route with a simpler configuration and support efficient power supply to the IoT terminal.
Brief Description of the Drawings
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
[0021] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to FIGS. 1 to 14.
[0022] [Configuration of Control System] First, with reference to FIG. 1, an outline of a control system including a control device 1 and a mobile terminal device (first device) 2 according to an embodiment of the present invention will be described.
[0023] The control system according to this embodiment includes a control device 1, a mobile terminal device 2, a power supply device (second device) 2B, a base station 3, and a core network 4. The control device 1 is communicably connected to the mobile terminal device 2 and the power supply device 2B via a wireless communication network NW compliant with a predetermined communication standard such as LTE / 4G, 5G, 6G, etc. The control system controls the path for the mobile terminal device 2 to move to a destination point set in the moving space. As shown in FIGS. 1 and 2, the moving space A where the mobile terminal device 2 moves can communicate, for example, by a 5G wireless communication method.
[0024] The mobile terminal device 2 includes a flying object capable of autonomous flight such as a mobile robot, a drone, an unmanned aerial vehicle, an autonomous driving vehicle, a ship, etc. When the remaining amount of the battery 211 of the mobile terminal device 2 becomes less than a threshold value, the mobile terminal device 2 obtains the optimal path to the power supply device 2B and autonomously moves to the position of the power supply device 2B which is the destination point. As shown in FIG. 2, when the mobile terminal device 2 is moving on an arbitrary route in the moving space A, the position when the remaining power becomes less than the threshold value is set as the initial point. Then, the mobile terminal device 2 moves from the position of the initial point to the position of the power supply device 2B which is the destination point and receives power supply from the power supply device 2B. The mobile terminal device 2 and the power supply device 2B can communicate by short-range wireless communication.
[0025] Hereinafter, the case where the mobile terminal device 2 is a mobile robot will be described as an example. The mobile terminal device 2 controls autonomous movement by a controller that processes information from a sensor 208 and the like described later and controls the rotation speed of the motor 209 and the drive mechanism 210. Also, the mobile terminal device 2 obtains the GPS position of its own device by a GPS receiver 207. The mobile terminal device 2 is configured as an IoT terminal having an IP address, and the mobile terminal device 2 can be uniquely identified by each IP address.
[0026] Further, the mobile terminal device 2 according to the present embodiment includes a SIM and has an IMSI (International Mobile Subscriber Identity) stored in the SIM. Details of the functional blocks and hardware configuration of the mobile terminal device 2 will be described later. Note that there may be a plurality of mobile terminal devices 2, and in that case, each of the mobile terminal devices 2 has the same configuration.
[0027] The power supply device 2B is fixedly arranged in the moving space. The power supply device 2B is realized by a computer including a processor, a main storage device, a communication interface, an auxiliary storage device, and an input / output I / O, and a program for controlling these hardware resources. The power supply device 2B further includes a charging device and a charge control unit for supplying power to the mobile terminal device 2. More specifically, the power supply device 2B includes a power management unit, a charging module, a DC-DC converter, and a contact charging terminal. As an example, the power supply device 2B is realized as a configuration integrated with a utility pole arranged in the moving space. A plurality of power supply devices 2B may be arranged in the moving space. For example, they may be arranged in a grid pattern at a predetermined interval within the moving space A in FIG. 2.
[0028] The power supply device 2B can uniquely identify the mobile terminal device 2 by an IP address. Also, the power supply device 2B according to the present embodiment includes a SIM and is also uniquely identified by the IMSI stored in the SIM. Further, the power supply device 2B includes a GPS receiver and acquires the GPS position of its own device.
[0029] The mobile terminal device 2 and the power supply device 2B transmit a location registration request signal to the core network 4 via the base station 3 at regular intervals. The mobile terminal device 2 and the power supply device 2B transmit the location registration request signal by adding their own GPS locations. The mobile terminal device 2 can add its own location to the location registration request signal and transmit it when crossing the area of the base station 3 or when power is turned on. Similarly, the power supply device 2B can add its own location to the location registration request signal and transmit it when power is turned on. The location registration request signal includes the IMSI of the transmitting device.
[0030] The mobile terminal device 2 also adds and transmits information on the occurrence of an abnormality related to the remaining battery power to the location registration request signal when the remaining amount of the battery 211 of its own device becomes less than the threshold value.
[0031] As shown in FIG. 2, the moving space A in which the mobile terminal device 2 moves is a three-dimensional matrix-shaped space composed of unit spaces divided into a plurality of spaces. Each unit space constituting the moving space A has the same volume. Further, each unit space has a node ID, and each unit space is represented by one position (x, y, z). As the position information, three-dimensional GPS position coordinates composed of latitude, longitude, and altitude can be used. For example, as the position of the unit space, a representative value such as the central position of the unit space can be used.
[0032] Also, as shown in FIG. 2, the mobile terminal device 2 moves from the position of the unit space corresponding to the initial location where the remaining amount of the battery 211 becomes less than the threshold value to the position of the unit space of the destination where the power supply device 2B is arranged, using each unit space as a waypoint.
[0033] The base station 3 is composed of radio base stations compliant with the 5G communication standard and relays communication between the mobile terminal device 2 and the power supply device 2B present in the communication area and the core network 4. The base station 3 is connected to the core network 4 via a backhaul link.
[0034] The core network 4 is connected to the control device 1 via a network NW such as a LAN, WAN, or the Internet. The core network 4 includes an AMF (Access and Mobility Management Function) 40 and a UDM (Unified Data Management) / UDR (Unified Data Repository) 41, which are nodes in the C-plane. In addition, the core network 4 includes a plurality of UPFs (User Plane Function) 42 in the U-plane. For functional nodes in the U-plane and C-plane other than those described above included in the core network 4, the illustration is omitted.
[0035] The AMF 40 is a node that provides a mobility control function and performs mobility control such as location registration, paging, and handover. The AMF 40 performs authentication, starts service connection, starts session management, etc. based on a location registration request signal from the mobile terminal device 2 and the power supply device 2B. The location information and abnormality occurrence information of the mobile terminal device 2 and the power supply device 2B added to the location registration request signal received by the AMF 40 are transmitted to the UDM / UDR 41.
[0036] The UDM / UDR 41 manages subscriber profiles, performs authentication, and mobility management. In the present embodiment, the UDM / UDR 41 adds fields for location information and abnormality occurrence information to the subscriber profile. The UDM / UDR 41 according to the present embodiment includes a communication interface 41a for communicating with the control device 1. In the UDM / UDR 41, the presence or absence of location information and abnormality occurrence information is stored using the IMSI included in the location registration request signal as a key. The transmission timestamp of the location registration request signal identifies when the mobile terminal device 2 and the power supply device 2B identified by the IMSI are present at the GPS location. In addition to the case where the UDM and the UDR are configured as one device, the UDM / UDR 41 may be a device in which the UDM and the UDR are separately arranged.
[0037] UPF42 is a user plane function that processes data between the base station 3 and a data network (DN) such as the Internet.
[0038] The control system according to the present embodiment learns an optimal route from the position of the initial point where the remaining amount of the battery 211 of the mobile terminal device 2 becomes less than the threshold value to the position of the destination point where the power supply device 2B is installed by reinforcement learning. Further, using the policy of the route of the mobile terminal device 2 obtained by reinforcement learning as teacher data, the relationship between the current position of the mobile terminal device 2 in the unit space and the policy of the route that the mobile terminal device 2 should sequentially follow is learned using a supervised learning model. Further, the learned supervised learning model is set in the mobile terminal device 2 as control information for controlling the route of the mobile terminal device 2. The mobile terminal device 2 performs the calculation of the learned supervised learning model with the current position of its own device in the unit space as an unknown input based on the set control information, and outputs a policy of the route that should be sequentially followed. Then, based on the output policy of the route that should be sequentially followed, the route from the initial point to the destination point is determined, and the movement to the power supply device 2B at the destination point is controlled.
[0039] The mobile terminal device 2 with the control information set changes the route in any direction indicated by each arrow from the mobile terminal device 2 at the initial point as shown in FIG. 2 according to the route determined based on the policy of the route obtained by the calculation of the learned supervised learning model, and moves in the direction that should be advanced for each unit space. The route can include various routes, that is, movement directions. In FIG. 2, the movement space A is described in a two-dimensional plane, but the route of the mobile terminal device 2 can be a three-dimensional route. By the control information set by the control device 1 in the mobile terminal device 2, the mobile terminal device 2 can reach the unit space where the power supply device 2B at the destination point is located from the unit space at the initial point.
[0040] Here, the route refers to the movement direction from the position of each unit space to the position of the adjacent unit space. Also, the path includes the entire route from the position of the initial point to the position of the destination point.
[0041] [Functional Blocks of the Control Device] As shown in FIG. 1, the control device 1 includes a first acquisition unit 10, a second acquisition unit 11, a third acquisition unit 12, a first learning unit 13, a second learning unit 14, a first storage unit (storage unit) 15, a second storage unit 16, a third storage unit 17, and a setting unit 18. The control device 1 controls the path of the mobile terminal device 2 to the position of the destination point set in the moving space.
[0042] The first acquisition unit 10 acquires abnormality occurrence information regarding the remaining battery level of the mobile terminal device 2 (first device). In the present embodiment, the first acquisition unit 10 acquires, via the network NW, the abnormality occurrence information added to the location registration request signal transmitted by the mobile terminal device 2 from the UDM / UDR 41 of the core network 4. The abnormality occurrence information is transmitted by adding it to the location registration request signal when the remaining amount of the battery 211 of the mobile terminal device 2 becomes less than the threshold value.
[0043] The second acquisition unit 11 acquires the position of the power supply device 2B (second device) as the position of the destination point. The second acquisition unit 11 acquires, via the network NW, the position information added to the location registration request signal transmitted by the power supply device 2B from the UDM / UDR 41 of the core network 4. As described above, the power supply device 2B adds its own GPS position to the location registration request signal at the time of installation and at a fixed cycle and transmits it to the core network 4. Further, the second acquisition unit 11 acquires the node ID of the unit space corresponding to the GPS position stored in the third storage unit 17 described later. When there are a plurality of power supply devices 2B in the moving space, based on the GPS position of the mobile terminal device 2 that is the transmission source of the location registration request signal to which the abnormality occurrence information acquired by the first acquisition unit 10 is added, the power supply device 2B with the shortest straight-line distance can be selected and set as the destination point.
[0044] The third acquisition unit 12 acquires the current location of the mobile terminal device 2 (first device) added to the location registration request signal transmitted by the mobile terminal device 2 (first device). The third acquisition unit 12 acquires the current location of the mobile terminal device 2 when the first acquisition unit 10 acquires the abnormality occurrence information. The third acquisition unit 12 acquires the location information added to the location registration request signal transmitted by the mobile terminal device 2 from the UDM / UDR 41 of the core network 4 via the network NW. As described above, the mobile terminal device 2 adds its own GPS location to the location registration request signal at a constant period and transmits it to the core network 4. The third acquisition unit 12 acquires the node ID of the unit space corresponding to the GPS location from the third storage unit 17. The third acquisition unit 12 acquires the unit space in which the mobile terminal device 2 is currently located as the current location, and the current location is the location of the unit space in which the mobile terminal device 2 exists at each time t.
[0045] The first learning unit 13 applies a reward function to the estimated result of calculating the sequential route that the mobile terminal device 2 (first device) should take from the initial position to the destination position, updates the route so as to maximize the reward for the mobile terminal device 2 to reach the destination position, and learns the route that the mobile terminal device 2 should take from its current position using a reinforcement learning model.
[0046] In this embodiment, as a course of action for the mobile terminal device 2 to proceed sequentially from the position of each unit space, the mobile terminal device 2 selects actions a related to movement in a predetermined number of directions (n is an integer of 2 or more) relative to the direction of travel. n The traveling direction is a direction based on the position of the unit space where the mobile terminal device 2 was located immediately before.
[0047] The first learning unit 13 uses a neural network model including an input layer s, a hidden layer h, and an output layer q as shown in FIG. 3 as a reinforcement learning model. In addition, as the neural network model, a state s t , and all the action value functions Q(s t ,a1), Q(s t ,a2), Q(s t, a3), ···, Q(s t , a n-1 ), Q(s t , a n ) outputs the Deep Q-Network (DQN), which is a neural network.
[0048] More specifically, the first learning unit 13 gives, as an input to the neural network model, the position of the current unit space indicating the position of the current mobile terminal device 2, performs the operation of the neural network model, and as the route that the mobile terminal device 2 should next proceed from the position of the current unit space, the action a related to each movement in n directions n outputs the first estimated value Q1 of the action value function representing the expected value of the cumulative value of the future reward obtained when taking.
[0049] The reward is given by the reward function r = r(s, a, s') of the state s indicating the current position of the mobile terminal device 2, the action a of the mobile terminal device 2 moving in a predetermined direction n , and the next position of the mobile terminal device 2, that is, the next state s'. In the present embodiment, the reward function includes, as a variable, the degree of reach to the position of the unit space of the power supply device 2B related to the destination point of the mobile terminal device 2. In addition, it can include, as a variable, the degree of reach to the position of the unit space corresponding to the space with an obstacle. For example, when approaching the destination point or reaching the power supply device 2B at the shortest distance by the action related to the movement of the mobile terminal device 2 in a predetermined direction, the reward, which is a scalar quantity, is set as a larger value.
[0050] On the other hand, when the mobile terminal device 2 moves away from the power supply device 2B which is the destination point, or reaches the unit space where there is an obstacle, it can be designed to give a negative reward value (for example, r = -1). In this way, by setting the reward of the unit space where there is an obstacle as a negative value, the mobile terminal device 2 can avoid these points and reach the position of the power supply device 2B.
[0051] Furthermore, the first learning unit 13 gives the position of the unit space that the mobile terminal device 2 next reaches as an input to the neural network model, performs the calculation of the neural network model, and outputs a second estimated value Q2 of the action value function. The first learning unit 13 learns the weight parameters of the neural network model so that the first estimated value Q1 becomes a target value calculated from the second estimated value Q2.
[0052] Assuming that the weight parameter of the neural network model is θ and the action value function is represented as Q(s,a;θ), the loss function to be minimized in learning is given by the following equation (1). L(θ)=1 / 2{r+γmax a’ Q(s’,a’;θ)-Q(s,a;θ)} 2 ···(1)
[0053] In the above equation (1), r is the reward (immediate reward), and γ indicates the discount rate. Q(s,a;θ) corresponds to the first estimated value Q1, and Q(s’,a’;θ) corresponds to the action value in the state s’ advanced by one step, that is, the second estimated value Q2. The target value is represented by r+γmax a’ Q(s’,a’;θ).
[0054] The first learning unit 13 can update the weight parameters of the neural network model by error backpropagation of the gradient of the loss function given by the above equation (1).
[0055] More specifically, as shown in FIG. 4, the first learning unit 13 can adopt a Fixed Target Q-Network using two neural networks, a main QN131 and a target QN133. The main QN131 selects an optimal action and updates the action value function Q. On the other hand, the target QN133 estimates and evaluates the value of the action a’ to be taken in the next state s’ as a result of the action. The main QN131 and the target QN133 have neural networks with the same layer structure, but the parameter of the main QN131 is “θ”, and the parameter of the target QN133 is “θ - ”.
[0056] The main QN131 receives the current position of the mobile terminal device 2 from the environment 130 as the state s. The environment 130 is a system of the moving space where the mobile terminal device 2 is placed. Under this environment 130, the mobile terminal device 2 moves to another unit space by taking an action a related to the movement in a predetermined direction, and while transitioning to the next state s’, it obtains a reward r from the environment 130.
[0057] The first learning unit 13 inputs the state s related to the current position of the mobile terminal device 2 to the main QN131 and obtains the action value function Q(s, a; θ). The first learning unit 13 calculates the action a using, for example, the ε-greedy method, or obtains the optimal action argmax a Q(s, a; θ) at the current time. In the environment 130, the mobile terminal device 2 takes the action argmax a Q(s, a; θ) related to the optimal route at the current time. The environment 130 observes the position of the unit space where the mobile terminal device 2 has moved as the next state s’ as a result of taking the action argmax a Q(s, a; θ), and outputs the reward r. The experience data 134 stores the experience (s, a, r, s’) output from the environment 130.
[0058] The first learning unit 13 obtains the loss function L in the DQN loss calculation 132, and updates the weights of the main QN131 with the gradient of the loss function L.
[0059] The first learning unit 13 copies the weights of the main QN131 to the target QN133 periodically for synchronization. The synchronization of the target QN133 is performed at a lower frequency than the update frequency of the weights of the main QN131. The first learning unit 13 extracts experiences from the experience data 134, inputs the past state to the target QN133, and outputs the estimated value max a’ Q(s’, a’; θ - ). The first learning unit 13 calculates the target value r + γmax a’ Q(s’, a’; θ - ) based on the estimated value max a’ Q(s’, a’; θ- ) is used to train the weights of the main QN 131 in the DQN loss calculation 132.
[0060] The route strategy that the mobile terminal device 2 should sequentially follow from the initial location where the remaining amount of the battery 211 becomes less than the threshold value to the location of the power supply device 2B obtained by the learning of the first learning unit 13, that is, the learned reinforcement learning model, is stored in the first storage unit 15. Also, the constructed learned reinforcement learning model is used as teacher data in the learning by the second learning unit 14.
[0061] The second learning unit 14 learns the relationship between the current position of the mobile terminal device 2 and the route strategy that the mobile terminal device 2 should sequentially follow from the current position using a supervised learning model.
[0062] FIG. 5 shows the structure of a neural network model adopted as an example of the supervised learning model used by the second learning unit 14. The neural network model includes an input layer x, a hidden layer h, and an output layer y. The second learning unit 14 gives the position of the current unit space of the mobile terminal device 2, that is, the position of the unit space corresponding to the GPS position of the mobile terminal device 2 at each time t, to the input layer of the neural network model, applies an activation function to the weighted sum of the inputs, and passes the output determined by the threshold process to the output layer. Each output node of the output layer outputs the predicted output of the model corresponding to the n action value functions Q of each partial route.
[0063] The second learning unit 14 introduces the objective function E shown in the following formula (2) so that the route strategy that should be sequentially followed from the current position, which is the predicted value from the neural network model for the current position of the mobile terminal device 2, becomes the value of the optimized strategy of the route that the mobile terminal device 2 should sequentially follow from the current position, which has been reinforced by the first learning unit 13 for each partial route, and learns the parameters of the neural network model.
[0064]
Equation
[0065] In the above formula (2), y1, y2, ···, y n represents the predicted output values of each output node. Also, Y1, Y2, ···, Y n is the teacher data. Here, it is the n optimized action value functions Q(s t , a1), Q(s t , a2), Q(s t , a3), ···, Q(s t , a n-1 ), Q(s t , a n ) obtained by reinforcement learning by the first learning unit 13 for the current position.
[0066] The value of the objective function E in the above formula (2) is the output values y1, y2, ···, y n-1 , y n for the position in the unit space corresponding to the GPS position of the mobile terminal device 2 at time t, which is the above input value x of the supervised learning model. When y n-1 , y n matches the target outputs Y1, Y2, ···, Y n-1 , Y n of the teacher data, it becomes 0. The second learning unit 14 adjusts the weight parameters of the neural network related to the supervised learning model so that the objective function E is minimized, that is, becomes 0. The second learning unit 14 can optimize the objective function E using the error backpropagation method or the like.
[0067] The first storage unit 15 stores the learned reinforcement learning model constructed by the reinforcement learning by the first learning unit 13.
[0068] The second storage unit 16 stores the learned supervised learning model constructed by the supervised learning by the second learning unit 14.
[0069] The third storage unit 17 stores the position information of the unit space constituting the moving space and the identification information of the mobile terminal device 2 and the power supply device 2B. The IP address or IMSI of the mobile terminal device 2 and the power supply device 2B can be used as the identification information.
[0070] The setting unit 18 sets the learned supervised learning model in the mobile terminal device 2 as control information for controlling the path from the initial location to the destination location of the mobile terminal device 2. For example, the setting unit 18 can transmit the control information to the mobile terminal device 2 via the network NW.
[0071] [Hardware Configuration of the Control Device] Next, an example of the hardware configuration for realizing the control device 1 having the above-described functions will be described with reference to FIG. 6.
[0072] As shown in FIG. 6, the control device 1 is, for example, a computer including a processor 102, a main storage device 103, a communication interface 104, an auxiliary storage device 105, and an input / output I / O 106 connected via a bus 101, and can be realized by a program for controlling these hardware resources. Further, the control device 1 can include a display device 107 connected via the bus 101.
[0073] The processor 102 is realized by a CPU, a GPU, an FPGA, an ASIC, or the like.
[0074] In the main storage device 103, programs for the processor 102 to perform various controls and calculations are stored in advance. The functions of the control device 1 such as the first acquisition unit 10, the second acquisition unit 11, the third acquisition unit 12, the first learning unit 13, the second learning unit 14, and the setting unit 18 shown in FIG. 1 are realized by the processor 102 and the main storage device 103.
[0075] The communication interface 104 is an interface circuit for network-connecting the control device 1 and various external electronic devices.
[0076] The auxiliary storage device 105 is composed of a readable and writable storage medium and a driving device for reading and writing various kinds of information such as programs and data to and from the storage medium. As the storage medium of the auxiliary storage device 105, a semiconductor memory such as a hard disk or a flash memory can be used.
[0077] The auxiliary storage device 105 has a program storage area for storing the control program executed by the control device 1. It also has a program storage area for storing the reinforcement learning program executed by the control device 1. Furthermore, the auxiliary storage device 105 has an area for storing the supervised learning program. The first storage unit 15, the second storage unit 16, and the third storage unit 17 described in FIG. 1 are realized by the auxiliary storage device 105. Also, the auxiliary storage device 105 has an area for storing the position coordinates of the moving space and the position coordinates of the unit space. Furthermore, the auxiliary storage device 105 has an area for storing identification information such as the IP addresses and IMSIs of the mobile terminal device 2 and the power supply device 2B. Furthermore, for example, it may have a backup area for backing up the above-mentioned data, programs, etc.
[0078] The input / output I / O 106 is an input / output device that inputs signals from external devices and outputs signals to external devices.
[0079] The display device 107 is composed of an organic EL display, a liquid crystal display, or the like. The display device 107 can display a map of the moving space, and the current positions and routes of the mobile terminal device 2 and the power supply device 2B.
[0080] [Function Blocks of Mobile Terminal Device] Next, the function blocks of the mobile terminal device 2 will be described with reference to FIG. 7.
[0081] The mobile terminal device 2 includes a transmission unit 20, a fourth storage unit 21, a fourth acquisition unit 22, a fifth storage unit 23, a fifth acquisition unit 24, an arithmetic unit 25, a determination unit 26, and a movement control unit 27. The mobile terminal device 2 determines the route to proceed next from its current position based on the control information set by the control device 1, and controls the movement of the device itself to the position of the power supply device 2B which is the destination.
[0082] When the remaining power of the device itself becomes less than the threshold value, the transmission unit 20 adds the abnormality occurrence information to the position registration request signal and transmits it. The abnormality occurrence information is information that the battery 211 of the mobile terminal device 2 needs to be charged, and requests the control device 1 for route control, that is, the learning process of the optimal route to the power supply device 2B. The mobile terminal device 2 monitors the remaining amount of the battery 211, and when it becomes less than the threshold value, transmits a position registration request signal with the abnormality occurrence information added to the core network 4 via the base station 3. The transmission unit 20 can add the position information of the device itself to the position registration request signal with the abnormality occurrence information added and transmit it. The threshold value set for the remaining power is set in consideration of the size of the movement space, the calculation load required for the mobile terminal device 2, the power required for movement control to the power supply device 2B, and the like.
[0083] Also, the transmission unit 20 adds the position information of the device itself to the position registration request signal and transmits it at a certain period. The position information is the current GPS position received by the GPS receiver 207.
[0084] The fourth storage unit 21 stores the control information set by the setting unit 18 of the control device 1. The control information is a learned supervised learning model in which the second learning unit 14 of the control device 1 learns the strategy of the route to sequentially proceed from the current position until the mobile terminal device 2 reaches the destination by supervised learning. The control information includes the position information of the power supply device 2B which is the destination.
[0085] The fourth acquisition unit 22 acquires the control information set by the control device 1. Specifically, the fourth acquisition unit 22 loads the control information stored in the fourth storage unit 21.
[0086] The fifth memory unit 23 stores map data including the position coordinates of the moving space, and information associating the position coordinates of the unit spaces constituting the moving space with the node IDs of the unit spaces. Further, the fifth memory unit 23 stores the position information of the destination point.
[0087] The fifth acquisition unit 24 acquires the current position of the own device. The fifth acquisition unit 24 acquires the position of the unit space where the own device is located at each time t based on the GPS position of the own device. The fifth acquisition unit 24 refers to the fifth memory unit 23 and acquires the position of the unit space corresponding to the current GPS position received by the GPS receiver 207 as the current position of the own device.
[0088] The calculation unit 25 gives the current position of the own device acquired by the fifth acquisition unit 24 as an unknown input to a learned supervised learning model, performs the calculation of the learned supervised learning model, and outputs a strategy for the route that the own device should sequentially proceed from the current position.
[0089] The determination unit 26 determines the route to be next traveled from the current position of the own device based on the strategy for the route that the own device should sequentially proceed from the current position, which is output by the calculation unit 25. More specifically, based on the route strategy output by the calculation unit 25, the position of the current unit space is set as the state s t and, for each state s t the route with the action a having the maximum value of the action value function Q is selected to determine the route to be sequentially traveled.
[0090] The movement control unit 27 controls the movement of the own device based on the strategy for the route that the own device should sequentially proceed from the current position, which is output by the calculation unit 25. Specifically, the movement control unit 27 controls the movement of the own device based on the route to be next traveled, which is output by the calculation unit 25 and determined by the determination unit 26. The movement control unit 27 can calculate a control command for the route to be next traveled from the current position and transmit a control command value to the motor 209. In this way, the mobile terminal device 2 is in each state s tAt [a certain point], by selecting the action a with the maximum value of the action value function Q, it is possible to move to the destination where the power supply device 2B is arranged along the optimal path.
[0091] [Hardware Configuration of Mobile Terminal Device] Next, an example of the hardware configuration for realizing the mobile terminal device 2 having the above-described functions will be described with reference to FIG. 8.
[0092] As shown in FIG. 8, the mobile terminal device 2 can be realized by, for example, a microcomputer including a processor 202, a main storage device 203, a communication interface 204, an auxiliary storage device 205, and an input / output I / O 206 connected via a bus 201, and a program for controlling these hardware resources, a GPS receiver 207, a sensor 208, a motor 209, a drive mechanism 210, a battery 211, and a power management module 212. A controller for controlling the autonomous movement of the mobile terminal device 2 is realized by a computer such as a microcomputer and a program.
[0093] In the main storage device 203, programs for the processor 202 to perform movement control and calculations are stored in advance. The functions of the mobile terminal device 2 such as the transmission unit 20, the fourth acquisition unit 22, the calculation unit 25, the determination unit 26, and the movement control unit 27 shown in FIG. 7 are realized by the processor 202 and the main storage device 203.
[0094] The communication interface 204 is an interface circuit for network-connecting the mobile terminal device 2 and the control device 1. The communication interface 204 also supports short-range wireless communication standards such as Bluetooth Low Energy (BLE, registered trademark) with the power supply device 2B.
[0095] The auxiliary storage device 205 is composed of a readable / writable storage medium and a drive device for reading and writing various information such as programs and data to and from the storage medium. As the storage medium, a semiconductor memory such as a hard disk or a flash memory can be used in the auxiliary storage device 205.
[0096] The auxiliary storage device 205 has a program storage area for storing the movement control program executed by the mobile terminal device 2. The auxiliary storage device 205 also has an area for storing an arithmetic program for performing the arithmetic operation of the learned supervised learning model. The auxiliary storage device 205 also has an area for storing a battery management program for performing battery management. The fourth storage unit 21 and the fifth storage unit 23 described with reference to FIG. 7 are realized by the auxiliary storage device 205. The auxiliary storage device 205 also has an area for storing identification information such as the IP address of the mobile terminal device 2. Furthermore, for example, it may have a backup area for backing up the above-described data, programs, and the like.
[0097] The input / output I / O 206 is an input / output device that inputs signals from external devices and outputs signals to external devices.
[0098] The GPS receiver 207 has a built-in antenna for receiving GPS signals. The fifth acquisition unit 24 in FIG. 7 is realized by the GPS receiver 207.
[0099] The sensor 208 is composed of various sensors such as an altitude sensor, an attitude sensor, a camera, a LiDAR, and a RADAR. In addition to the GPS receiver 207, the fifth acquisition unit 24 in FIG. 6 is realized by the altitude sensor. Also, based on the various sensor data measured by the sensor 208, the movement controller controls the movement of the mobile terminal device 2. Furthermore, the sensor 208 includes a voltage sensor, a current sensor, a temperature sensor, etc. for measuring the remaining amount of the battery 211.
[0100] The motor 209 rotates by rotational drive and drives a drive mechanism 210 attached to the rotation axis of the motor 209.
[0101] The battery 211 is an internal battery such as a lithium-ion battery and supplies power to the configuration of the mobile terminal device 2.
[0102] The power management module 212 includes a power input circuit, a voltage regulator, a charging management circuit, and a battery monitoring circuit. The power management module 212 monitors the remaining amount of the battery 211.
[0103] Furthermore, the mobile terminal device 2 includes a SIM and has the IMSI (International Mobile Subscriber Identity) of the SIM.
[0104] The hardware configuration of the power supply device 2B includes the configuration of the mobile terminal device 2 other than the motor 209 and the drive mechanism 210. The battery 211 of the power supply device 2B includes an internal battery and a charging battery. Also, the power management module 212 manages and controls the charging process.
[0105] [Operation of the control system] Next, the operation of the control system including the control device 1 and the mobile terminal device 2 having the above-described configuration will be described with reference to the sequence of FIG. 9. It is assumed that the power supply device 2B is installed at a fixed position within the moving space. It is assumed that the mobile terminal device 2 is executing a predetermined task while moving within the moving space.
[0106] First, the power supply device 2B transmits a location registration request signal with its own GPS position added to the UDM / UDR 41 of the core network 4 via the base station 3 (step S100). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS position in association with the IMSI of the power supply device 2B. Thereafter, the first acquisition unit 10 of the control device 1 acquires the GPS position of the power supply device 2B from the UDM / UDR 41 as the destination point (step S1).
[0107] After that, when the power management module 212 of the mobile terminal device 2 detects that the remaining amount of the battery 211 has fallen below the threshold, the transmission unit 20 transmits a location registration request signal with the abnormal occurrence information and the GPS location added to the core network 4 (step S101). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal, the abnormal occurrence information, and the GPS location in association with the IMSI of the mobile terminal device 2. Then, the second acquisition unit 11 of the control device 1 acquires the abnormal occurrence information from the UDM / UDR 41 (step S2). Triggered by the acquisition of the abnormal occurrence information in step S2, the processes from step S3 to step S8 below are executed.
[0108] Next, the mobile terminal device 2 transmits a location registration request signal with its own GPS location added to the core network 4 (step S102). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS location in association with the IMSI of the mobile terminal device 2. Subsequently, the third acquisition unit 12 of the control device 1 acquires the GPS location of the mobile terminal device 2 as the current location from the UDM / UDR 41 (step S3).
[0109] In step S3, the third acquisition unit 12 acquires the location of the unit space where the mobile terminal device 2 is currently located at each time t. As the location of the initial point, the GPS location added to the location registration request signal together with the abnormal occurrence information in step S101 can be used. The third acquisition unit 12 acquires the location of the unit space corresponding to the current GPS location received by the GPS receiver 207 of the mobile terminal device 2 as the current location of the mobile terminal device 2.
[0110] Next, the first learning unit 13 performs the first learning process (step S4). In the first learning process, the first learning unit 13 applies a reward function to the estimation result of calculating the route that the mobile terminal device 2 should sequentially follow from the location of the initial point to the location of the destination point until it reaches the destination point, updates it so that the reward for the mobile terminal device 2 to reach the location of the destination point is maximized, and learns the policy of the route that the mobile terminal device 2 should sequentially follow from the current location using a reinforcement learning model. The details of the first learning process will be described later.
[0111] Thereafter, the first storage unit 15 stores the learned reinforcement learning model obtained in step S4 (step S5). Next, the second learning unit 14 learns the relationship between the current position of the mobile terminal device 2 and the policy of the route that the mobile terminal device 2 should sequentially proceed from the current position obtained in the first learning process in step S5, using a supervised learning model (second learning process) (step S6).
[0112] Specifically, the second learning unit 14 repeatedly adjusts and updates parameters such as weights and thresholds so that the error between the predicted output value of the policy of the route to be sequentially advanced when the position of the unit space corresponding to the current GPS position of the mobile terminal device 2, that is, the current state, is given as an input value to the supervised learning model and the teacher data is minimized for the objective function E in the above formula (2), and determines the values of these parameters. In step S6, the second learning unit 14 can determine the parameters that minimize the objective function E by the error backpropagation method or the like.
[0113] The teacher data used in step S6 is the policy of the route that should be sequentially advanced from the current position of the unit space obtained by the learned reinforcement learning model constructed in the first learning process of step S4.
[0114] Next, the second storage unit 16 stores the learned supervised learning model constructed in step S6 (step S7). Thereafter, the setting unit 18 sets the learned supervised learning model in the mobile terminal device 2 as control information (step S8). In step S8, the setting unit 18 can transmit the learned supervised learning model to the mobile terminal device 2 via the network NW. Thereafter, in the mobile terminal device 2, as described later with reference to FIG. 12, movement control is performed based on the control information.
[0115] Next, the first learning process (step S4 in FIG. 9) by the control device 1 will be described using the flowcharts of FIGS. 10 and 11. First, step S3 described in FIG. 9 is executed. First, the first learning unit 13 gives, as an input to the neural network model, the position of the unit space where the mobile terminal device 2 is currently located, which is the current state of the mobile terminal device 2 obtained in step S3, and performs the calculation of the neural network model. As the route to proceed next from the current position of the unit space of the mobile terminal device 2, the first estimated value Q1 of the action value function representing the expected value of the cumulative value of the future reward obtained when each action related to the movement in a predetermined direction with respect to the traveling direction is taken is output (step S20).
[0116] Subsequently, the first acquisition unit 10 acquires the position of the unit space of the mobile terminal device 2 at the next time t as the next state s' (step S21). The position of the unit space that the mobile terminal device 2 has reached next is determined based on the GPS position of the mobile terminal device 2 acquired by the third acquisition unit 12 at each time step. Further, the first learning unit 13 gives, as an input to the neural network model, the position of the unit space that the mobile terminal device 2 has reached next, which is acquired in step S21, performs the calculation of the neural network model, and outputs the second estimated value Q2 of the action value function (step S22).
[0117] Next, the first learning unit 13 calculates a target value from the second estimated value Q2 (step S23). Subsequently, the first learning unit 13 learns the weight parameters of the neural network model so that the first estimated value Q1 becomes the target value calculated from the second estimated value Q2 (step S24). Specifically, the first learning unit 13 updates the weight parameters of the neural network model so as to minimize the loss function of the above formula (1).
[0118] Thereafter, the first storage unit 15 stores the learned reinforcement learning model obtained in step S24 (step S5).
[0119] Next, referring to FIG. 11, the first learning process by the first learning unit 13 when adopting a Fixed Target Q-Network using two neural networks of a main QN131 and a target QN133 will be described.
[0120] The process of step S3 is the same as the steps of the first learning process described in FIG. 10. After that, the first learning unit 13 gives the position of the unit space where the mobile terminal device 2 is currently located, which is acquired in step S3, as an input to the main QN131, performs the operation of the neural network model, outputs the action value function Q, and calculates the route a to proceed next (step S120).
[0121] Next, the first learning unit 13 returns the action of the mobile terminal device 2 to the environment 130 on the route a obtained in step S130, and obtains the position of the unit space where the mobile terminal device 2 has advanced and the reward r, which are the next state s' of the mobile terminal device 2 (step S121).
[0122] The first learning unit 13 stores the experience (s, a, r, a') obtained in step S121 in the experience data 134 (step S122). Next, the first learning unit 13 obtains the loss function L in the DQN loss calculation 132 and updates the weights of the main QN131 with the gradient of the loss function L (step S123). The first learning unit 13 repeats the processes from step S120 to step S123 for a set number of times.
[0123] After that, the first learning unit 13 periodically copies the weights of the main QN131 to the target QN133 for synchronization (step S124). The synchronization of the target QN133 is performed at a frequency lower than the update frequency of the weights of the main QN131. Next, the first learning unit 13 extracts an experience from the experience data 134, inputs the past state to the target QN133, and estimates the value max a’ Q(s’,a’;θ - ) to be output (step S126).
[0124] Next, the first learning unit 13 estimates the value max output by the target QN133 a’Q(s’, a’; θ - ) - based target value r + γmax a’ Q(s’, a’; θ - ) is calculated (step S127). Next, the first learning unit 13 calculates the loss function L in the DQN loss calculation 132 using the target value calculated in step S127 (step S128). Next, the first learning unit 13 performs learning of the weights of the main QN 131 so as to minimize the loss given by the loss function L (step S129). After that, the learned reinforcement learning model is stored in the first storage unit 15 (step S5).
[0125] Next, the operation of the mobile terminal device 2 having the above-described configuration will be described with reference to the flowchart of FIG. 12. Hereinafter, each process after the control information is set in the mobile terminal device 2 in step S8 of FIG. 10 will be described.
[0126] First, the fourth acquisition unit 22 of the mobile terminal device 2 acquires the learned supervised learning model from the fourth storage unit 21 (step S30). The fourth acquisition unit 22 reads out the control information transmitted from the control device 1 and stored in the fourth storage unit 21, that is, the learned supervised learning model.
[0127] Next, the fifth acquisition unit 24 acquires the position of the current unit space of its own device as the current position (step S10). Specifically, the position of the unit space corresponding to the GPS position received by the GPS receiver 207 can be acquired as the current position of its own device. Next, the calculation unit 25 uses the control information acquired in step S30, gives the position of the current unit space of its own device acquired in step S31 as an unknown input, performs the calculation of the learned supervised learning model, and outputs the policy of the route to be sequentially advanced from the current unit space position (step S32).
[0128] For example, in the movement space A of FIG. 2, when the position of the initial point of the mobile terminal device 2 is input to the learned supervised learning model as the current position at time t = 1, n action value functions Q to be advanced next from the initial point at time t = 1 are output.
[0129] Next, the determination unit 26 determines the route to be sequentially followed by selecting the route that takes the action a with the maximum value among the values of the n action value functions Q output in step S32 (step S33). Next, the movement control unit 27 controls the movement of the own device based on the route to be followed next determined in step S33 (step S34). More specifically, the movement control unit 27 can calculate a control command for the route to be followed next from the current position and transmit a control command value to the motor 209.
[0130] The mobile terminal device 2 repeats the processes from step S31 to step S34 until it reaches the unit space where the power supply device 2B is arranged at the destination from the initial point (step S35: NO). After that, when the mobile terminal device 2 reaches the position of the unit space of the destination (step S35: YES), it connects to the power supply device 2B and performs charging (step S36). In this way, the mobile terminal device 2 can autonomously move from the initial point where the remaining amount of the battery 211 has become less than the threshold value to the destination where the power supply device 2B is arranged and perform charging by executing the processes from step S30 to step S36 using the learned supervised learning model which is control information.
[0131] As described above, according to the control device 1 according to the present embodiment, the optimal route strategy of the mobile terminal device 2 from the initial point to the destination where the power supply device 2B is arranged is learned by reinforcement learning, and the route strategy obtained by the reinforcement learning is used as teacher data to learn the relationship between the current position of the mobile terminal device 2 and the route strategy to be sequentially followed using a supervised learning model. Further, the learned supervised learning model is set in the mobile terminal device 2 as control information for controlling the route. Therefore, an optimal movement route can be obtained with a simpler configuration, and efficient power supply to the IoT terminal can be supported.
[0132] Further, according to the control device according to the present embodiment, using the abnormal occurrence information indicating that the remaining battery level has fallen below the threshold value and the position information of the mobile terminal device 2, which are added to the location registration request signal transmitted by the mobile terminal device 2, the optimal route strategy of the mobile terminal device 2 from the initial point to the destination point where the power supply device 2B is arranged is learned by reinforcement learning. Therefore, even when the moving space is wider, an optimal moving route can be obtained with a simpler configuration.
[0133] Further, according to the control system according to the present embodiment, in order to set control information in the mobile terminal device 2, it is possible to realize a mobile terminal device 2 capable of autonomous movement along an optimal route from the initial point where the remaining battery level has fallen below the threshold value to the destination point where the power supply device 2B is arranged.
[0134] [Modification Example] Next, a modification example of the embodiment of the present invention will be described. In the above-described embodiment, the configuration in which control information indicating the optimal route from the initial point where the remaining battery level of the mobile terminal device 2 has fallen below the threshold value to the position of the power supply device 2B fixedly arranged in the moving space is set in the mobile terminal device 2 has been described. That is, in the above-described embodiment, the case where the first device that moves is the mobile terminal device 2 and the second device fixedly arranged is the power supply device 2B has been described.
[0135] On the other hand, in this modification example, as shown in FIG. 13, in response to the remaining battery level of the terminal device 2' falling below the threshold value, control information indicating the optimal route to the terminal device 2' fixedly arranged in the moving space A is set in the mobile power supply device 2B'. Therefore, in this modification example, the first device that moves is the mobile power supply device 2B', and the second device fixedly arranged is the terminal device 2'.
[0136] The configuration of the control device 1 of this modified example and the configuration of the control system are the same as those of the control device 1 and the control system described in FIG. 1. Also, the mobile power supply device 2B' included in the control system according to this modified example has a functional block similar to the functional block of the mobile terminal device 2 according to this embodiment described in FIG. 7. In this modified example, the transmission unit 20 of the mobile power supply device 2B' adds its own position information to the location registration request signal and transmits it at a fixed period, but the abnormal occurrence information is transmitted by the terminal device 2'.
[0137] The hardware configuration of the mobile power supply device 2B' has the same configuration as the hardware configuration of the mobile terminal device 2 according to this embodiment described in FIG. 8. In this modified example, the battery 211 includes an internal battery and a charging battery for charging the terminal device 2'. Also, the power management module 212 manages and controls the charging process.
[0138] The terminal device 2' is fixedly arranged in the mobile space and includes the transmission unit 20 among the functional blocks of the mobile terminal device 2 described in FIG. 7. Also, the hardware configuration of the terminal device 2' has a configuration other than the motor 209 and the drive mechanism 210 among the configurations of the mobile terminal device 2 described in FIG. 8.
[0139] [Operation of the control system] Next, the operation sequence of the control system according to this modified example will be described with reference to the sequence diagram of FIG. 14. First, the terminal device 2' is fixedly arranged at a predetermined position in the mobile space and monitors the remaining power of its own device. The mobile power supply device 2B' is arranged to be movable to any position within the mobile space.
[0140] First, the terminal device 2' transmits a location registration request signal with its own GPS position added to the UDM / UDR 41 of the core network 4 via the base station 3 (step S110). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS position in association with the IMSI of the terminal device 2'. Thereafter, the first acquisition unit 10 of the control device 1 acquires the GPS position of the terminal device 2' from the UDM / UDR 41 as the destination (step S1A).
[0141] Thereafter, when it is detected that the remaining battery level of the terminal device 2' has fallen below the threshold value, a location registration request signal with the abnormal occurrence information and the GPS position added is transmitted to the core network 4 (step S111). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal, the abnormal occurrence information, and the GPS position in association with the IMSI of the terminal device 2'. Thereafter, the second acquisition unit 11 of the control device 1 acquires the abnormal occurrence information from the UDM / UDR 41 (step S2A). Triggered by the acquisition of the abnormal occurrence information in step S2A, the processes from step S3A to step S8A below are executed.
[0142] Next, the mobile power supply device 2B' transmits a location registration request signal with its own GPS position added to the core network 4 (step S112). The UDM / UDR 41 stores the transmission timestamp of the location registration request signal and the GPS position in association with the IMSI of the mobile power supply device 2B'. Subsequently, the third acquisition unit 12 of the control device 1 acquires the GPS position of the mobile power supply device 2B' from the UDM / UDR 41 as the current position (step S3A).
[0143] Next, the first learning unit 13 performs first learning processing (step S4). In the first learning processing, the first learning unit 13 applies a reward function to the estimated result of calculating the route that the mobile power supply device 2B' should sequentially follow until it reaches the position of the destination point from the position of the initial point of the mobile power supply device 2B', and updates it so that the reward for the mobile power supply device 2B' to reach the position of the destination point is maximized. The policy of the route that the mobile power supply device 2B' should sequentially follow from the current position is learned using a reinforcement learning model.
[0144] Thereafter, the first storage unit 15 stores the learned reinforcement learning model obtained in step S4 (step S5). Next, the second learning unit 14 learns the relationship between the current position of the mobile power supply device 2B' and the policy of the route that the mobile power supply device 2B' should sequentially follow from the current position obtained in the first learning processing in step S5 using a supervised learning model (second learning processing) (step S6).
[0145] Next, the second storage unit 16 stores the learned supervised learning model constructed in step S6 (step S7). Thereafter, the setting unit 18 sets the learned supervised learning model in the mobile power supply device 2B' as control information (step S8A). In step S8A, the setting unit 18 can transmit the learned supervised learning model to the mobile power supply device 2B' via the network NW. Thereafter, in the mobile power supply device 2B', movement control is performed based on the control information.
[0146] The movement control performed by the mobile power supply device 2B' based on the control information corresponds to the processing from step S30 to step S35 of the mobile terminal device 2 in the present embodiment described with reference to FIG. 12. When the mobile power supply device 2B' reaches the position of the terminal device 2' which is the destination point, it connects to the terminal device 2' and charges the terminal device 2'.
[0147] As described above, according to the modification example of the present embodiment, when the remaining power of the terminal device 2' becomes less than the threshold value, the optimal route policy of the mobile power supply device 2B' from the initial point, which is the position of the mobile power supply device 2B' at that time, to the destination point where the terminal device 2' is arranged is learned by reinforcement learning. Using the route policy obtained by reinforcement learning as teacher data, the relationship between the current position of the mobile power supply device 2B' and the route policy to be sequentially advanced is learned using a supervised learning model. Furthermore, the learned supervised learning model is set in the mobile power supply device 2B' as control information for controlling the route. Therefore, an optimal movement route can be obtained with a simpler configuration, and efficient power supply to the IoT terminal can be supported.
[0148] In the described embodiment, the reinforcement learning model used by the first learning unit 13 is exemplified by the DQN related to the Fixed Target Q-Network configured by a multi-layer neural network. However, as the reinforcement learning model, a CNN, a multi-layer perceptron, etc. can be used. In addition to the DQN exemplified as the reinforcement learning model, Double DQN, Dueling DQN, Actor-Critic (AC) method, Soft Actor-Critic (SAC), Deep Deterministic Policy Gradient (DDPG), Q-learning, etc. can be used.
[0149] Also, in the described embodiment, the supervised learning model used by the second learning unit 14 is exemplified by the case of using a multi-layer neural network. However, as the supervised learning model, a multi-layer perceptron, a decision tree-based model such as a random forest, a support vector machine, etc. can be used.
[0150] The embodiments of the control device, control method, and control system of the present invention have been described above. However, the present invention is not limited to the described embodiments, and various modifications that those skilled in the art can assume within the scope of the invention described in the claims are possible.
Explanation of Reference Numerals
[0151] 1... Control device, 10... First acquisition unit, 11... Second acquisition unit, 12... Third acquisition unit, 13... First learning unit, 14... Second learning unit, 15... First memory unit, 16... Second memory unit, 17... Third memory unit, 18... Setting unit, 2... Mobile terminal device, 2B... Power supply device, 20... Transmitter, 21... Fourth memory unit, 22... Fourth acquisition unit, 23... Fifth memory unit, 24... Fifth acquisition unit, 25... Arithmetic unit, 26... Decision unit, 27... Movement control unit, 101, 201... Bus, 102, 202... Processor, 103, 203... Main memory device, 104, 204... Communication interface, 105, 205... Auxiliary memory device, 106, 206... Input / output I / O, 107... Display device, 207... GPS receiver, 208... Sensor, 209... Motor, 210... Driving mechanism, 211... Battery, 212... Power management module, 130... Environment, 131... Main QN, 132... DQN loss calculation, 133... Target QN, 134... Experience data, NW... Network.
Claims
1. A control device that controls a route of a first device that moves to a destination point set in a moving space, a first acquisition unit configured to acquire, via a core network, abnormality occurrence information related to a remaining power supply of the first device or a second device disposed in the moving space, the abnormality occurrence information being added to a first location registration request signal transmitted at a constant cycle by the first device or a second location registration request signal transmitted at a constant cycle by the second device; a second acquisition unit configured to acquire a location of the second device as the location of the destination point via the core network; a third acquisition unit configured to acquire, via the core network, a current location of the first device added to the first location registration request signal transmitted by the first device at a certain period or when the first device is powered on; a first learning unit configured to apply a reward function to an estimation result of calculating a course to be taken by the first device from an initial position of the first device until the first device reaches the destination position, and to update the course so as to maximize a reward for the first device to reach the destination position, and to learn a course plan for the first device to take from the current position using a reinforcement learning model; A second learning unit configured to learn, using a supervised learning model, a relationship between the current position of the first device and a course plan for the first device to follow from the current position until the first device reaches the destination point, the course plan being obtained by learning by the first learning unit; A storage unit configured to store the trained supervised learning model constructed by the second learning unit; A control device comprising:
2. 2. The control device according to claim 1, The second acquisition unit acquires, as the location of the destination point, the location of the second device added to the second location registration request signal transmitted by the second device at a certain period or when the second device is powered on. A control device comprising:
3. 2. The control device according to claim 1, When the remaining power of the first device becomes less than a threshold value, the first device adds the abnormality occurrence information to the first location registration request signal that is transmitted at a constant period and transmits the signal; the second device is a power supply device that supplies power to the first device, The third acquisition unit acquires the current location of the first device when the first acquisition unit acquires the abnormality occurrence information. A control device comprising:
4. 2. The control device according to claim 1, the first device is a power supply device that supplies power to the second device, When the remaining power of the second device becomes less than a threshold value, the second device adds the abnormality occurrence information to the second location registration request signal that is transmitted at a constant period and transmits the signal; The third acquisition unit acquires the current location of the first device when the first acquisition unit acquires the abnormality occurrence information. A control device comprising:
5. 2. The control device according to claim 1, The method further includes a setting unit configured to set the trained supervised learning model in the first device as control information for controlling a route from the position of the initial point to the position of the destination point. Control device.
6. A method for controlling a route of a first device moving to a destination point set in a moving space, comprising: a first acquisition step of acquiring, via a core network, abnormality occurrence information related to the remaining power supply of the first device or the second device arranged in the moving space, the abnormality occurrence information being added to a first location registration request signal transmitted at a constant cycle by the first device or a second location registration request signal transmitted at a constant cycle by the second device; a second acquisition step of acquiring a location of the second device as the location of the destination point via the core network; a third acquisition step of acquiring, via the core network, a current location of the first device added to the first location registration request signal transmitted by the first device at a certain period or when the first device is powered on; a first learning step of applying a reward function to an estimation result of calculating a course that the first device should take from an initial position of the first device to the destination position, updating the course so as to maximize a reward for the first device to reach the destination position, and learning a course plan for the first device to take from the current position using a reinforcement learning model; a second learning step of learning, using a supervised learning model, a relationship between the current position of the first device and a course plan for the first device to follow from the current position until the first device reaches the destination point, the course plan being obtained by learning in the first learning step; a storage step of storing the trained supervised learning model constructed in the second learning step in a storage unit; A control method comprising:
7. In the control method according to claim 6, the second acquisition step acquires the position of the second device added to the second position registration request signal transmitted by the second device at a fixed period or when power is turned on, as the position of the destination point. A control method characterized by this.
8. In the control method according to claim 6, when the remaining power of the first device becomes less than a threshold value, the first device adds the abnormality occurrence information to the first position registration request signal transmitted at a fixed period and transmits it, the second device is a power supply device that supplies power to the first device, the third acquisition step acquires the current position of the first device, triggered by the acquisition of the abnormality occurrence information in the first acquisition step. A control method characterized by this.
9. In the control method according to claim 6, the first device is a power supply device that supplies power to the second device, when the remaining power of the second device becomes less than a threshold value, the second device adds the abnormality occurrence information to the second position registration request signal transmitted at a fixed period and transmits it, the third acquisition step acquires the current position of the first device, triggered by the acquisition of the abnormality occurrence information in the first acquisition step. A control method characterized by this.
10. In the control method according to claim 6, further, a setting step of setting the learned supervised learning model in the first device as control information for controlling the route from the position of the initial point to the position of the destination point by the first device is provided. A control method.
11. In the control method according to claim 6, further, the first device has a fourth acquisition step of acquiring the learned supervised learning model constructed in the second learning step, a fifth acquisition step of acquiring the current position of the device itself, gives the current position of the device itself acquired in the fifth acquisition step as an unknown input to the learned supervised learning model, performs the calculation of the learned supervised learning model, and outputs a strategy for the route to be sequentially advanced from the current position of the device itself. An operation step, and a movement control step of controlling the movement of the device itself from the initial point to the destination point based on the strategy for the route to be sequentially advanced from the current position of the device itself output in the operation step. A control method comprising this.
12. The control device according to any one of claims 1 to 5, The first device and A control system comprising: The first device is configured to: A fourth acquisition unit configured to acquire the learned supervised learning model constructed by the control device; A fifth acquisition unit configured to acquire the current position of the own device; An operation unit configured to provide the current position of the own device acquired by the fifth acquisition unit as an unknown input to the learned supervised learning model, perform the operation of the learned supervised learning model, and output a policy for a route to be sequentially advanced from the current position of the own device; A movement control unit configured to control the movement of the own device from the initial point to the destination point based on the policy for the route to be sequentially advanced from the current position of the own device output by the operation unit A control system comprising.
Citation Information
Patent Citations
Charging control system for mobile robot
JP1991284103A
Travel route teaching system and method for autonomous mobile body
JP2016206876A
Radio communication system between mobiles
JP2023045195A
Route management device, route management method, and route management system
JP7572588B1
Traveling vehicle system
WO2024070621A1