Electromagnetic space and geospatial joint navigation optimization method for single mobile user

By optimizing electromagnetic and geospatial navigation through a deep neural network model, the problem of communication interruption for mobile users during navigation is solved. This enables rapid arrival at the destination in directions with high electromagnetic intensity while maintaining communication quality and reducing energy consumption.

CN120403594BActive Publication Date: 2025-10-24SICHUAN HUATENG FUTURE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510433086.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-10-24
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

During navigation for mobile users, existing technologies struggle to simultaneously guarantee communication quality and rapid arrival at the destination, especially in joint navigation in electromagnetic and geographic spaces, where weak signal areas can lead to communication interruptions.

Method used

By employing a deep neural network model that combines electromagnetic and geospatial information, and through prediction network blocks, fusion units, observation encoders, and Q-network modules, joint navigation optimization based on electromagnetic and geospatial information is achieved. Deep reinforcement learning and supervised learning algorithms are used to guide mobile users to move in directions with high electromagnetic intensity.

Benefits of technology

In complex environments, joint navigation optimization enables mobile users to reach their destinations faster and maintain stable communication quality, while reducing energy consumption and route planning complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120403594B_ABST
    Figure CN120403594B_ABST
Patent Text Reader

Abstract

The application provides a kind of electromagnetic space and geographical space joint navigation optimization method for single mobile user, receives the current state and terminal point information of mobile user agent, current state includes the current position of agent, path obstacle, measured electromagnetic information and walkable path;Current state containing measured electromagnetic information is input to deep neural network, so that the output predicted moving direction tends to move in the direction of sufficient electromagnetic intensity to maintain communication.The application not only considers the navigation of geographical space, but also further considers the joint navigation of electromagnetic and geographical space, so that the mobile user in space can reach the destination faster while ensuring the communication quality.The application combines the characteristics of reinforcement learning during the training of the navigation model based on deep network, compares the real electromagnetic map after the movement of the agent with the predicted map for supervised learning, so that deep reinforcement learning and machine learning are organically combined together.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to navigation technology for mobile users, in particular to electromagnetic space and geographical space joint navigation technology. BACKGROUND

[0002] Wireless networks have been widely deployed and become very large in scale due to the rapid development of wireless communication technology. The main function of wireless networks is to provide stable and reliable services for mobile users. In the scenario of simultaneous communication and movement, the communication behavior and movement behavior of users need to be jointly planned in electromagnetic space and geographical space. Considering electromagnetic space, if the user always stays in an area with strong signal strength during movement, it will help to obtain more stable network connection and reduce the risk of signal loss and communication interruption; considering geographical space, it is necessary to ensure that the user can avoid obstacles in the environment and find the shortest path as the optimal path during path planning.

[0003] The concurrent interweaving of movement and communication complicates the joint navigation problem. For example, in addition to geographical obstacles, users should also avoid areas with weak signals, because in these areas, the communication quality of users will be severely degraded. Therefore, we should consider both geographical maps and electromagnetic maps when navigating, otherwise, if only path planning is considered, communication interruption of mobile users will occur when encountering weak signal areas. On the other hand, if only moving in strong signal areas is considered without considering path planning, it is difficult for mobile users to quickly reach the destination. SUMMARY

[0004] The technical problem to be solved by the present application is to provide a navigation scheme that can enable mobile users in space to reach the destination faster while ensuring communication quality.

[0005] The technical solution adopted by the present application to solve the above technical problem is to provide an electromagnetic space and geographical space joint navigation optimization method for a single mobile user, receiving current state and end point information of a mobile user agent, the current state including the current position of the agent, path obstacles, measured electromagnetic information and walkable paths; inputting the current state containing the measured electromagnetic information into a deep neural network, so that the output predicted movement direction tends to move in the direction with sufficient electromagnetic intensity to maintain communication; the deep neural network performs supervised learning by comparing the measured electromagnetic information after the agent performs an action with the predicted electromagnetic information during the training process;

[0006] The prediction process of the deep neural network includes the following steps:

[0007] Output the predicted electromagnetic information of the next state according to the measured electromagnetic information in the current state;

[0008] Fuse the predicted electromagnetic information of the next state with the current state to obtain electromagnetic space and geographic space joint information;

[0009] According to the electromagnetic space and geographic space joint information, output joint encoding information;

[0010] The joint encoding information is input into the Q network module, and the action corresponding to the maximum value of the Q function output by the Q network module is taken as the predicted moving direction.

[0011] The application also provides an electromagnetic space and geographic space joint navigation model for a single mobile user based on a deep neural network, which realizes the above steps through a prediction network block, a fusion unit, an observation encoder and a Q network module therein.

[0012] The Q network module adopts the architecture of a Dueling DQN duel network; the Q value calculation is divided into two branches of state value and advantage function, which are calculated through two independent fully connected layers FC, and the results of the two fully connected layers FC are combined to output the Q value.

[0013] The training of the deep neural network adopts a double network structure, and a target deep neural network with the same structure as the deep neural network but different parameter updates is introduced.

[0014] The application not only considers geographic space navigation, but also further considers electromagnetic and geographic space joint navigation, so that the mobile user in space can reach the destination faster while ensuring the communication quality. The electromagnetic and geographic space joint navigation problem is modeled as a challenging non-convex problem, and is solved through deep reinforcement learning and supervised learning algorithms.

[0015] The application has the beneficial effects that according to the characteristics of the electromagnetic map (electromagnetic information), a prediction network block is added, so that the agent can predict the change of electromagnetic intensity, and the agent can move in the direction of high electromagnetic intensity on the premise of finding the destination. When training the navigation model based on the deep network, the characteristics of reinforcement learning are combined, the real electromagnetic map after the agent moves is compared with the predicted map for supervised learning, so that deep reinforcement learning and machine learning are organically combined together. Compared with the existing method which only considers geographic space navigation, the user can consume less energy in the electromagnetic and geographic joint space, and can quickly plan a path for the user in a complex and changing environment. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The navigation model based on the deep network used in the application. DETAILED DESCRIPTION

[0017] Consider a wireless communication network where a base station is trying to serve a user and there are two interference sources in the network. We assume that the network area is a square, defined as [x L ,x U ]×[y L ,y U ], where the subscripts L and U represent the lower and upper bounds of the square’s two-dimensional coordinates, respectively. M represents the total number of cells in the network area, and the cell side length is Δs, n=1,..., N represents the time step. Since the mobile user moves one cell in each time step, we define b n ∈{m=1,...,M} is the cell where the user moves in time step n, and m is the cell number. n represents the location of the mobile user at time step n, Represents the trajectory of mobile users.

[0018]

[0019] in is the initial position, x I ,y I is the two-dimensional coordinate of the initial position, is the end position. Indicates the lower boundary of the region and Indicates the upper right boundary of the region, ≤ is the element-by-element inequality sign, Indicates any.

[0020] set up Represents the time step n, from the base station to the unit b n The channel loss, g m represents the channel loss from the interference source in cell m to the base station. The channel loss depends on the large-scale path loss and shadowing of the local environment as well as small-scale fading. Considering the uplink, if the signal to interference plus noise ratio (SINR) transmitted by the mobile user is lower than the threshold γ th , that is, when SINR n <γ th When , the mobile user is considered to be interrupted at time step n:

[0021]

[0022] in Indicates that at time step n, the user arrives at unit b n The power transmitted by the user, |σ n | 2 Represents the noise interference at time step n, usually white noise. m is the power emitted by the interference source in cell m. We define the communication energy consumption as E c ,but

[0023]

[0024] In the process of mobile user movement, we not only need to minimize the task communication energy consumption E c , but also minimize the weighted distance sum of the user's current location to the destination at each time, so as to move towards the destination as much as possible. The navigation modeling optimization problem P0 is as follows:

[0025]

[0026] p bn ≤P max (4c)

[0027] q0=q I ,q N =q F (4d)

[0028]

[0029] where γ is the discount factor, P max is the maximum transmit power of the user, |||| represents the 2-norm, Δs is the cell side length, V n represents the feasible movement direction set of the mobile user, V n ={1,-1,i,-i}, 1 and -1 represent moving right and left respectively, and i and -i represent moving up and down respectively; represents the feasible movement direction of the mobile user at time step n. is a geographical map, which takes values of 0 or 1. 0 represents passable, and 1 represents existence of obstacles. satisfy cond. represents that the conditions explicitly written in the subsequent cond need to be satisfied. cond: represents condition definition.

[0030] If P0 has an optimal solution, then the QoS constraint condition in the above formula (4b) must satisfy the equation at the optimal solution. Therefore, P0 can be rewritten as P1.

[0031]

[0032] Solving the navigation modeling optimization problem P1:

[0033] P1 is solved by reinforcement learning. First, P1 is regarded as a Markov decision process (MDP) with partial observability. MDP is a mathematical framework for modeling sequential decision problems, which describes a process in which an agent interacts with the environment through actions to maximize cumulative rewards in an environment with Markov properties. The MDP of the embodiment is represented by a four-tuple

[0034] State set Each element of the state set is a state The information within the agent's viewing range, including the agent's current position b n , path obstacles, measuring electromagnetic information and traversable paths.

[0035] The current location of the agent can be displayed using a binary local map. In addition to displaying the current agent location, the binary local map can also display the locations of other agents within the surrounding radius.

[0036] Measured electromagnetic information is used to display spectrum information within the visible range, which can be a continuous normalized electromagnetic map. Path obstacles are used to display geographic obstacle information within the visible range, which can be a binary obstacle map. Traversable paths include information about already traveled paths and the probability of reaching the destination in four different directions of movement. Specifically, a heuristic algorithm can be used to mark the probability of reaching the destination in each of the four directions of movement.

[0037] During actual training, we observed that the agent was swaying severely, so we also provided the path it had walked within the view. This is to give a negative reward, or penalty, if the agent is judged to have walked past it again.

[0038] Action Collection The action space corresponds to the user's movement direction and is divided into four directions: front, back, left, and right.

[0039] Among them, 1 and -1 represent movement to the right and left respectively, and i and -i represent movement up and down respectively.

[0040] State transition probability The state transition matrix is ​​deterministic because the next position is directly determined by the agent's current position and the currently selected action.

[0041] Reward Collection Each reward r in the reward set is determined based on the state of the agent after performing the action:

[0042] 1) If the agent performs an action After that, it can move, then give a negative reward value R m (q n );

[0043] 2) If the agent performs an action If the communication is normal, a negative reward value R is given d (b n );

[0044] 3) If the agent performs an action If an obstacle is encountered, a negative reward value R is given. o ;

[0045] 4) If the agent performs action after the communication interruption, a negative reward value R n is given.

[0046] 5) If the agent performs action after which a repeated path is walked, a negative reward value R c is given.

[0047] 6) If the agent performs action after which the end point is reached, a positive reward value R f is given.

[0048] In order to let the agent reach the destination with fewer steps during training, the path planning problem is only given a positive reward when the destination is reached, and a negative reward is given in other cases. In order to ensure that the path of the agent is movable and the communication is normal, |R m (q n )|<|R o |, |R d (b n )|<|R n |. The proportion between movable and normal communication can be adjusted according to actual needs. In this embodiment, it is considered that the movable and the electromagnetic intensity sufficient to maintain the communication are equally important, and therefore R o =R n . After the agent performs action , it does not reach the end point, and is in a repeated path that can be moved and the communication is normal, then r=R m (q n )+R d (b n )+R c .

[0049] Specifically, the reward value R m (q n )= -γ||q n -q F || 2 .

[0050] Specifically, the method for judging whether the communication is interrupted is: according to (5b) to get the power n of the agent in unit b , if the power is less than P , it is considered that the communication is interrupted;

[0051] Specifically, the reward value Specifically, the size of the reward value R c for walking a repeated path can be set to the reward value R m (q n ) of the movable.) and the reward value R for encountering an obstacle o Between |R m (q n )|<|R c |<|R o |, and is proportional to the number of times the agent has passed the current point.

[0052] By building and training a deep neural network, a joint navigation system of electromagnetic space and geospatial space for a single mobile user is constructed, including a prediction network block, a fusion unit, an observation encoder and a Q network module. Figure 1 shown.

[0053] The prediction network block is used to output the predicted electromagnetic information p of the next state based on the current measured electromagnetic information d.

[0054] The prediction network block used in the embodiment includes two independent convolutional layers, the 1st convolutional layer and the 4th convolutional layer, and a residual block. The 1st convolutional layer, the residual block and the 4th convolutional layer are connected once. The input of the residual block is the output of the 1st convolutional layer, which is added to the original input of the residual block after the output of the 2nd and 3rd convolutional layers. The addition result is then activated by ReLU to form the output of the residual block. The 1st, 2nd and 3rd convolutional layers are all structures in which a 3*3 convolution kernel, a batch normalization layer BN and a ReLU activation are connected in sequence. The 4th convolutional layer is a structure in which a 3*3 convolution kernel, a batch normalization layer BN and a Sigmod activation are connected in sequence.

[0055] The electromagnetic information d is measured and predicted by the prediction network block to make four actions in the current time step, namely up, down, left and right. The predicted electromagnetic information p of the next state corresponding to the next state is obtained, so that the intelligent agent can know the electromagnetic map at the next moment and have the ability to speculate, making it easier to move in the direction of higher electromagnetic intensity that is sufficient to maintain communication.

[0056] The fusion unit is used to fuse the predicted electromagnetic information p of the next state with the current state s and output the combined electromagnetic space and geographic space information c. The fusion unit of the embodiment is an adder. The current state is the information within the visual range of the intelligent agent, including the current position b of the intelligent agent. n , path obstacles, measuring electromagnetic information d and traversable paths.

[0057] The observation encoder is used to output joint coding information e according to the joint information c of electromagnetic space and geographic space.

[0058] The observation encoder adopted in the embodiment includes two independent convolutional layers, the 1st convolutional layer and the 8th convolutional layer, and three residual blocks. The input of the first residual block is the output of the 1st convolutional layer. The input of the second residual block is the output of the first residual block. The input of the third residual block is the output of the first residual block. Each residual block includes 2 convolutional layers. The first residual block includes the 2nd convolutional layer and the 3rd convolutional layer. The second residual block includes the 4th convolutional layer and the 5th convolutional layer. The third residual block includes the 6th convolutional layer and the 7th convolutional layer. In each residual block, the output of the input through the two convolutional layers is added to the original input of the residual block, and the addition result is then activated by ReLU to form the output of the residual block. The 8 convolutional layers are all structures with a 3*3 convolution kernel, a batch normalization layer BN and a ReLU activation connected in sequence.

[0059] The Q network module is used to output the q value based on the joint coding information e to predict the user's movement direction The Q network module is a neural network structure used to approximate the Q function in deep reinforcement learning. The embodiment adopts the Dueling DQN dueling network architecture, which decomposes the Q value into state value and advantage function, calculates it through two independent FC fully connected layers, and finally merges the output Q value. The optimal action is selected by the maximum value q

[0060] To stabilize the training process and avoid oscillation and divergence in Q-value estimation, a dual-network architecture is used for deep neural network training. This involves introducing a target deep neural network with the same structure as the deep neural network, but with asynchronous parameter updates. The deep neural network used after training is called an online deep neural network. The model parameters of the target deep neural network are periodically copied or delayed from the online deep neural network to calculate the target Q-value.

[0061] The training process of a dual network structure is given as follows:

[0062] Step 1: Input the state s into the online deep neural network with initial weights

[0063] First, the historical experience of the interaction between the planned agent and the environment is collected and stored in the experience playback memory. The historical experience is a four-tuple information s is the current state, is the execution action, r is the corresponding action Rewards, s ′ To perform an action The next state after

[0064] The acquisition process of the historical experience is as follows: the agent has the probability ∈ to randomly select an action from the action set There is a 1-∈ probability according to Select the best action Where ω is the weight of the target deep neural network. The target deep neural network outputs the maximum value q of the Q function based on the input s and d in s, thereby selecting the optimal execution action of the agent Agent-to-action Evaluate to get reward r, then collect and execute actions The next state s obtained ′ , continuously collect historical experience of the interaction between the planned intelligent agent and the environment, and store it in the experience replay memory to obtain the training sample set D.

[0065] Step 2: Randomly sample quadruple information to update the deep neural network

[0066] First, randomly sample quadruple information from the training sample set D in the experience replay memory Input into the deep neural network and the target deep neural network and do the following processing:

[0067] Initially, the parameters of the deep neural network and the target deep neural network are the same;

[0068] Substitute the state s into the deep neural network for feedforward operation to obtain the predicted estimated Q value corresponding to all possible actions

[0069] The state s corresponds to the state s in its four-tuple information ′ Substitute it into the target deep neural network for feedforward operation to calculate the maximum value of the Q function output by the network in For state s ′ The action corresponding to the maximum Q value after substituting into the target deep neural network, ω - is the weight of the observation encoder and Q network module; get the target value Where γ is the discount factor.

[0070] The loss function is constructed based on the output of the deep neural network and the target deep neural network:

[0071] Combine electromagnetic information p with the action performed by Agent Electromagnetic information after p The comparison is supervised learning, and the loss function is: where ω p are the weights of the prediction network block.

[0072] The random gradient descent method is applied to update the weights of the deep neural network and the target deep neural network, wherein the weights in the deep neural network are updated in real time, and the weights in the target deep neural network are updated every set time step; and when the selected generation steps are reached, the trained deep neural network is obtained.

Claims

1. A method for joint electromagnetic and geographical space navigation optimization for a single mobile user, characterized in that, The current state of the mobile user agent and the terminal point information are received, the current state including the current position of the agent, path obstacles, measured electromagnetic information, and walkable paths; the current state containing the measured electromagnetic information is input into a deep neural network, so that the output predicted moving direction tends to move in a direction with sufficient electromagnetic intensity to maintain communication; the deep neural network compares the measured electromagnetic information after the agent performs an action with the predicted electromagnetic information during the training process to perform supervised learning; the prediction process of the deep neural network includes the following steps: Output the predicted electromagnetic information of the next state according to the measured electromagnetic information in the current state; Fuse the predicted electromagnetic information of the next state with the current state to obtain electromagnetic space and geographic space joint information; Output joint encoding information according to the electromagnetic space and geographic space joint information; Input the joint encoding information into a Q network module, and take the action corresponding to the maximum value of the Q function output by the Q network module as the predicted moving direction.

2. The method of claim 1, wherein, The current position of the agent uses a binary local map display; The measured electromagnetic information is used to show the spectrum information within the visual range, which is a continuous normalized electromagnetic map; The path obstacles are used to show the geographic obstacle information within the visual range, which is a binary obstacle map; The walkable path includes the path information that has been walked through and the possibility of reaching the terminal point in the four moving directions of up, down, left, and right.

3. The method of claim 1, wherein, During the training process of the deep neural network, the reward r is obtained by evaluating the action performed by the agent, and the training process is constrained as much as possible to maximize the reward r; The reward r is determined according to the specific situation after the agent performs an action, and the rules are as follows: 1) The agent performs an action If the agent can move, then a negative reward value R is given m (q n ); 2) the agent performs an action If the post-communication is normal, a negative reward value R is given d (b n ); 3) the agent performs an action Upon encountering an obstacle, a negative reward value R is given o ; 4) the agent performs an action After the communication is interrupted, a negative reward value R is given n ; 5) the agent performs an action If the agent walks the same path again, it is given a negative reward value R c ; 6) the agent performs an action A positive reward value R is given if the agent reaches the goal last f .

4. The method of claim 3, wherein, The reward r rule satisfies: |R(b n )| < |R n |; |R m (q n )| < |R c | < |R o |; R o = R n .

5. The method of claim 3, wherein, The method to determine whether the communication is interrupted is to get the power of the transmission of the agent at the current location When The communication is considered interrupted when The communication is considered normal; P max is the maximum transmission power for the user, γ th is a preset threshold, M represents the total number of cells in the network area, m is the cell serial number, b n is the cell where the mobile user is located at time step n, represents the channel loss from the base station to cell b n at time step n, g m represents the channel loss from the interference source in cell m to the base station, P m is the transmission power of the interference source in cell m, |σ n | 2 represents the noise interference at time step n.

6. The method of claim 5, wherein, reward value R m (q n ) = -γ||q n -q F || 2 ; reward value where γ is a discount factor, q n denotes the position where the mobile user is at time step n, q F denotes the terminal position.

7. Electromagnetic space and geospatial joint navigation model for a single mobile user, characterized in that, The model is based on a deep neural network, receives the current state of the mobile user agent and the terminal point information, and the current state includes the current position of the agent, path obstacles, measured electromagnetic information, and walkable paths; the current state containing the measured electromagnetic information is input into a deep neural network, so that the output predicted moving direction tends to move in a direction with sufficient electromagnetic intensity to maintain communication; the deep neural network compares the measured electromagnetic information after the agent performs an action with the predicted electromagnetic information during the training process to perform supervised learning; The deep neural network includes a prediction network block, a fusion unit, an observation encoder, and a Q network module; The prediction network block is used to receive the measured electromagnetic information in the input current state and output the predicted electromagnetic information of the next state; The fusion unit is used to fuse the predicted electromagnetic information of the next state with the current state, and outputs the electromagnetic space and geographic space joint information; The observation encoder is used to receive the input electromagnetic space and geographic space joint information, and outputs the joint encoding information; The Q network module is used to receive the input joint encoding information, and outputs the action corresponding to the maximum value of the Q function, which is taken as the predicted moving direction.

8. The navigation model of claim 7, wherein, The Q network module adopts the architecture of the Dueling DQN duel network; the Q value calculation is divided into two branches of state value and advantage function, which are calculated through two independent fully connected layers FC, and the results of the two fully connected layers FC are combined to output the Q value.

9. The navigation model of claim 8, wherein, The training of the deep neural network adopts a double network structure, and a target deep neural network which has the same structure as the deep neural network but different parameter updating is introduced.

10. The navigation model of claim 9, wherein, In the training process of the deep neural network, the reward r is obtained by evaluating the action performed by the agent, and the training process is constrained as much as possible to maximize the reward r. The reward r is determined according to the specific situation after the agent performs the action, and the rules are as follows: 1) The agent performs an action If the agent can move, then a negative reward value R is given m (q n ); 2) the agent performs an action If the communication is normal, a negative reward value R is given d (b n ); 3) the agent performs an action Upon encountering an obstacle, a negative reward value R is given o ; 4) the agent performs an action After the communication is interrupted, a negative reward value R is given n ; 5) the agent performs an action If the agent walks the same path again, it is given a negative reward value R c ; 6) the agent performs an action Upon reaching the goal, a positive reward value R is given f .

Citation Information

Patent Citations

  • Electromagnetic map reconstruction system and method based on graph structure data

    CN118470228A

  • IT9120110686A1