Communication navigation perception depth enhancement space-time collaborative networking perception optimization method and device
In the absence of air self-organized communication network nodes, the deep reinforcement learning model is used to optimize the space-time and space-time deployment of ground communication, perception and data reception nodes, and the challenges of complex geographical environments to ground high-frequency wireless self-organized network communication are solved, and the communication speed and deployment location are significantly improved, and the communication guarantee capabilities in disaster emergency are improved.
Patent Information
- Application Number
- CN202510050394.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Complex geographical environments present challenges to ground high-frequency wireless ad hoc network communication, especially in the absence of aerial ad hoc communication network nodes, how to achieve effective communication rate and deployment location optimization.
By establishing a digital surface model of the affected area and a communication fitting model of the ad hoc network, a perception node, a communication node and a data receiving node are set up, and a deep reinforcement learning model based on ResNet-DDPG is adopted to optimize the space-time deployment of communication, perception and data receiving nodes on the ground.
It significantly improves the effectiveness of communication speed and deployment location, and improves communication support and command and dispatch capabilities in disaster emergency response, which is of great significance especially in harsh climate conditions.
Smart Images

Figure CN119997030A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information science, and more specifically, to a method and device for optimizing communication navigation perception depth enhancement and spatiotemporal collaborative networking perception. Background Art
[0002] Floods, earthquakes and typhoons not only cause serious damage to crops and infrastructure, but also cause a large number of casualties. Faced with these increasingly serious challenges, humans must increase their efforts in disaster prevention and mitigation to ensure the safety of life and property and the sustainable development of society. However, extreme natural disasters often lead to the paralysis of communication infrastructure, making it impossible for emergency personnel to communicate in real time, affecting the real-time transmission of multimodal sensing data such as location, video, photos and laser, which in turn hinders the effective dispatch of emergency command and seriously affects the efficiency of rescue response. Self-organizing networks (ANETs) have become an important means of disaster emergency communications because they do not rely on pre-built infrastructure and can quickly deploy emergency communication networks. However, how to effectively deploy ANETs to provide high-quality and stable communications is still a problem that needs to be studied in depth.
[0003] In recent years, some emergency communication solutions have been proposed to address the above-mentioned problems. The key technical path of the integrated communication, navigation, and remote sensing space-based information real-time service system has laid a theoretical and technical foundation for emergency response in ad hoc networks. The integrated communication, navigation, and remote sensing monitoring chain for tailings emergency monitoring provides a corresponding execution model for emergency monitoring of tailings. The low-energy consumption deployment solution of the emergency drone ad hoc network based on the integration of communication, navigation, and sensing provides decision algorithms for user matching, communication resource matching, and drone scheduling.
[0004] However, the existing technical solutions have some shortcomings. First, the existing technology does not take into account the emergency scenarios caused by bad weather and other reasons, such as the lack of assistance from aerial self-organizing communication network nodes. Second, complex geographical environment elements including vegetation and buildings have a huge impact on communication, navigation and remote coordination. Summary of the invention
[0005] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a communication navigation perception depth enhanced spatiotemporal collaborative networking perception optimization method and device, aiming to overcome the challenges of complex geographical environment to ground high-frequency wireless ad hoc network communication. In the absence of aerial ad hoc communication network nodes, this method achieves a significant improvement in communication rate and deployment location effectiveness by optimizing the spatiotemporal collaborative deployment of ground communication, perception and data receiving nodes.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0007] A communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method, comprising:
[0008] Build a digital surface model of the disaster area and mark the obstacles;
[0009] Build a communication fitting model for the ad hoc network in the disaster area;
[0010] Set up several sensing nodes, communication nodes and data receiving nodes on the ground in the disaster area to simulate the task planning of sensing nodes in the ad hoc network;
[0011] According to the sensing node task planning, build a deep reinforcement learning model based on ResNet-DDPG, including constructing a reinforcement learning state space for spatio-temporal optimization of the ad hoc network based on the ResNet network, constructing a reinforcement learning action space for spatio-temporal optimization of the ad hoc network, designing a reinforcement learning reward that integrates geographical space features based on the communication fitting model, performing reinforcement learning training based on the DDPG model, and outputting the optimal solution of the ad hoc network relative to the collaborative deployment of sensing and communication nodes after the training converges.
[0012] Furthermore, the method for building a digital surface model of the disaster area and marking the obstacles is as follows:
[0013] Utilize the three-dimensional terrain model of the disaster area to extract the geographical elements of the obstacles and generate a three-dimensional geographical entity model and its digital surface model;
[0014] Divide the two-dimensional plane of the disaster area into grids, resample the actual height of the obstacles in the digital surface model, and then mark and assign values in the grids.
[0015] Furthermore, the method for building a communication fitting model for the ad hoc network in the disaster area is as follows:
[0016] Randomly design multiple measurement points in the disaster area and place communication nodes;
[0017] Construct a communication fitting model, and the communication fitting model is a backpropagation neural network;
[0018] Construct a communication fitting model, the communication fitting model is a BP neural network model, set the input as the distance of the penetrated obstacle, and the output as the loss coefficient of the network transmission speed between two communication nodes. The loss coefficient of the network transmission speed between two communication nodes is the ratio of the communication rate between two communication nodes to the theoretical maximum communication rate between two communication nodes. Train the communication fitting model to obtain a trained communication fitting model.
[0019] Furthermore, the method for setting up several sensing nodes, communication nodes and data receiving nodes in the disaster area to simulate the task planning of sensing nodes in the ad hoc network is as follows:
[0020] The perception path planning of the perception nodes is carried out by adopting area traversal, grid scanning or target surround. The perception path of each perception node is represented by the displacement per unit time step.
[0021] Acquire the coordinate position of the sensing node in real time in the environment, and schedule the communication node to provide a continuous and stable data transmission chain between the sensing node and the fixed-position data receiving node;
[0022] Based on the simulation of the positions of several perception nodes, communication nodes and data receiving nodes, the overall communication perception nodes are deployed in a coordinated manner to complete the digital modeling of perception node task planning.
[0023] Furthermore, the method for constructing a spatiotemporal optimized reinforcement learning state space based on a self-organizing network of ResNet network is as follows:
[0024] The digital surface model of the disaster area is represented as a raster image as the first layer of raster; the geographical features of obstacles are extracted and marked as the second layer of raster; a raster with the same size as the geographical space of the disaster area is created to generate the location mask of the node, and the longitude and latitude coordinates of the sensing nodes, communication nodes and data receiving nodes in the ad hoc network at a certain moment are converted into corresponding raster coordinates as the third layer of raster, and different labels are used to represent different node types;
[0025] The image formed by splicing three layers of grids is input into the ResNet18 network with pre-trained weights to obtain the reinforcement learning state space of the self-organizing network with spatiotemporal optimization.
[0026] Furthermore, the method for constructing a spatiotemporal optimized reinforcement learning action space for ad hoc networks is:
[0027] The action of each communication node at a certain moment is represented by polar coordinates in two-dimensional space;
[0028] Polar coordinates are converted into node displacements in the horizontal and vertical directions to update the communication node position while avoiding building obstacles. The calculation formula is:
[0029]
[0030] Among them, Δx, Δy are polar coordinates converted into the horizontal and vertical displacements of the node, r is the velocity constant, ρ is the distance controlled by the communication node, and θ is the radian system used to control the direction of movement of the communication node;
[0031] The final action vector at a moment will be a vector composed of the displacements of all communication nodes. Each communication node will change its position according to its own displacement vector. When the agent receives the action vector, it will calculate the displacement of each communication node, and the environment will regenerate the position mask to complete the state update.
[0032] The location of the sensing node is updated according to the sensing node task planning.
[0033] Furthermore, the method for designing reinforcement learning rewards that incorporate geospatial features is:
[0034] Through the constructed communication fitting model, the communication rate between any two nodes in the disaster area is calculated taking into account the terrain and the occlusion of obstacles.
[0035] When the environment takes action and the position status of each node in the ad hoc network changes, a weighted graph is formed, where the weight is the communication rate between two nodes on the link;
[0036] For each path, the link segment with the lowest rate is used as the communication rate value of the entire link. The sensing node selects the path with the highest communication rate among all the paths to the data receiving node as the communication rate of the sensing node.
[0037] Set reward item 1 (t), the calculation formula is:
[0038]
[0039] Among them, c ij is the communication rate between the i-th sensing node and the j-th communication node, c jk is the communication rate between the jth communication node and the kth data receiving node, c ik is the communication rate between the i-th sensing node and the k-th data receiving node, N s is the total number of sensing nodes, N c is the total number of communication nodes, N b is the total number of data receiving nodes, For normalization processing;
[0040] Set the penalty term r 2 (t), the calculation formula is:
[0041]
[0042] Wherein, count(·) is a count, which is used to count the number of sensing nodes whose communication rate is less than the preset minimum communication rate threshold;
[0043] The reinforcement learning reward r(t) integrating geospatial features is:
[0044] r(t)=(1-r 2 (t))·r 1 (t).
[0045] Furthermore, the method of reinforcement learning training based on the DDPG model is as follows:
[0046] Generate a location mask according to the location information of the communication node and the sensing node;
[0047] The digital surface model of the disaster area is spliced with the location mask, and then the feature vector is output by the pre-trained ResNet18 model as the state of retaining spatial information features;
[0048] A DDPG model is established. The DDPG model includes a policy network and an evaluation network. The agent generates an action strategy through the policy network according to the environment state. The action strategy is represented by an action vector. The action vector records the moving distance and direction of all communication nodes and is represented in the form of a polar coordinate system. The action strategy is scored by the evaluation network. The policy network uses gradient descent for training and learning to improve the score. The evaluation network calculates the time difference loss between the evaluation of the state and action at a certain moment and the evaluation of the state and action at the next moment, and uses rewards closer to those given by the environment as its learning goal for learning and training.
[0049] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method is implemented.
[0050] A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method.
[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0052] The present invention proposes a method and device for optimizing communication and navigation perception by enhancing spatiotemporal collaborative networking, aiming to overcome the challenges of complex geographical environment to ground high-frequency wireless ad hoc network communication. In the absence of aerial ad hoc communication network nodes, this method achieves a significant improvement in communication rate and deployment location effectiveness by optimizing the spatiotemporal collaborative deployment of ground communication, perception and data receiving nodes. Experimental results show that this method is of great significance in communication support and command and dispatch in disaster emergency response, especially under severe weather conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present embodiment. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0054] Figure 1 It is a schematic flow chart of the communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method of the present invention.
[0055] Figure 2 This is an application scenario diagram of the present invention.
[0056] Figure 3 This is the algorithm model diagram of the present invention.
[0057] FIG. 4( a ) is a schematic diagram of a node layout method and an obstacle blocking distance calculation method of the present invention.
[0058] FIG4( b ) is a network structure diagram of the communication fitting model of the present invention.
[0059] FIG4( c ) is a graph showing the change in loss value during the communication fitting model training process of the present invention.
[0060] Figure 5 This is a state design diagram of the present invention.
[0061] Figure 6 It is the action design diagram of the present invention.
[0062] Figure 7 The figure is a flow chart of the algorithm of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] First embodiment
[0065] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0066] In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.
[0067] It should be understood that the present invention can be implemented in different forms and should not be interpreted as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the disclosure thorough and complete and to fully convey the scope of the present invention to those skilled in the art.
[0068] In extreme natural disasters, emergency personnel often face the dilemma of being unable to communicate due to the damage of public network communication infrastructure. In the process of emergency, multimodal perception data such as location, video, photo, laser, etc. cannot be transmitted in real time, which leads to the inability to command and dispatch front-line emergency personnel to effectively perform emergency tasks in real time during the rescue command process, seriously affecting the development of emergency response. Ad Hoc Network (ANET) can quickly build an emergency communication network by being mounted on a mobile device or portable platform, and because the ANET network does not require pre-built network facilities (such as communication base stations, etc.), it is suitable for environments where it is difficult or impossible to deploy public communication network facilities. Therefore, it is possible to quickly build a wireless communication network in the disaster area when the public network communication infrastructure is damaged to provide communication coverage for the emergency area. However, how to carry out the effective deployment of ANET to provide high-quality stable communication is still a problem that needs to be studied and solved in depth. In this regard, this embodiment proposes a communication navigation perception depth enhanced spatiotemporal collaborative networking perception optimization method to overcome the problem of ground high-frequency wireless ad hoc network communication in complex geographical environment, and to promote the spatiotemporal collaborative optimization of ground communication-navigation-perception nodes in the absence of air ad hoc communication nodes. This effort aims to significantly enhance the real-time online disaster emergency perception capabilities in extreme disaster environments, thereby providing strong support for efficient emergency response. After experimental verification, the method of this embodiment has shown obvious advantages over similar methods in terms of communication rate improvement and deployment location effectiveness, and can effectively support disaster emergency response that only relies on ground emergency resources. It is of great significance for disaster emergency perception and command and dispatch under severe weather conditions.
[0069] The communication node, perception node and data receiving node provided in this embodiment are defined.
[0070] Communication node: responsible for data transmission, such as routers, bridges, switches, modems, hubs, etc., used for data forwarding and processing.
[0071] Data receiving node: a data receiving node that performs unified scheduling and receives data, such as a server or base station.
[0072] Perception node: A perception node that independently performs perception tasks and obtains video and image data, such as sensors or intelligent measurement and control equipment.
[0073] The first embodiment provides a communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method, such as Figure 1 and Figure 2 As shown, including:
[0074] Step S1: Building a digital surface model of the disaster area and marking obstacles;
[0075] Step S2: Establishing a communication fitting model of the ad hoc network in the disaster-stricken area;
[0076] Step S3: setting up a number of sensing nodes, communication nodes and data receiving nodes in the disaster-stricken area to simulate the task planning of sensing nodes in the ad hoc network;
[0077] Step S4: According to the task planning of the perception nodes, a deep reinforcement learning model based on ResNet-DDPG is established, including constructing a spatiotemporal optimized reinforcement learning state space of the ad hoc network based on the ResNet network, constructing a spatiotemporal optimized reinforcement learning action space of the ad hoc network, designing a reinforcement learning reward that integrates geographic spatial features based on the communication fitting model, and performing reinforcement learning training based on the DDPG model. After the training converges, the optimal solution for the collaborative deployment of the ad hoc network relative to the perception communication nodes is output.
[0078] The present invention proposes a method and device for optimizing communication and navigation perception by enhancing spatiotemporal collaborative networking, aiming to overcome the challenges of complex geographical environment to ground high-frequency wireless ad hoc network communication. In the absence of aerial ad hoc communication network nodes, this method achieves a significant improvement in communication rate and deployment location effectiveness by optimizing the spatiotemporal collaborative deployment of ground communication, perception and data receiving nodes. Experimental results show that this method is of great significance in communication support and command and dispatch in disaster emergency response, especially under severe weather conditions.
[0079] The above steps are described in detail below.
[0080] In step S1, the digital surface model is established to calculate the communication rate between nodes on the one hand, and to enable the communication nodes to perceive the surrounding environment information and dynamically plan their own positions on the other hand. The application scenarios in the disaster area are as follows: Figure 3 As shown. Since obstacles have an impact on signal transmission, it is necessary to mark the obstacles, and set the obstacles as buildings and vegetation. In this embodiment, the three-dimensional terrain model of the disaster-stricken area is used to extract the two types of geographical elements, buildings and vegetation, to generate a three-dimensional geographic entity model of the disaster-stricken area and its digital surface model (Digital Surface Model, DSM for short). The two-dimensional plane of the disaster-stricken area is divided into grids, and the actual digital surface model (DSM) heights of buildings and vegetation are resampled, and finally the grids are marked and assigned values.
[0081] In step S2, establishing a communication fitting model can effectively quantify the impact of buildings and vegetation on communication signals, calculate the communication rate between nodes, and the calculation of rewards is also based on the communication rate. The method for establishing a communication fitting model for the self-organizing network in the disaster-stricken area is:
[0082] Step S201: randomly design multiple measurement points in the disaster-stricken area and place communication nodes of the ad hoc network;
[0083] Step S202: construct a communication fitting model, which is a BP neural network model. The input is set to be the distance of the obstacle penetrated, and the output is the loss coefficient of the network transmission speed between the two communication nodes. The loss coefficient of the network transmission speed between the two communication nodes is the ratio of the communication rate between the two communication nodes to the theoretical maximum communication rate between the two communication nodes. The communication fitting model is trained to obtain a trained communication fitting model.
[0084] Specifically, as shown in Figure 4(a), the BP neural network model performs supervised learning on the data, with the input being the distance d of the building penetrated building and the distance of vegetation penetrated d tree , the output is the loss coefficient L of the network transmission speed between the two communication nodes b , the representation of the BP neural network model is:
[0085] L b =θ(d building , d tree )
[0086] The loss coefficient L of the network transmission speed between two communication nodes b The calculation formula is:
[0087]
[0088] Among them, C max is the theoretical maximum communication rate between two communication nodes, and c is the communication rate c between two communication nodes, which is obtained through measurement.
[0089] As shown in FIG4(b), the BP neural network model is a back propagation neural network. In the specific implementation of this embodiment, two input neurons are set in the first layer as the input layer, four neurons are set in the second layer as the hidden layer, and one neuron is included in the third layer as the output layer. As shown in FIG4(c), the data for network training in the specific implementation of this embodiment is 8820 bars (d building , d tree , L b ) triplet data can achieve better convergence effect.
[0090] When the BP neural network model is trained, the loss coefficient L of the network transmission speed between two communication nodes is obtained by measuring the penetration of electromagnetic waves by buildings and vegetation. b, the network bandwidth between the two communication nodes is measured multiple times through iperf3, and the average value of the network speed during a period of stability is taken as the communication rate c between the two communication nodes. The coordinates and elevations of the two communication nodes are recorded to obtain the communication rate c between the two communication nodes measured in different environments.
[0091] In step S3, based on the digital surface model established in step S1, a number of sensing nodes, communication nodes and data receiving nodes are set in the disaster-stricken area according to actual needs, and task planning of different types of sensing nodes in the ad hoc network is simulated, specifically including:
[0092] Step S301: Performing perception path planning for the perception node using three path planning methods: area traversal, grid scanning, or target surround;
[0093] Step S302: acquiring the coordinate position of the sensing node in real time in the environment, and scheduling the communication node to provide a continuous and stable data transmission link between the sensing node and the data receiving node with a fixed position;
[0094] Step S304: Based on the simulation of the positions of a number of perception nodes, communication nodes and data receiving nodes, the overall communication perception node collaborative deployment is realized to complete the digital modeling of the perception node task planning.
[0095] Specifically, during the path planning process, the perception path of each perception node is expressed according to the displacement per unit time step, and the calculation formula is:
[0096] P i = {P i (t 1 ), P i (t 2 ),...,P i (t N ))
[0097] Among them, P i represents the trajectory of the sensing node i, P i (t n ) indicates that the sensing node i is at t n displacement of time;
[0098] p(t n+1 )=p(t n )+P i (t n )·v,n∈[1,N]
[0099] Among them, p(t n+1 ) is t n+1 The position of the sensing node at the moment, p(t n ) is t n The position of the sensing node at the moment, v is the velocity constant.
[0100] In the specific implementation of this embodiment, a self-organizing network is set up, and the spatiotemporal dynamic environment is constructed based on gym. In a two-dimensional scene simulating reality of 500m*500m, it is divided into grids with a side length of 1m, so the number of grids is 500*500=2.5e5. Among them, the proportion of buildings and vegetation is approximately 21.34% and 34.58% respectively. The actual DSM heights of buildings and vegetation are resampled and the grids are marked and assigned. Four sensing nodes, two communication nodes and one data receiving node are set in the self-organizing network. Two of the four sensing nodes adopt the target surround method, and the other two adopt the area traversal and raster scanning methods respectively.
[0101] In this embodiment, in step S4, Figure 5 As shown in the figure, a spatiotemporal optimized reinforcement learning state space of a self-organizing network based on the ResNet network is constructed, and a state space encoding method integrating spatial encoding features is designed to construct a spatiotemporal optimized reinforcement learning state space of a self-organizing network. Specifically:
[0102] Step S401: The digital surface model of the disaster area is represented as a raster image as the first layer of raster; the geographical elements of obstacles are extracted and marked as the second layer of raster; a raster with the same size as the geographical space of the disaster area is created to generate a node position mask, and the longitude and latitude coordinates of the sensing nodes, communication nodes and data receiving nodes in the ad hoc network at a certain moment are converted into corresponding raster coordinates as the third layer of raster, and different labels are used to represent different node types;
[0103] Step S402: Input the image formed by splicing the three layers of grids into the ResNet18 network with pre-trained weights, so as to obtain the reinforcement learning state space of the self-organizing network with spatiotemporal optimization.
[0104] Specifically, the digital surface model (DSM) of the disaster-stricken area is represented as a raster image as the first layer of raster, where the horizontal and vertical coordinates represent their positions in the projected coordinate system, and the raster values are the elevations of geographic features. The elevations are normalized and multiplied by 255 to facilitate the subsequent extraction of spatial features.
[0105] ArcGIS was used to extract features such as buildings and vegetation on the terrain, and different labels were used to mark them as the second layer of raster.
[0106] Create a grid of the same size as the geographic space to generate the node location mask, convert the longitude and latitude coordinates of the sensing nodes, communication nodes, and data receiving nodes in the ad hoc network at a certain moment into corresponding grid coordinates, and use different labels to represent different node types.
[0107] The three layers of grids are stitched together to form an image I of size H×W×3, where each pixel has a value of [0-255];
[0108] The image of the three-layer grid is input into the ResNet18 network with pre-trained weights to obtain the reinforcement learning state space of the self-organizing network spatiotemporal optimization, where the state feature vector at a certain time is s(t).
[0109] In this embodiment, in step S4, a reinforcement learning action space for spatiotemporal optimization of the ad hoc network is constructed, such as Figure 6 As shown, specifically:
[0110] The action of each communication node at a certain time t is expressed as m(t) using polar coordinates in two-dimensional space:
[0111] m(t)=(ρ, θ), ρ∈[-1, 1], θ∈[-1, 1]
[0112] Among them, ρ is the distance that the communication node moves forward, θ is the radian system used to control the direction of movement of the communication node, a positive number indicates movement in the original direction, and a negative number indicates movement in the opposite direction.
[0113] Polar coordinates are converted into node displacements in the horizontal and vertical directions Δx and Δy to update the position of the communication node. When calculating the actual change of the coordinates, a speed constant r is multiplied to represent the maximum displacement that the entity can make per unit time, so that its movement distance is consistent with the speed in reality. At the same time, when encountering a building, it avoids the building by turning randomly:
[0114]
[0115] Finally, the action vector a(t) at a certain moment will be the vector formed by the displacement of all communication nodes. Each communication node will change its position according to its own displacement vector:
[0116]
[0117] Among them, m j (t) refers to the displacement vector of communication node jj at time t. j here represents the enumeration of communication nodes, referring to a certain communication node, such as the 1st, 2nd, jth, ... Nth c N c Indicates the total number of communication nodes.
[0118] When the agent receives the action vector, it calculates the displacement of each communication node, regenerates the position mask of the environment, and performs subsequent processes to complete the state update.
[0119] The location of the perception node is updated according to the perception node task planning, which is different from the update method of the communication node.
[0120] In this embodiment, in step S4, a reinforcement learning reward integrating geographic spatial features is designed, specifically:
[0121] Step S411: Using the communication fitting model constructed in step S2, the communication rate between any two nodes in the disaster-stricken area is calculated taking into account the terrain and the occlusion of obstacles (buildings, vegetation, etc.).
[0122] Step S412: When the environment takes action and the position status of each node in the ad hoc network changes, the communication rate between two nodes is calculated to form a weighted graph, where the weight is the communication rate between the two nodes on the link.
[0123] Step S413: For each path, the link segment with the lowest rate is used as the communication rate value of the entire link. The sensing node selects the path with the highest communication rate among all paths to the data receiving node as the communication rate of the sensing node.
[0124] Step S414: Set reward item r 1 (t), the reward term is the average of the communication rates of all sensing nodes and is normalized according to the theoretical maximum communication rate of the communication node. The calculation formula of the reward term is:
[0125]
[0126] Among them, c ij is the communication rate between the i-th sensing node and the j-th communication node, c jk is the communication rate between the jth communication node and the kth data receiving node, c ik is the communication rate between the i-th sensing node and the k-th data receiving node, N s is the total number of sensing nodes, N c is the total number of communication nodes, N b is the total number of data receiving nodes, For normalization processing.
[0127] Step S415: To make the communication rate of all nodes greater than a threshold c thresold To avoid the model unilaterally increasing the rate of some nodes and ignoring the overall connectivity rate, set the penalty term r 2 (t) is used to obtain the communication rate less than the preset minimum communication rate threshold C threshold The proportion of perception nodes, the penalty term r 2 The calculation formula of (t) is:
[0128]
[0129] Wherein, count(·) is a count used to count the number of sensing nodes whose communication rate is less than a preset minimum communication rate threshold.
[0130] Step S416: The reinforcement learning reward r(t) integrating geographic spatial features is:
[0131] r(t)=(1-r 2 (t))·r 1 (t).
[0132] In this embodiment, in step S4, Figure 2 , Figure 5 and Figure 6 As shown in the figure, reinforcement learning training is performed based on the DDPG (deep deterministic policy gradient) model, specifically:
[0133] Step S421: Generate a location mask according to the location information of the communication node and the sensing node;
[0134] Step S422: splicing the digital surface model (DSM) of the disaster area with the location mask, and then outputting the feature vector through the pre-trained ResNet18 model as a state of retaining spatial information features;
[0135] Step S423: Establish a DDPG model. The DDPG model includes a policy network (Actor network) and an evaluation network (Critic network). The policy network (Actor network) is the actor, and the evaluation network (Critic network) is the evaluator. The agent generates an action strategy through the policy network according to the environmental state. The action strategy is represented by an action vector. The action vector records the moving distance and direction of all communication nodes and is represented in the form of a polar coordinate system. The action strategy is given to the evaluation network for scoring as an evaluation of the quality of the action. The policy network is trained and learned using a gradient descent method to maximize the score. The evaluation network calculates the state and action (s at a certain time t) t , a t ) and the state and action (s) at the next moment t+1 t+1 , a t+1 )’s q-values, and takes the reward (reward) given by the environment as its own learning goal for learning and training, thereby improving the scoring ability and accuracy.
[0136] like Figure 7As shown, step S423 specifically includes: in the specific training process, firstly, after the environment, model and experience replay pool are initialized, the agent (Agent) continuously interacts in the environment based on the current model, and saves the experience quadruple in the experience replay pool. When enough experience is collected, the training process starts;
[0137] During the training process, the policy network includes the online policy network and the target policy network, and the evaluation network includes the online evaluation network and the target evaluation network. t Input into the online policy network and get action a=μ(s t |θ μ ), and then s t and a are used as the input of the online evaluation network, and the q value of the online evaluation network is calculated as the evaluation of the current state and action, q = Q (s t , a|θ Q ), where s t represents the state at time t, θ represents the parameters of the neural network model, μ represents the online policy network, and θ μ represents the parameters of the online policy network, μ(·) represents the output of the online policy network, and the content in brackets represents the input of the online policy network; θ Q represents the parameters of the online evaluation network, Q represents the online evaluation network, Q(·) represents the output of the online evaluation network, and the content in brackets represents the input of the online evaluation network.
[0138] Use -q as loss, where N represents the number of samples in a batch, and update the policy network through policy gradient to achieve the goal of maximizing the q value:
[0139]
[0140] Among them, L(μ) represents the loss function of the online evaluation network.
[0141] Through the target evaluation network, s t+1 Score with a′ and get q′=Q′(s t+1 ,μ′(s t+1 |θ′)), calculate the time difference loss value L(Q) of q, q′ and the actual reward value r(t) at time t, and update the evaluation network:
[0142] y i =r(t)+γ·Q′(s t+1 ,μ′(s t+1 |θ μ′ )|θ Q′ )
[0143]
[0144] Among them, s t+1 represents the updated state, a′ represents the updated action, q′ represents the evaluation of the updated state and action by the target evaluation network, θ′ represents the parameters of the updated neural network model, μ′ represents the target policy network, θ μ′ represents the parameters of the target policy network, μ′(·) represents the output of the target policy network, and the content in brackets represents the input of the target policy network; θ Q′ represents the parameters of the target evaluation network, Q′ represents the target evaluation network, Q′(·) represents the output of the target evaluation network, the content in brackets represents the input of the target evaluation network, and γ represents the discount factor, which is a concept in reinforcement learning and indicates the importance of future rewards r to the current decision. The higher the factor, the more importance is attached to subsequent rewards.
[0145] After taking the derivatives of both, the gradient of network update can be calculated:
[0146]
[0147] in, It means to find the gradient of A to update the parameters of B, then It means to find the gradient of L(μ) to update the parameters of θμ. It means to find the gradient of Q to update the parameters of a. It means to find the gradient of μ to update the parameters of θ. Indicates that the gradient of L(Q) is sought to update θ Q Parameters, Indicates that the gradient of Q is used to update θ Q Parameters, Indicates finding y i The gradient of θ is used to update θ Q Parameters.
[0148] According to the gradient, the network parameters of the online policy network and the online evaluation network are updated respectively, where lr a and lr c The learning rates of the online policy network and the online evaluation network are respectively. In practice, the Adam optimizer is used:
[0149]
[0150] The soft update method is used to update the two target networks, namely the target policy network and the target evaluation network, where τ is the updated weight coefficient, indicating the proportion of the model parameters updated each time:
[0151]
[0152] This embodiment uses a target network to reduce the impact of bootstrapping during reinforcement learning training, and adopts a soft update network update method to update only part of the parameters from the real-time network to the target network each time.
[0153] In the specific implementation manner of this embodiment, the hyperparameters of the training process are set as follows: Grid range: set to [500, 500].
[0154] Speed of RANET nodes: set to 5m / s. Speed of CNS nodes: set to 5m / s. Size of a mini-batch: set to 240. Actor's learning rate: set to 1e-6. Critic's learning rate: set to 1e-4. Action exploration: set to decay from 2 to 1e-6. Gamma: set to 0.99. Size of replay buffer: set to 2400. Training epochs: set to 500. Minimum communication rate threshold: set to 4Mb / s. Number of CNS nodes: set to 4. Number of RANET nodes: set to 2. Number of EC nodes: set to 1. Soft update weight coefficient (τ): set to 0.1.
[0155] In the specific implementation of this embodiment, as shown in Table 1 below, which is a comparison table of the average communication rate and improvement rate of the whole process of the present invention, it can be seen that the average communication rate of the whole process of the method of the present invention is improved compared with the communication nodes without relays, centralized deployment, and traditional DDPG methods, and the improvement rate of the method of the present invention is greatly improved compared with the communication nodes without relays.
[0156] Table 1. Comparison of average communication rate and improvement rate during the whole process
[0157]
[0158] This embodiment has the following beneficial effects:
[0159] This embodiment is based on the communication navigation perception depth enhancement space-time collaborative networking perception optimization method. The provided networking method shows obvious advantages in deployment location effectiveness and communication rate improvement compared with similar algorithms. At the same time, this research can effectively support emergency response in extreme disaster scenarios that lack the assistance of aerial self-organizing communication network nodes and rely only on ground self-organizing network emergency communication, which is of great significance for disaster emergency perception, command and dispatch, and life rescue under severe weather conditions.
[0160] Second embodiment
[0161] The second embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method when executing the program.
[0162] Third embodiment
[0163] The third embodiment provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method.
[0164] The memory in the embodiment of the present invention is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.
[0165] The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method disclosed in the embodiment of the present invention can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method can be completed by an integrated logic circuit of hardware or software instructions in the processor. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory, and the processor reads the information in the memory, and combines its hardware to complete the steps of the communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method provided in the embodiment of the present invention.
[0166] In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), FPGA, general purpose processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0167] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memories described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0168] The above embodiments are merely examples of the technical solutions of the present invention. The methods involved in the present invention are not limited to the contents described in the above embodiments, but are subject to the scope defined in the claims. Any modification, supplement or equivalent replacement made by a person skilled in the art based on the embodiment is within the scope of protection required by the claims of the present invention.
Claims
1. A communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method, characterized in that: include: Create a digital surface model of the affected area and mark obstacles; Establish a communication fitting model for the ad hoc network in the disaster-affected area; Several sensing nodes, communication nodes and data receiving nodes are set up on the ground in the disaster-stricken area to simulate the task planning of sensing nodes in the ad hoc network; According to the task planning of perception nodes, a deep reinforcement learning model based on ResNet-DDPG is established, including constructing a spatiotemporal optimized reinforcement learning state space of a self-organizing network based on the ResNet network, constructing a spatiotemporal optimized reinforcement learning action space of a self-organizing network, designing reinforcement learning rewards that integrate geographic spatial features based on the communication fitting model, and conducting reinforcement learning training based on the DDPG model. After training convergence, the optimal solution for the collaborative deployment of the self-organizing network relative to the perception communication nodes is output.
2. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: The method to build a digital surface model of the disaster area and mark obstacles is as follows: Using the 3D terrain model of the disaster-affected area, the geographical elements of obstacles are extracted to generate a 3D geographic entity model and its digital surface model; The two-dimensional plane of the disaster area is divided into grids, and the actual heights of obstacles in the digital surface model are resampled and marked and assigned in the grids.
3. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: The method for establishing the communication fitting model of the ad hoc network in the disaster-stricken area is: Randomly design multiple measurement points in the disaster-stricken area and place communication nodes; Construct a communication fitting model, which is a back propagation neural network; A communication fitting model is constructed. The communication fitting model is a BP neural network model. The input is set as the distance of the obstacle penetrated, and the output is the loss coefficient of the network transmission speed between the two communication nodes. The loss coefficient of the network transmission speed between the two communication nodes is the ratio of the communication rate between the two communication nodes to the theoretical maximum communication rate between the two communication nodes. The communication fitting model is trained to obtain a trained communication fitting model.
4. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: Several sensing nodes, communication nodes and data receiving nodes are set up in the disaster-stricken area. The method of simulating the task planning of sensing nodes in the ad hoc network is as follows: The perception path planning of the perception nodes is carried out by adopting area traversal, grid scanning or target surround. The perception path of each perception node is represented by the displacement per unit time step. Acquire the coordinate position of the sensing node in real time in the environment, and schedule the communication node to provide a continuous and stable data transmission chain between the sensing node and the fixed-position data receiving node; Based on the simulation of the positions of several perception nodes, communication nodes and data receiving nodes, the overall communication perception nodes are deployed in a coordinated manner to complete the digital modeling of perception node task planning.
5. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: The method for constructing a spatiotemporal optimized reinforcement learning state space for a self-organizing network based on the ResNet network is: The digital surface model of the disaster area is represented as a raster image as the first layer of raster; the geographical features of obstacles are extracted and marked as the second layer of raster; a raster with the same size as the geographical space of the disaster area is created to generate the location mask of the node, and the longitude and latitude coordinates of the sensing nodes, communication nodes and data receiving nodes in the ad hoc network at a certain moment are converted into corresponding raster coordinates as the third layer of raster, and different labels are used to represent different node types; The image formed by splicing three layers of grids is input into the ResNet18 network with pre-trained weights to obtain the reinforcement learning state space of the self-organizing network with spatiotemporal optimization.
6. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 5 is characterized in that: The method to construct a reinforcement learning action space for spatiotemporal optimization of ad hoc networks is: The action of each communication node at a certain moment is represented by polar coordinates in two-dimensional space; Polar coordinates are converted into node displacements in the horizontal and vertical directions to update the communication node position while avoiding building obstacles. The calculation formula is: Among them, Δx, Δy are polar coordinates converted into the horizontal and vertical displacements of the node, r is the velocity constant, ρ is the distance controlled by the communication node, and θ is the radian system used to control the direction of movement of the communication node; The final action vector at a moment will be a vector composed of the displacements of all communication nodes. Each communication node will change its position according to its own displacement vector. When the agent receives the action vector, it will calculate the displacement of each communication node, and the environment will regenerate the position mask to complete the state update. The location of the sensing node is updated according to the sensing node task planning.
7. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: The approach to designing reinforcement learning rewards that incorporate geospatial features is: Through the constructed communication fitting model, the communication rate between any two nodes in the disaster area is calculated taking into account the terrain and the occlusion of obstacles. When the environment takes action and the position status of each node in the ad hoc network changes, a weighted graph is formed, where the weight is the communication rate between two nodes on the link; For each path, the link segment with the lowest rate is used as the communication rate value of the entire link. The sensing node selects the path with the highest communication rate among all the paths to the data receiving node as the communication rate of the sensing node. Set the reward term r1(t), and the calculation formula is: Among them, c ij is the communication rate between the i-th sensing node and the j-th communication node, c jk is the communication rate between the jth communication node and the kth data receiving node, c ik is the communication rate between the i-th sensing node and the k-th data receiving node, N s is the total number of sensing nodes, N c is the total number of communication nodes, N b is the total number of data receiving nodes, For normalization processing; Set the penalty term r2(t), and the calculation formula is: Wherein, count(·) is a count, which is used to count the number of sensing nodes whose communication rate is less than the preset minimum communication rate threshold; The reinforcement learning reward r(t) integrating geospatial features is: r(t)=(1-r2(t))·r1(t).
8. The communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method according to claim 1 is characterized in that: The specific method for reinforcement learning training based on the DDPG model is: Generate a location mask according to the location information of the communication node and the sensing node; The digital surface model of the disaster area is spliced with the location mask, and then the feature vector is output by the pre-trained ResNet18 model as the state of retaining spatial information features; A DDPG model is established. The DDPG model includes a policy network and an evaluation network. The agent generates an action strategy through the policy network according to the environment state. The action strategy is represented by an action vector. The action vector records the moving distance and direction of all communication nodes and is represented in the form of a polar coordinate system. The action strategy is scored by the evaluation network. The policy network uses gradient descent for training and learning to improve the score. The evaluation network calculates the time difference loss between the evaluation of the state and action at a certain moment and the evaluation of the state and action at the next moment, and uses rewards closer to those given by the environment as its learning goal for learning and training.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method as described in any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the communication navigation perception depth enhancement spatiotemporal collaborative networking perception optimization method is implemented.
Citation Information
Patent Citations
Mobile self-organizing network node moving method based on double virtual forces
CN110213782A
Map-free obstacle avoidance navigation method based on distribution estimation and reinforcement learning
CN111707270A
Unmanned aerial vehicle cluster optimization deployment method and system based on deep learning
CN117615384A
Emergency scene-oriented air base station deployment method
CN117793723A
Space-time dynamic deployment method for air-ground ad hoc communication network based on deep reinforcement learning
CN117835463A