Mobile sensing method based on multi-agent reinforcement learning in industrial wireless network
Through the multi-agent reinforcement learning method, multi-mobile robots collect signal strength information, combined with classification and reinforcement learning models, the problems of insufficient coverage and poor real-time performance in complex environments are solved, and higher accuracy and stable indoor positioning are achieved.
Patent Information
- Application Number
- CN202510359761.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional indoor positioning technology has limited coverage scenarios in complex environments, poor positioning real-time performance, and high maintenance costs, making it difficult for the existing technology to adapt to environmental changes.
Multi-agent reinforcement learning method is adopted, and signal strength information is collected through multi-mobile robots, and classification model and reinforcement learning model are combined to plan the robot path to improve positioning accuracy and stability.
It realizes a larger coverage and higher precision positioning in complex indoor environments, improves positioning speed and accuracy, and reduces the impact of environmental changes on the positioning system.
Smart Images

Figure CN120456230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless positioning technology, and in particular to a mobile perception method based on multi-agent reinforcement learning in an industrial wireless network. Background Art
[0002] With the widespread adoption of mobile devices and the continuous advancement of intelligent technology, the demand for indoor location-based services has increased significantly. For example, applications such as crowd monitoring in shopping malls, locating individuals, and navigating underground garages urgently require high-precision indoor positioning technology. Because the performance of the Global Positioning System (GPS) degrades significantly in indoor environments, researchers are exploring alternatives based on short-range wireless signals.
[0003] Currently, a wide range of indoor positioning technologies are emerging, primarily including methods that leverage Wi-Fi, Bluetooth, and ultrasonic signals to achieve high-precision positioning. These technologies leverage the diverse characteristics of indoor wireless signals and have become a key means of meeting the needs of modern indoor location-based services (LBS). Among these various technical solutions, Wi-Fi fingerprint positioning has garnered particular attention due to its wide network coverage and compatibility with existing devices.
[0004] Traditional indoor positioning systems can be broadly categorized into two types: geometric positioning and fingerprint positioning. Geometric methods analyze parameters such as received signal strength (RSS), time of flight (TOF), and angle of arrival (AOA) to build a wireless signal propagation model to estimate the target's location. However, due to significant signal fluctuations and multipath effects in indoor environments, constructing an accurate propagation model is difficult. Therefore, positioning methods that utilize the unique "fingerprint" of Wi-Fi signals at each discrete spatial point are more effective.
[0005] While fingerprint positioning can theoretically achieve high positioning accuracy, practical applications often require deploying multiple Wi-Fi access points indoors to collect sufficient data and build a comprehensive fingerprint database. This technology is even more challenging to implement in areas with sparse AP deployment. Furthermore, any changes to the indoor environment or adjustments to AP configurations can render the existing fingerprint database invalid, necessitating re-collection and updating of data, which undoubtedly increases maintenance costs and workload.
[0006] Both traditional single-station positioning and mobile single-station positioning solve the problem of wireless positioning requiring the deployment of a large amount of infrastructure. However, traditional single-station positioning is mostly based on line of sight and is not suitable for complex indoor environments. Mobile single-station positioning faces problems such as limited coverage scenarios and poor real-time positioning. Summary of the Invention
[0007] In view of this, the purpose of the embodiments of the present application is to provide a mobile perception method based on multi-agent reinforcement learning in an industrial wireless network, which can improve the problems of limited coverage scenarios and poor real-time positioning of traditional mobile single stations.
[0008] To achieve the above technical objectives, the technical solutions adopted in this application are as follows:
[0009] A mobile perception method based on multi-agent reinforcement learning in an industrial wireless network is applied to a multi-agent positioning system. The method includes:
[0010] When the multi-agent positioning system is in an online stage, the states of all mobile robots are obtained, where the states are either moving states or stopped states, wherein all the mobile robots are located in an indoor environment, the indoor environment includes a plurality of road sections, and a node is provided between each two road sections. When the mobile robot is located on a node, the mobile robot determines, based on a planning strategy, whether to keep moving or stop moving, so that the mobile robot is in a moving state or a stopped state. When it is determined to keep moving, the mobile robot moves to the road section corresponding to the planning strategy.
[0011] When at least one of the mobile robots is in the stopped state, all the mobile robots are stopped from moving at the node where they are located;
[0012] Acquire first signal strength information of all target points collected by the mobile robot, where the first signal strength information is collected by the mobile robot on the previous road section corresponding to the current node, and all the first signal strengths form a first signal strength matrix;
[0013] Inputting the first signal strength matrix into a trained classification model, so that the classification model outputs a similarity between the first signal strength matrix and each preset reference point;
[0014] Based on the similarity and the position of the corresponding preset reference point, an estimated position of the target point is obtained.
[0015] Furthermore, when the multi-agent positioning system is in an offline stage, the method further includes:
[0016] Obtaining a current state space of the mobile robot, the current state space including identification information of a node currently located by the mobile robot, identification information of all the road sections currently moved by the mobile robot, and second signal strength information of a test point corresponding to the current node, the second signal strength information being collected by the mobile robot in a previous road section corresponding to the current node, wherein the mobile robot and the test point are both located in an indoor environment;
[0017] Obtaining a current reward based on the second signal strength information of the test point corresponding to the current node and the reward function;
[0018] Inputting the current state space into a reinforcement learning model so that the reinforcement learning model outputs an action in the current state;
[0019] The current state space, current state action, next moment state space and current reward are used as first training data for training the reinforcement learning model. The reinforcement learning model is configured on the mobile robot so that in the online stage, after the mobile robot inputs the current state space into the reinforcement learning model, the reinforcement learning model outputs the planning strategy, wherein the next moment state space includes the node where the mobile robot is located at the next moment, the identification information of all the road sections where the mobile robot completes the movement at the next moment, and the second signal strength information corresponding to the node at the next moment. The next moment state space is obtained through the current state action. The mobile robot moves to the road section corresponding to the action in the current state based on the action in the current state.
[0020] Furthermore, the reward function is:
[0021]
[0022] Among them, CurrentStep represents the number of sections completed by the current mobile robot, and k represents the experimental parameter, which is 0.05-0.15;
[0023] in, Indicates the estimated position coordinates of the test point in the offline stage;
[0024] p is the real position coordinate of the test point;
[0025] Obtaining the estimated position coordinates of the test point, including:
[0026] Acquire second signal strength information of the test points collected by all the mobile robots at the nodes corresponding to the stopped state to form a second signal strength matrix;
[0027] Inputting the second signal strength matrix into the classification model, so that the classification model inputs a second similarity between the second signal strength matrix and the signal information of each of the preset reference points;
[0028] Based on the second similarity and the position of the corresponding preset reference point, the .
[0029] Furthermore, before obtaining the current state space of the mobile robot, the method further includes:
[0030] Obtaining an RSS fingerprint signal of each of the preset reference points;
[0031] The RSS fingerprint signals of all the preset reference points are used as the second training data and input into the pre-trained classification model to obtain the trained classification model, so that in the online stage, the step of inputting the first signal strength matrix into the trained classification model can be performed, so that the classification model outputs the similarity between the first signal strength matrix and the signal information of each preset reference point.
[0032] Furthermore, obtaining the RSS fingerprint signal of each preset reference point includes:
[0033] Acquire third signal strength information of all the preset reference points collected by the mobile robot on each road section;
[0034] Based on all the third signal strength information, an RSS fingerprint signal of each preset reference point is obtained.
[0035] Furthermore, the position of the target point is obtained according to a positioning algorithm based on all the matching values and the position of the preset reference point, including:
[0036] Obtaining the coordinates of all the preset reference points and the similarities corresponding to the preset reference points;
[0037] Inputting the coordinates of the preset reference point and the corresponding similarity into a positioning algorithm to obtain the coordinates of the position of the target point;
[0038] The positioning algorithm is:
[0039]
[0040] in, Sort in descending order;
[0041] P rpj Indicates the similarity of the corresponding preset reference point;
[0042] C j is the position coordinate of the j-th preset reference point;
[0043] N represents the number of preset reference points;
[0044] M represents a value between 3 and N, where M is a natural number and 3≤M<N;
[0045] estimate_position represents the position coordinates of the target point.
[0046] The invention adopting the above technical solution has the following advantages:
[0047] In the technical solution provided by the present application, first signal strength information is collected by multiple mobile robots, all the first signal strength information is input into the classification model to obtain all similarities, and then the position of the target point is obtained based on all the similarities and the position of the preset reference point. The present application achieves the purpose of positioning multiple mobile robots. Compared with the existing technology, multiple mobile robots can cover a larger scene, and more road sections can be set up in indoor environments, thereby improving positioning accuracy. The simultaneous collection of first signal strength information by multiple mobile robots can increase the dimension of features in time, thereby improving positioning speed. At the same time, based on the first signal strength information collected by multiple mobile robots, multiple first signal strength information can reflect the different characteristics of the indoor environment. By comprehensively collecting the first signal strength information, the signal strength distribution of the target point in the indoor environment can be more accurately reflected, thereby achieving more accurate positioning; at the same time, by collecting the first signal strength information at multiple positions, the present technical solution has better accuracy and stability compared to the single-station positioning mode of the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The present application may be further illustrated by the non-limiting embodiments provided in the accompanying drawings. It should be understood that the following drawings illustrate only certain embodiments of the present application and are therefore not to be construed as limiting the scope of the present application. It is understood that a person skilled in the art can derive other relevant drawings from these drawings without inventive effort.
[0049] Figure 1 This is a flowchart provided for an embodiment of the present application.
[0050] Figure 2 Schematic diagram of the indoor environment provided in an embodiment of the present application.
[0051] Figure 3 This is a block diagram of the network architecture proposed in the embodiments of the present application.
[0052] Figure 4 A schematic diagram of the offline phase process provided in an embodiment of the present application.
[0053] Figure 5 This is a sub-flowchart of S150 provided in an embodiment of the present application.
[0054] Figure 6 A schematic diagram of the online stage process provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts in the drawings or descriptions are numbered the same. Implementations not shown or described in the drawings are known to those of ordinary skill in the art. In the description of this application, the terms "first," "second," etc. are used solely to distinguish descriptions and are not to be construed as indicating or implying relative importance.
[0056] Please refer to Figure 1 The present application also provides a method for mobile perception based on multi-agent reinforcement learning in an industrial wireless network. The method for mobile perception based on multi-agent reinforcement learning in an industrial wireless network may include the following steps:
[0057] S110, when the multi-agent positioning system is in an online stage, obtaining the status of all mobile robots, wherein the status is a moving state or a stopped state, wherein all the mobile robots are located in an indoor environment, the indoor environment includes a plurality of road sections, and a node is provided between each two road sections. When the mobile robot is located at a node, the mobile robot determines whether to keep moving or stop moving based on a planning strategy, so that the mobile robot is in a moving state or a stopped state. When it is determined to keep moving, the mobile robot moves to the road section corresponding to the planning strategy;
[0058] S120, when at least one of the mobile robots is in the stopped state, causing all the mobile robots to stop moving at the node where they are located;
[0059] S130, obtaining first signal strength information collected by all the mobile robots, where the first signal strength information is collected by the mobile robots at the previous road section corresponding to the node where the mobile robots stopped moving;
[0060] S140, inputting the first signal strength matrix into a trained classification model, so that the classification model outputs a similarity between the first signal strength matrix and signal information of each preset reference point;
[0061] S150: Obtain an estimated position of the target point based on the similarity and the position of the corresponding preset reference point.
[0062] S110-S150 in this embodiment are all steps implemented in the online phase. In the offline phase before S110, it is necessary to train the reinforcement learning model and the classification model so that the mobile robot can determine the next path to take based on the planning strategy based on the reinforcement learning model. In this embodiment, a multi-mobile robot reinforcement learning framework is used to formulate multiple mobile robot path planning problems, improve the robustness of the path planning algorithm, and reduce positioning errors. The reinforcement learning model is trained as follows:
[0063] First, an example of a road segment is Figure 2 As shown, A1-A9 represent nodes, and the straight lines between adjacent nodes represent road sections. In this embodiment, after the mobile robot completes each road section and reaches the corresponding node, it selects the next road section to travel based on a greedy strategy. While moving along the road section, it collects signal strength information from preset reference points or test points.
[0064] In this embodiment, after the mobile robot completes each section of the road, it is recorded as a time t, and the state s is obtained from the environment. t ∈S, the mobile robot executes the ε-greedy strategy π(a t |s t ) Determine action a t ∈A, the environment transitions to the next state s t+1 ∈S, get reward r t , change {(s t ,a t ,r t ,s t+1 ,d t )} is saved in the experience replay pool and used to train the reinforcement learning model.
[0065] State space: Different mobile robots have different state spaces. Each mobile robot can get its own observation value in the indoor environment, which is specifically described as the node where the current mobile robot is located, using P i Indicate which sections of the road have been traveled, use E i It is initially filled with 0, has a maximum upper limit of step M, and the signal strength information of the road just walked, then the state of a mobile robot at a certain moment can be described as:
[0066] State={P i ,[E1,E2,…,E i …,E M ],RSSInfo}
[0067] RSS_Info={(MeanRss1,MinRss1,MaxRss1,StdRss1,StdAllRss,MeanX1,MeanY1),……,(MeanRss d ,MinRss d ,MxaRss d ,StdRss d ,StdAllRss,MeanX d ,MeanY d ),……,(MeanRss 30,MinRss 30 ,MeanRss 30 ,StdRss 30 ,StdAllRss,MeanX 30 ,MeanY 30 )}
[0068] Wherein, d (1<=d<=30) means that in this embodiment, a road section is divided into 30 segments, MeanRss d Refers to the average signal strength on the dth segment, MinRss d Refers to the minimum signal strength on segment d, MxaRss d Refers to the maximum signal strength on the dth segment, StdRss d Refers to the variance of the signal strength on the dth segment, StdAllRss refers to the variance of the signal strength on the entire path, MeanX d ,MeanY d Refers to the average values of the X-axis and Y-axis on the dth segment respectively.
[0069] Action space (action): Assume that this embodiment selects 26 road segments from the entire indoor environment as the paths that the mobile robot can take, and also includes a stop action. The action of a mobile robot at a certain moment can be described as follows:
[0070] Action={E1,E2,…,E i …,E 26 ,Stop}
[0071] Reward function: The reward function is designed based on the optimization goal, reducing the average number of steps required for positioning and improving positioning accuracy. It can be described as follows:
[0072]
[0073] Among them, PositonLoss represents the accuracy error obtained by calculating the weight of MSFMM output after the mobile robot completes the movement on the current road section. CurrentStep refers to the number of sections completed by the current mobile robot. k is an experimental parameter, and the value of this embodiment is 0.1. Figure 3 As shown. Among them:
[0074] The reward function is:
[0075]
[0076] Among them, CurrentStep represents the number of sections completed by the current mobile robot, and k represents the experimental parameter, which is 0.05-0.15;
[0077] in, Indicates the estimated position coordinates of the test point in the offline stage;
[0078] p is the real position coordinate of the test point;
[0079] Obtaining the estimated position coordinates of the test point, including:
[0080] Acquire second signal strength information of the test points collected by all the mobile robots at the nodes corresponding to the stopped state to form a second signal strength matrix;
[0081] Inputting the second signal strength matrix into the classification model, so that the classification model inputs a second similarity between the second signal strength matrix and the signal information of each of the preset reference points;
[0082] Based on the second similarity and the position of the corresponding preset reference point, the
[0083] In this embodiment, based on the following formula, Specifically:
[0084]
[0085] in, Sort in descending order;
[0086] P fp Indicates the similarity between the second signal strength and the signal strength of the corresponding preset reference point;
[0087] C q is the position coordinate of the qth preset reference point;
[0088] W represents the number of preset reference points;
[0089] H represents a value between 3 and N, M is a natural number, and 3≤M<N;
[0090] estimate_position1 means The location coordinates of .
[0091] In this embodiment, a reinforcement learning model is used to plan the path of a mobile robot. During the offline phase, the mobile robot follows a greedy strategy. After completing each section of a route, it inputs its current state space into its own reinforcement learning model. Each reinforcement learning model then uses this state to determine the next section of the route to take, i.e., the next action. This determines the next state, and a reward function assigns a reward to the mobile robot based on this action. The reward, the next state, the current state, and the current action are then fed into the reinforcement learning model as training data. Based on this training data, the reinforcement learning model outputs a planning strategy for each mobile robot.
[0092] In the online stage, when the mobile robot completes a section of the road and reaches the corresponding node, the mobile robot inputs the current state into the reinforcement learning model. The reinforcement learning model outputs the next appropriate section of the mobile robot. The mobile robot then moves according to the "appropriate next section", that is, the mobile robot moves on the sections in the indoor environment according to the planned strategy.
[0093] In this embodiment, each mobile robot is equipped with a reinforcement learning model. When positioning is finally required (that is, when the maximum number of steps is reached or an agent chooses to stop moving), a central server integrates the historical RSS_Info of all agents, inputs it into the classification network, and then gives the final positioning result.
[0094] The classification model effectively integrates convolutional neural networks (CNNs) with long short-term memory networks (LSTMs), supplemented by residual modules and self-attention mechanisms, to achieve efficient extraction and fusion of one-dimensional data features, significantly improving classification accuracy and system robustness. Specifically, the model uses two parallel feature extraction branches:
[0095] 1. CNN-LSTM branch
[0096] First, preliminary convolutional feature extraction is performed on the input data through a series of residual blocks (Restblock) and dropout layers to form a preliminary feature representation.
[0097] This feature is then transformed and fed into an LSTM layer to capture the temporal dynamics of the data. This branch extracts temporal correlation information using LSTM, and the final step uses a dropout layer to process the LSTM output to prevent overfitting.
[0098] 2.CNN-Attention branch:
[0099] In the other branch, the residual block is also used to process the original input. The difference is that residual modules with different convolution kernel sizes are used to capture multi-scale features.
[0100] Next, a custom-designed self-attention module (Attentionblock) was introduced, which uses 1×1 convolution to generate query (Q), key (K) and value (V), and calculates attention weights through a multi-head self-attention mechanism to effectively capture long-distance dependencies and global context information.
[0101] This branch is finally subjected to global average pooling to obtain a feature vector with high-dimensional semantic information.
[0102] Finally, the feature vectors extracted by the two branches (the temporal features from the LSTM branch and the global features from the CNN-Attention branch, respectively) are concatenated and input into a fully connected layer for final classification. In general, by combining the CNN-LSTM branch and the CNN-Attention branch in parallel, a complementary enhancement of the data's local temporal features and global contextual information is achieved. Simultaneously, the use of lightweight 1×1 convolutions to construct a multi-head self-attention mechanism in one-dimensional signal processing scenarios not only ensures computational efficiency but also effectively improves the ability to model complex multipath effects and long-range dependencies. Furthermore, the combined use of residual connections and multi-layer dropout techniques effectively alleviates the vanishing gradient problem in deep network training and prevents overfitting, significantly improving the model's generalization performance.
[0103] In this embodiment, Figure 4 As shown in the figure, the offline phase includes modules such as data collection, data cleaning, feature extraction, and deep reinforcement learning model training. In the offline phase, the first task is to collect RSS fingerprint signals for preset reference points. This step specifically includes:
[0104] The third signal strength information of all preset reference points collected by the mobile robot on each road section is obtained, and then the data is preprocessed, which mainly includes outlier removal and data fusion. The isolation forest method is used for outlier removal.
[0105] Based on all the third signal strength information, the RSS fingerprint signal of each preset reference point is obtained and stored in the fingerprint database. Therefore, in this step, this embodiment establishes a complete fingerprint database. Subsequently, in this embodiment, the entire fingerprint database is used to train the aforementioned classification network, which aims to match the similarity between the test point and each preset reference point. This completes the first step of the offline phase. The acquired fingerprint database is input into the aforementioned pre-trained classification model to obtain a trained classification model. The input of the classification model is K*RSS_Info, and the output is the matching probability between the target point and a single preset reference point.
[0106] In this example, reinforcement learning (QMIX) was used to implement a multi-agent indoor positioning system. In this multi-agent scenario, mobile robots work together to determine a location. During the offline phase, using a reinforcement learning model, the mobile robots learn which edges yield the best positioning accuracy for different test points, as well as when to stop the positioning process, ultimately providing an estimated position.
[0107] like Figure 4 As shown, all mobile robots move on a road section in an indoor environment (the indoor environment includes several test points). Every time a mobile robot completes the current road section, the state space and reward of all mobile robots are obtained, as well as the action selected by the mobile robot based on the greedy strategy according to the current reward. The state space at the next moment after executing the action is input into the reinforcement learning model for training the reinforcement learning model as training data.
[0108] Reinforcement learning models can also be trained as follows:
[0109] Obtaining a current state space of the mobile robot, the current state space including identification information of a node currently located by the mobile robot, identification information of all the road sections currently moved by the mobile robot, and second signal strength information of a test point corresponding to the current node, the second signal strength information being collected by the mobile robot in a previous road section corresponding to the current node, wherein the mobile robot and the test point are both located in an indoor environment;
[0110] Obtaining a current reward based on the second signal strength information of the test point corresponding to the current node and the reward function;
[0111] Inputting the current state space into a reinforcement learning model so that the reinforcement learning model outputs an action in the current state;
[0112] The current state space, current state action, next moment state space and current reward are used as first training data for training the reinforcement learning model. The reinforcement learning model is configured on the mobile robot so that in the online stage, after the mobile robot inputs the current state space into the reinforcement learning model, the reinforcement learning model outputs the planning strategy, wherein the next moment state space includes the node where the mobile robot is located at the next moment, the identification information of all the road sections where the mobile robot completes the movement at the next moment, and the second signal strength information corresponding to the node at the next moment. The next moment state space is obtained through the current state action. The mobile robot moves to the road section corresponding to the action in the current state based on the action in the current state.
[0113] In this embodiment, when training the classification model, preset reference points are set in an indoor environment, and their coordinates are input into a central server. When training the reinforcement learning model, the preset reference points are removed, and test points are set in the indoor environment. By setting up these test points, the mobile robot can perform path planning based on the reinforcement learning model. It should be noted that both the preset reference points and the test points have signal strength.
[0114] The following is a detailed description of the steps of the multi-agent reinforcement learning-based mobile perception method in industrial wireless networks:
[0115] S110 is the online phase of the process. This phase can refer to the operational phase, where the mobile robot is equipped with a device or system capable of determining its current location. Alternatively, the mobile robot can be a ROS car, which can return coordinates. These coordinates are essentially calculated using radar and odometry.
[0116] In this embodiment, the indoor environment includes multiple road sections, which are pre-designed and arranged according to preset shapes. As the mobile robot moves along the road sections, it continuously collects signals from target points and obtains signal strength information. Before this step, the coordinates of each road section and node are annotated. This allows the mobile robot to provide feedback on its coordinate position at any location, thereby achieving positioning.
[0117] In terms of obtaining the status of mobile robots, a mechanism can be designed to record the current status of each mobile robot. The mechanism can be: the mobile robot sends its own status to the control system in real time. When it is in a stopped state, the mobile robot sends a signal representing the stopped state. When the multi-agent positioning system receives the signal, it considers that the mobile robot is in a stopped state.
[0118] In this embodiment, the mobile robot has deep learning capabilities.
[0119] In S120, in this embodiment, the conditions for the mobile robot to stop moving at least meet one of the following conditions:
[0120] 1. The number of sections completed by the mobile robot reaches the preset upper limit;
[0121] 2. The planning strategy based on the output of the reinforcement learning model assumes that the mobile robot has received the maximum reward after reaching the current node.
[0122] Therefore, when any mobile robot stops moving, it means that the stopped mobile robot believes that it has obtained the maximum reward, that is, the path taken by all mobile robots in the past is most conducive to locating the position of the test point. At this time, all other mobile robots are ordered to stop moving.
[0123] In S130, the first signal strength information of the target point collected by all mobile robots on the previous road section is obtained. The first signal strength information can be represented as RSS_Info1
[0124] RSS_Info1={(MeanRss1,MinRss1,MaxRss1,StdRss1,StdAllRss,MeanX1,MeanY1),……,(MeanRss d ,MinRss d ,MxaRss d ,StdRss d ,StdAllRss,MeanX d ,MeanY d ),……,(MeanRss M ,MinRss M ,MxaRss M ,StdRss M ,StdAllRss,MeanX M ,MeanY M )}
[0125] Among them, d (1 <= d <= M) means that an edge is divided into M segments, MeanRss d Refers to the average signal strength on the dth segment, MinRss d Refers to the minimum signal strength on segment d, MxaRss d Refers to the maximum signal strength on the dth segment, StdRss d Refers to the variance of the signal strength on the dth segment, StdAllRss refers to the variance of the signal strength on the entire path, MeanX d ,MeanY dRefers to the average values of the X-axis and Y-axis on the dth segment respectively.
[0126] The second signal strength information can be expressed as:
[0127] RSS_Info2={(MeanRss1,MinRss1,MaxRss1,StdRss1,StdAllRss,MeanX1,MeanY1),……,(MeanRss d ,MinRss d ,MxaRss d ,StdRss d ,StdAllRss,MeanX d ,MeanY d ),……,(MeanRss N ,MinRss N ,MxaRss N ,StdRss N ,StdAllRss,MeanX N ,MeanY N )}.
[0128] The classification model of this embodiment is used to match the similarity between the test point and the preset reference point.
[0129] In S140, the data output by the classification model is output, where output is represented by:
[0130]
[0131] in, Represents the similarity between the first signal strength matrix and the signal strength of the corresponding preset reference point;
[0132] N represents the number of preset reference points.
[0133] For example, the fingerprint signal database includes the positions of three preset reference points. The classification model can output the similarity between the test point and the three preset reference points, and the similarity is expressed as probability.
[0134] In S150, Figure 5 As shown, the following steps are included:
[0135] S151: Acquire the coordinates of all the preset reference points and the similarities corresponding to the preset reference points;
[0136] S152: Inputting the coordinates of the preset reference point and the corresponding similarity into a positioning algorithm to obtain the coordinates of the position of the target point;
[0137] The positioning algorithm is:
[0138]
[0139] N represents the number of preset reference points;
[0140] C j is the position coordinate of the jth preset reference point;
[0141] M represents a value between 3 and N, where M is a natural number and 3≤M<N;
[0142] estimate_position represents the location coordinates of the target point.
[0143] In this embodiment, the method can be divided into two stages: an offline stage and an online stage.
[0144] S110-S150 is the line stage in this embodiment, and its specific logic diagram is as follows: Figure 6 In this embodiment, multiple mobile robots collect signal strength information of a target point on different sections of an indoor environment, and different mobile robots receive different first signal strength information each time.
[0145] Based on the current state space, the trained reinforcement learning model determines the next path the mobile robot should take, or whether it should stop. If no stop action is chosen, the mobile robot will continue to advance in the indoor environment until it reaches the set maximum step or stops. If a mobile robot chooses to stop, all other mobile robots will stop at that step. The historical signal strength information from all mobile robots is then combined (horizontally stacked) to output different matching probabilities for all preset reference points to the trained classification model. Finally, the estimated position of the test point is obtained using the Wknn algorithm.
[0146] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A mobile perception method based on multi-agent reinforcement learning in industrial wireless networks, characterized by: Applied to a multi-agent positioning system, the method includes: When the multi-agent positioning system is in an online stage, the states of all mobile robots are obtained, where the states are either moving states or stopped states, wherein all the mobile robots are located in an indoor environment, the indoor environment includes a plurality of road sections, and a node is provided between each two road sections. When the mobile robot is located on a node, the mobile robot determines, based on a planning strategy, whether to keep moving or stop moving, so that the mobile robot is in a moving state or a stopped state. When it is determined to keep moving, the mobile robot moves to the road section corresponding to the planning strategy. When at least one of the mobile robots is in the stopped state, all the mobile robots are stopped from moving at the node where they are located; Acquire first signal strength information of all target points collected by the mobile robot, where the first signal strength information is collected by the mobile robot on the previous road section corresponding to the current node, and all the first signal strengths form a first signal strength matrix; Inputting the first signal strength matrix into a trained classification model, so that the classification model outputs the similarity between the first signal strength matrix and the signal information of each preset reference point; Based on the similarity and the position of the corresponding preset reference point, an estimated position of the target point is obtained.
2. The method according to claim 1, characterized in that When the multi-agent positioning system is in an offline stage, the method further includes: Obtaining a current state space of the mobile robot, the current state space including identification information of a node currently located by the mobile robot, identification information of all the road sections currently moved by the mobile robot, and second signal strength information of a test point corresponding to the current node, the second signal strength information being collected by the mobile robot in a previous road section corresponding to the current node, wherein the mobile robot and the test point are both located in an indoor environment; Obtaining a current reward based on the second signal strength information of the test point corresponding to the current node and the reward function; Inputting the current state space into a reinforcement learning model so that the reinforcement learning model outputs an action in the current state; The current state space, the action in the current state, the state space at the next moment and the current reward are used as the first training data for training the reinforcement learning model. The reinforcement learning model is configured on the mobile robot so that in the online stage, after the mobile robot inputs the current state space into the reinforcement learning model, the reinforcement learning model outputs the planning strategy, wherein the state space at the next moment includes the node where the mobile robot is located at the next moment, the identification information of all the sections where the mobile robot completes the movement at the next moment, and the second signal strength information corresponding to the node at the next moment. The state space at the next moment is obtained through the action in the current state. The mobile robot moves to the section corresponding to the action in the current state based on the action in the current state.
3. The method according to claim 2, characterized in that The reward function is: Among them, CurrentStep represents the number of sections completed by the current mobile robot, and k represents the experimental parameter, which is 0.05-0.15; in, Indicates the estimated position coordinates of the test point in the offline stage; p is the real position coordinate of the test point; Obtaining the estimated position coordinates of the test point, including: Acquire second signal strength information of the test points collected by all the mobile robots at the nodes corresponding to the stopped state to form a second signal strength matrix; Inputting the second signal strength matrix into the classification model, so that the classification model inputs a second similarity between the second signal strength matrix and the signal information of each of the preset reference points; Based on the second similarity and the position of the corresponding preset reference point, the 4. The method according to claim 2, characterized in that Before obtaining the current state space of the mobile robot, the method further includes: Obtaining an RSS fingerprint signal of each of the preset reference points; The RSS fingerprint signals of all the preset reference points are used as the second training data and input into the pre-trained classification model to obtain the trained classification model, so that in the online stage, the step of inputting the first signal strength matrix into the trained classification model can be performed, so that the classification model outputs the similarity between the first signal strength matrix and the signal information of each preset reference point.
5. The method according to claim 4, characterized in that The obtaining of the RSS fingerprint signal of each preset reference point includes: Acquire third signal strength information of all the preset reference points collected by the mobile robot on each road section; Based on all the third signal strength information, an RSS fingerprint signal of each preset reference point is obtained.
6. The method according to claim 1, wherein The step of obtaining the position of the target point based on all the matching values and the position of the preset reference point according to a positioning algorithm includes: Obtaining the coordinates of all the preset reference points and the similarities corresponding to the preset reference points; Inputting the coordinates of the preset reference point and the corresponding similarity into a positioning algorithm to obtain the coordinates of the position of the target point; The positioning algorithm is: in, Sort in descending order; P rpj Represents the similarity between the first signal strength matrix and the signal strength of the corresponding preset reference point; C j is the position coordinate of the j-th preset reference point; N represents the number of preset reference points; M represents a value between 3 and N, where M is a natural number and 3≤M<N; estimate_position represents the position coordinates of the target point.