Fusion positioning method and system based on magnetic sensor and vision

By building a time series network that integrates the positioning data set and trains a magnetic perceptron, visual perceptron and multi-source locator, using magnetic field signals and point cloud data for dual memory updates, the problem of poor positioning effect in closed scenarios is solved, and efficient and accurate target positioning is achieved.

CN120256911APending Publication Date: 2025-07-04UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328284.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In closed scenarios, satellite positioning methods cannot be applied, while wireless positioning methods based on UWB and 5G require the deployment of a large number of base stations and the positioning cost is high, and wireless signals cannot be transmitted through buildings, resulting in poor positioning effect.

Method used

Using a fusion positioning method based on magnetic sensors and vision, a time-series network of magnetic perceptrons, vision perceptrons and multi-source locators is trained, and dual memory updates are used to achieve real-time and accurate positioning of the targets.

Benefits of technology

The target positioning accuracy and efficiency are significantly improved, and the problems of low positioning accuracy of a single magnetic sensor and high calculation cost of a single visual positioning are avoided, thereby achieving more accurate and efficient target positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256911A_ABST
    Figure CN120256911A_ABST
Patent Text Reader

Abstract

The invention provides a fusion positioning method and system based on a magnetic sensor and vision, and the method comprises the steps: constructing a fusion positioning data set, comprising a three-dimensional point cloud model of the building, magnetic field signals output by a magnetic field sensor array at multiple moments, peripheral point cloud data scanned by point cloud scanning equipment at multiple moments, and communication terminal position information at multiple moments; inputting the data and training a fusion positioning model comprising a magnetic sensor, a visual sensor and a multi-source positioner; the magnetic sensor is used for learning and extracting valuable magnetic field positioning information from the magnetic field signal; the visual sensor is used for learning and extracting valuable visual positioning information from the peripheral point cloud data, fusing the magnetic field positioning information and the visual positioning information and outputting multi-source positioning information; and the multi-source positioner is used for learning the valuable positioning information and the multi-source positioning information in the positioning memory unit and outputting an accurate positioning value of the target at the current moment. According to the invention, the target can be positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target positioning, and in particular to a fusion positioning method and system based on a magnetic sensor and vision. Background Art

[0002] For enclosed scenarios such as indoors, tunnels, and underground parking lots, satellite-based positioning methods are not applicable. Wireless positioning based on UWB, 5G, etc. requires the deployment of a large number of positioning base stations, resulting in high positioning costs. Moreover, wireless signals cannot penetrate buildings for transmission, leading to poor positioning effects. Therefore, there is an urgent need for a new positioning method. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a fusion positioning method based on a magnetic sensor and vision. The technical solution is as follows:

[0004] On the one hand, a fusion positioning method based on a magnetic sensor and vision is provided. The method includes:

[0005] S1. Construct a fusion positioning data set. The data in the data set includes: the three-dimensional point cloud model of the building itself, the magnetic field signals output by the magnetic field sensor array at multiple times, the surrounding point cloud data scanned by the point cloud scanning device at multiple times, and the position information of the communication terminal at multiple times. Divide the data set into a training set and a test set;

[0006] S2. Input the data in the training set into and train a fusion positioning model. The fusion positioning model is a time series network that uses magnetic sensing signals and visual sensing signals for dual memory update, including a magnetic sensor, a visual sensor, and a multi-source locator;

[0007] The magnetic sensor is used to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it;

[0008] The visual sensor is used to learn and extract valuable visual positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the visual positioning information to output the current valuable multi-source positioning information;

[0009] The multi-source locator is used to learn the valuable positioning information in the positioning memory unit and the multi-source positioning information based on the magnetic field and vision through a convolutional neural network, and finally output the accurate positioning value of the target at the current moment;

[0010] S3. Use the trained fusion positioning model to position the target to be positioned.

[0011] Optionally, the construction of the S1 fusion positioning data set specifically includes:

[0012] Set multiple magnetic positioning base stations in the building to transmit magnetic positioning beacon signals containing the position information of the base stations;

[0013] Scan the building based on the 3D point cloud device to construct its 3D point cloud model G;

[0014] Let the communication terminal carrying the magnetic field sensor array and the point cloud scanning device move to different positions in the building along an arbitrary trajectory, and record the magnetic field signals output by the magnetic field sensor array in real time. The data dimension is K×N×M, where K is the number of magnetic field sensors, N is the number of base stations, and M is the number of output parameters of a single sensor. The surrounding point cloud data scanned by the point cloud scanning device has a data dimension of 6×S, where S is the number of point clouds, and 6 represents the three-dimensional spatial coordinates and RGB values of the point cloud. The position data of the communication terminal includes its spatial coordinates (x, y, z) in the building and the angle θ between it and the due north direction;

[0015] After obtaining a sufficient amount of data, the magnetic field signals output by the communication terminal at different times, the surrounding point cloud data, and the position data are made to correspond one by one to form a fusion positioning data set.

[0016] Optionally, the input of the magnetic sensor includes two parts: the magnetic field signal Ct output by the magnetic field sensor array at the current moment; the position information Rt-1 of the communication terminal at the previous moment;

[0017] In the magnetic sensor, operations of two branches are performed. The two branches respectively learn and extract relevant positioning information in the magnetic field signal and the characteristics of the true positioning information at the previous moment from multiple dimensions, and give valuable magnetic field positioning information at the current moment based on the learned characteristics;

[0018] Among them, the operation of the first branch:

[0019] The magnetic field signal Ct is respectively input into two correlation perception layers. In the first correlation perception layer, dimensionality compression is performed on the base station number dimension N of the magnetic field signal Ct with a dimension of K×N×M by taking the mean of all values, obtaining a feature with a dimension of K×1×M. On the sensor output parameter dimension M, the Euclidean distances between each parameter and the other M-1 parameters are calculated respectively, and the M-1 Euclidean distances corresponding to each parameter are sequentially concatenated to the dimension where the mean is located, obtaining a feature Ct10 with a dimension of K×(1 + M - 1)×M, that is, K×M×M;

[0020] In the second correlation-aware layer, dimension compression is performed by taking the mean of all values of the parameter dimension M of the magnetic field signal Ct, resulting in features of dimension K×N×1. On the dimension N of the number of sensor output base stations, the Euclidean distances between each base station and the other N - 1 base stations are calculated respectively, and the N - 1 Euclidean distances corresponding to each base station are sequentially concatenated to the dimension where the mean is located, obtaining features Ct20 of dimension K×N×(1 + N - 1), that is, K×N×N.

[0021] Subsequently, the outputs Ct10 and Ct20 of the two correlation-aware layers are respectively input into two fully connected layers with a kernel of 1024 for feature extraction, obtaining two outputs Ct11 and Ct21 of dimension K×1024 respectively.

[0022] Then, Ct11 and Ct21 are respectively input into a fully connected layer with a kernel of S / 2 for further feature extraction and dimension adjustment, obtaining two outputs Ct12 and Ct22 of dimension K×S / 2 respectively.

[0023] Subsequently, Ct12 and Ct22 are feature concatenated to obtain an output feature Ct3 of dimension K×S.

[0024] Operations of the second branch:

[0025] First, the magnetic field signal Ct of dimension K×N×M is input into a global pooling layer, and mean pooling is performed on the last two-dimensional features, obtaining features Ct4 of dimension K×1.

[0026] Ct4 is matrix-multiplied with the communication terminal position information Rt - 1 of the previous moment of dimension 1×4, obtaining features Ct5 of dimension K×4.

[0027] Then, Ct5 is input into the global pooling layer, obtaining features Ct6 of dimension K×1.

[0028] Then, the feature Ct3 output by the first branch is multiplied by the transpose of the feature Ct6 output by the second branch to obtain features of dimension K×S, and they are input into the sigmoid activation layer to obtain the current output CL0 of the magnetic sensor.

[0029] Optionally, the input of the visual sensor includes three parts: the point cloud data Vt output by the point cloud scanning device at the current moment; the own three-dimensional point cloud model G of the building; the current output CL0 of the magnetic sensor.

[0030] In the visual sensor, first, the current point cloud data Vt is input into the global pooling layer for global pooling operation, obtaining features Vt1 of dimension 1×S.

[0031] Perform k-means clustering on the 3D point cloud model G of the building itself, so that the internal point cloud forms S point cloud clusters, and take the center points of the point cloud clusters as S feature points of G;

[0032] Multiply the feature Vt1 by the S feature points of G respectively to form an S×S-dimensional feature matrix Vt2;

[0033] Multiply the feature Vt1 by the K×S-dimensional output CL0 of the magnetic sensor to obtain an S×K-dimensional feature matrix Vt3;

[0034] At this time, the feature matrix Vt2 includes global point cloud information, and Vt3 includes magnetic positioning information. Then, perform channel dimension splicing on Vt2, Vt3, and Vt to obtain a feature Vt4 with a dimension of S×(S + K + 6);

[0035] Subsequently, input the feature Vt4 into two consecutive fully connected layers to perform relative relationship learning between the current point cloud information and the global point cloud information, and joint optimization of the magnetic positioning information and the visual positioning information, and output a feature Vt6 with a dimension of S×K;

[0036] Finally, input Vt6 into the sigmoid activation layer to obtain the current multi-source positioning information VL0.

[0037] Optionally, the input of the multi-source locator includes two parts: the current positioning memory unit Lt and the current multi-source positioning information VL0 output by the visual sensor. Among them, the dimension of the current positioning memory unit Lt is Q×K, Q is a preset value, which is obtained by multiplying the feature CLt-1 by the current multi-source positioning information VL0 output by the visual sensor. The feature CLt-1 has a dimension of Q×S and is obtained by multiplying the positioning memory unit Lt-1 with a dimension of Q×K at the previous moment by the current output CL0 of the magnetic sensor. The processing process of the multi-source locator is as follows:

[0038] First, perform a fully connected operation with a kernel of K / 2 on the current positioning memory unit Lt with a dimension of Q×K and the multi-source positioning information VL0 with a dimension of S×K respectively to obtain features Dt0 and Dt1 with a dimension of K×K / 2;

[0039] Concatenate the features Dt0 and Dt1 to obtain a feature matrix Dt2 with a dimension of K×K;

[0040] Then, triple-copy Dt2 in the channel dimension and input it into the lightweight backbone network MobileNet-V2 for feature extraction, and output a 1×1280-dimensional vector Dt3;

[0041] Input Dt3 into the local pooling layer with a kernel of 320, that is, take the average value of every 320 values, to obtain the accurate positioning value Rt of the target at the current moment with a dimension of 1×4.

[0042] Optionally, the processing process of the fusion positioning model is as follows:

[0043] The positioning memory unit Lt-1 with dimension Q×K at the previous moment is multiplied by the current output CL0 of the magnetic sensor to obtain a feature CLt-1 with dimension Q×S, so that the positioning memory unit updates its own state information according to the current magnetic sensing signal;

[0044] Then, the feature CLt-1 is multiplied by the current output VL0 of the visual sensor to obtain the current positioning memory unit Lt with dimension Q×K, so that the positioning memory unit updates its own state information according to the current visual sensing signal, which contains valuable positioning information learned by the network so far;

[0045] Then, Lt and the visual sensor output VL0 are simultaneously input into the multi-source locator to obtain the accurate positioning value Rt of the target at the current moment.

[0046] Optionally, the training process of the fusion positioning model is as follows:

[0047] First, the overall fusion positioning model is trained based on the fusion positioning dataset;

[0048] After the training converges relatively, other internal parameters are fixed, and only the lightweight backbone network MobileNet-V2 in the multi-source locator is trained;

[0049] Finally, the parameter fixation is cancelled, and then the overall fusion positioning model is trained based on the fusion positioning dataset. After the training converges, the final fusion positioning model is obtained.

[0050] On the other hand, a fusion positioning system based on magnetic sensors and vision is provided. The system includes:

[0051] A construction module for constructing a fusion positioning dataset. The data in the dataset includes: the own three-dimensional point cloud model of the building, the magnetic field signals output by the magnetic field sensor array at multiple moments, the surrounding point cloud data scanned by the point cloud scanning device at multiple moments, and the position information of the communication terminal at multiple moments. The dataset is divided into a training set and a test set;

[0052] A training module for inputting and training the fusion positioning model with the data in the training set. The fusion positioning model is a time series network that uses magnetic sensing signals and visual sensing signals for dual memory update, and includes a magnetic sensor, a visual sensor, and a multi-source locator;

[0053] The magnetic sensor is used to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it;

[0054] The visual sensor is used to learn and extract valuable visual positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the visual positioning information to output the current valuable multi-source positioning information;

[0055] The multi-source locator is used to learn the valuable positioning information in the positioning memory unit and the multi-source positioning information based on the magnetic field and vision through a convolutional neural network, and finally output the accurate positioning value of the target at the current moment;

[0056] The positioning module is used to use the trained fusion positioning model to position the target to be positioned.

[0057] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned fusion positioning method based on the magnetic sensor and vision.

[0058] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned fusion positioning method based on the magnetic sensor and vision.

[0059] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0060] 1. The core technology of the fusion positioning method based on the magnetic sensor and vision of the present invention is the fusion positioning model. Compared with other traditional methods, it realizes the real-time and accurate positioning of the target by fusing two sensing information, namely the magnetic positioning beacon signal and the point cloud data, and can avoid the problems of low accuracy of the single magnetic sensor positioning method and high calculation cost of the single vision positioning method, and significantly improve the target positioning accuracy and efficiency.

[0061] 2. The fusion positioning model of the present invention is a time series network that uses magnetic sensing signals and visual sensing signals for dual memory update. Among them, the magnetic sensor accurately captures the valuable positioning information in the magnetic positioning beacon signal through multiple correlation perceptions. The visual sensor is guided by the interaction between the magnetic sensing positioning information and the global point cloud model to quickly realize the extraction of target multi-source positioning information. The multi-source locator mines the multi-source positioning information based on a convolutional neural network and accurately outputs the current position of the target. This structure can simultaneously learn the valuable positioning information in the magnetic sensor and the visual sensor, and guide the output of the current positioning value based on historical positioning knowledge, so as to achieve more accurate and efficient target positioning than a single sensor. Description of the Drawings

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0063] Figure 1 is a flowchart of a fusion positioning method based on a magnetic sensor and vision provided by an embodiment of the present invention;

[0064] Figure 2 is an overall block diagram of a fusion positioning method based on a magnetic sensor and vision provided by an embodiment of the present invention;

[0065] Figure 3 is a schematic structural diagram of a magnetic sensor provided by an embodiment of the present invention;

[0066] Figure 4 is a schematic structural diagram of a vision sensor provided by an embodiment of the present invention;

[0067] Figure 5 is a schematic structural diagram of a multi-source locator provided by an embodiment of the present invention;

[0068] Figure 6 is a block diagram of a fusion positioning system based on a magnetic sensor and vision provided by an embodiment of the present invention;

[0069] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0070] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments.

[0071] An embodiment of the present invention provides a fusion positioning method based on a magnetic sensor and vision. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 As shown in the flowchart of this method, Figure 2 As shown in the overall block diagram of this method, the processing flow can include the following steps:

[0072] S1. Construct a fusion positioning data set. The data in the data set includes: the three-dimensional point cloud model of the building itself, the magnetic field signals output by the magnetic field sensor array at multiple moments, the surrounding point cloud data scanned by the point cloud scanning device at multiple moments, and the position information of the communication terminal at multiple moments. Divide the data set into a training set and a test set;

[0073] Optionally, the construction of the fusion positioning data set in S1 specifically includes:

[0074] Set up multiple magnetic positioning base stations inside the building to transmit magnetic positioning beacon signals containing the location information of the base stations;

[0075] Scan the building based on a three-dimensional point cloud device to construct its three-dimensional point cloud model G;

[0076] Let a communication terminal equipped with a magnetic field sensor array and a point cloud scanning device move to different positions inside the building along an arbitrary trajectory, and record in real time the magnetic field signals output by the magnetic field sensor array. The data dimension is K×N×M, where K is the number of magnetic field sensors, N is the number of base stations, and M is the number of output parameters of a single sensor. The surrounding point cloud data scanned by the point cloud scanning device has a data dimension of 6×S, where S is the number of point clouds, and 6 represents the three-dimensional spatial coordinates and RGB values of the point cloud. The position data of the communication terminal includes its spatial coordinates (x, y, z) inside the building and the angle θ between it and the due north direction;

[0077] After obtaining a sufficient amount of data, match the magnetic field signals output by the communication terminal at different times, the surrounding point cloud data, and the position data one by one to form a fused positioning data set.

[0078] S2. Input the data in the training set into and train a fused positioning model. The fused positioning model is a temporal network that uses magnetic sensing signals and visual sensing signals for dual memory update, including a magnetic sensor, a visual sensor, and a multi-source locator;

[0079] The magnetic sensor is used to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it;

[0080] The visual sensor is used to learn and extract valuable visual positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the visual positioning information to output the current valuable multi-source positioning information;

[0081] The multi-source locator is used to learn the valuable positioning information in the positioning memory unit and the multi-source positioning information based on magnetic field and vision through a convolutional neural network, and finally output the accurate positioning value of the target at the current moment;

[0082] Optionally, as Figure 3 shown, the input of the magnetic sensor includes two parts: the magnetic field signal Ct output by the magnetic field sensor array at the current moment; the position information Rt-1 of the communication terminal at the previous moment;

[0083] In the magnetic sensor, two branches of operations are performed. The two branches respectively learn from multiple dimensions and extract relevant positioning information and the characteristics of the true positioning information at the previous moment from the magnetic field signal, and give valuable current magnetic field positioning information based on the learned characteristics.

[0084] Among them, the operation of the first branch:

[0085] The magnetic field signal Ct is respectively input into two correlation perception layers. In the first correlation perception layer, dimensionality compression is performed by taking the mean of all values in the base station number dimension N of the magnetic field signal Ct with dimensions K×N×M, obtaining a feature with dimensions K×1×M. On the sensor output parameter dimension M, the Euclidean distances between each parameter and the other M - 1 parameters are calculated respectively, and the M - 1 Euclidean distances corresponding to each parameter are sequentially concatenated to the dimension where the mean is located, obtaining a feature Ct10 with dimensions K×(1 + M - 1)×M, that is, K×M×M.

[0086] In the second correlation perception layer, dimensionality compression is performed by taking the mean of all values in the parameter dimension M of the magnetic field signal Ct, obtaining a feature with dimensions K×N×1. On the sensor output base station number dimension N, the Euclidean distances between each base station and the other N - 1 base stations are calculated respectively, and the N - 1 Euclidean distances corresponding to each base station are sequentially concatenated to the dimension where the mean is located, obtaining a feature Ct20 with dimensions K×N×(1 + N - 1), that is, K×N×N.

[0087] Subsequently, the outputs Ct10 and Ct20 of the two correlation perception layers are respectively input into two fully connected layers with a kernel of 1024 for feature extraction, and two outputs Ct11 and Ct21 with dimensions K×1024 are respectively obtained.

[0088] Then Ct11 and Ct21 are respectively input into a fully connected layer with a kernel of S / 2 for further feature extraction and dimensionality adjustment, and two outputs Ct12 and Ct22 with dimensions K×S / 2 are respectively obtained.

[0089] Subsequently, Ct12 and Ct22 are feature - concatenated to obtain an output feature Ct3 with dimensions K×S.

[0090] The operation of the second branch:

[0091] First, the magnetic field signal Ct with dimensions K×N×M is input into a global pooling layer, and average pooling is performed on the last two - dimensional features, obtaining a feature Ct4 with dimensions K×1.

[0092] Ct4 is matrix - multiplied with the communication terminal position information Rt - 1 with dimensions 1×4 to obtain a feature Ct5 with dimensions K×4.

[0093] Then, Ct5 is input into the global pooling layer to obtain a feature Ct6 with a dimension of K×1;

[0094] Then, the feature Ct3 output by the first branch is multiplied by the transpose of the feature Ct6 output by the second branch to obtain a feature with a dimension of K×S, and it is input into the sigmoid activation layer to obtain the current output CL0 of the magnetic sensor.

[0095] Optionally, as Figure 4 shown, the input of the visual sensor includes three parts: the point cloud data Vt output by the point cloud scanning device at the current moment; the own three-dimensional point cloud model G of the building; the current output CL0 of the magnetic sensor;

[0096] In the visual sensor, first, the current point cloud data Vt is input into the global pooling layer for global pooling operation to obtain a feature Vt1 with a dimension of 1×S;

[0097] Then, the k-means clustering operation is performed on the own three-dimensional point cloud model G of the building to form S point cloud clusters for the internal point clouds, and the center points of the point cloud clusters are taken as the S feature points of G;

[0098] The feature Vt1 is multiplied by the S feature points of G respectively to form an S×S-dimensional feature matrix Vt2;

[0099] The feature Vt1 is multiplied by the K×S-dimensional output CL0 of the magnetic sensor to obtain an S×K-dimensional feature matrix Vt3;

[0100] At this time, the feature matrix Vt2 includes the global point cloud information, and Vt3 includes the magnetic positioning information. Then, Vt2, Vt3, and Vt are concatenated in the channel dimension to obtain a feature Vt4 with a dimension of S×(S+K+6);

[0101] Subsequently, the feature Vt4 is input into two consecutive fully connected layers to perform the relative relationship learning of the current point cloud information and the global point cloud information, and the joint optimization of the magnetic positioning information and the visual positioning information, and output a feature Vt6 with a dimension of S×K;

[0102] Finally, Vt6 is input into the sigmoid activation layer to obtain the current multi-source positioning information VL0.

[0103] Optionally, as Figure 5 shown, the input of the multi-source locator includes two parts: the current positioning memory unit Lt and the current multi-source positioning information VL0 output by the visual sensor, where as Figure 2As shown, the dimension of the current positioning memory unit Lt is Q×K. Q is a preset value, which is obtained by multiplying the feature CLt-1 and the current multi-source positioning information VL0 output by the visual sensor in matrix form. The feature CLt-1, with a dimension of Q×S, is obtained by multiplying the positioning memory unit Lt-1 with a dimension of Q×K at the previous moment and the current output CL0 of the magnetic sensor in matrix form (when t = 1, Lt-1 does not exist, that is, there is no positioning memory unit in the initial state, CLt-1 is the CL0 output by the magnetic sensor when t = 1, and Lt is obtained by multiplying the feature CLt-1 and the multi-source positioning information VL0 output by the visual sensor when t = 1). The processing process of the multi-source locator is as follows:

[0104] First, perform a fully connected operation with a kernel of K / 2 on the current positioning memory unit Lt with a dimension of Q×K and the multi-source positioning information VL0 with a dimension of S×K respectively, to obtain features Dt0 and Dt1 with a dimension of K×K / 2;

[0105] Concatenate the features Dt0 and Dt1 to obtain a feature matrix Dt2 with a dimension of K×K;

[0106] Then, triple-copy Dt2 in the channel dimension and input it into the lightweight backbone network MobileNet-V2 for feature extraction, and output a 1×1280-dimensional vector Dt3;

[0107] Input Dt3 into a local pooling layer with a kernel of 320, that is, take the average value of every 320 values, to obtain an accurate positioning value Rt of the target at the current moment.

[0108] Optionally, as Figure 2 shown, the processing process of the fusion positioning model is as follows:

[0109] Multiply the positioning memory unit Lt-1 with a dimension of Q×K at the previous moment and the current output CL0 of the magnetic sensor in matrix form to obtain a feature CLt-1 with a dimension of Q×S, so that the positioning memory unit updates its own state information according to the current magnetic sensing signal;

[0110] Then multiply the feature CLt-1 and the current output VL0 of the visual sensor in matrix form to obtain the current positioning memory unit Lt with a dimension of Q×K, so that the positioning memory unit updates its own state information according to the current visual sensing signal, including the valuable positioning information learned by the network so far;

[0111] Then input Lt and the visual sensor output VL0 into the multi-source locator at the same time to obtain an accurate positioning value Rt of the target at the current moment.

[0112] Compared with traditional time-series neural networks, the fusion positioning model of the embodiments of the present invention can simultaneously learn and update the dual positioning information of magnetic fields and vision, and fuse the correlation relationships between different base stations and sensor parameters multiple times in the magnetic sensor, and guide the vision to quickly and accurately locate based on the global point cloud model and magnetic sensing information in the vision sensor, so as to achieve more accurate and rapid target positioning.

[0113] Optionally, the training process of the fusion positioning model is as follows:

[0114] First, perform overall training on the fusion positioning model based on the fusion positioning data set;

[0115] After the training converges relatively, fix other internal parameters and only perform training on the lightweight backbone network MobileNet-V2 in the multi-source locator;

[0116] Finally, cancel the parameter fixation, and then perform overall training on the fusion positioning model based on the fusion positioning data set. After the training converges, obtain the final fusion positioning model.

[0117] S3. Use the trained fusion positioning model to locate the target to be located.

[0118] As Figure 6 shown, the embodiments of the present invention also provide a fusion positioning system based on magnetic sensors and vision. The system includes:

[0119] A construction module 610, configured to construct a fusion positioning data set. The data in the data set includes: the self three-dimensional point cloud model of a building, the magnetic field signals output by a magnetic field sensor array at multiple moments, the surrounding point cloud data scanned by a point cloud scanning device at multiple moments, and the position information of communication terminals at multiple moments. Divide the data set into a training set and a test set;

[0120] A training module 620, configured to input the data in the training set and train a fusion positioning model. The fusion positioning model is a time-series network that uses magnetic sensing signals and vision sensing signals to perform dual memory updates, and includes a magnetic sensor, a vision sensor, and a multi-source locator;

[0121] The magnetic sensor is configured to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it;

[0122] The vision sensor is configured to learn and extract valuable vision positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the vision positioning information to output the current valuable multi-source positioning information;

[0123] The multi-source locator is used to learn valuable positioning information in the positioning memory unit and the multi-source positioning information based on magnetic field and vision through a convolutional neural network, and finally output an accurate positioning value of the target at the current moment;

[0124] The positioning module 630 is used to use the trained fusion positioning model to position the target to be positioned.

[0125] The functional structure of a fusion positioning system based on a magnetic sensor and vision provided by an embodiment of the present invention corresponds to a fusion positioning method based on a magnetic sensor and vision provided by an embodiment of the present invention, and will not be elaborated here.

[0126] Figure 7 FIG. 10 is a schematic structural diagram of an electronic device 700 provided by an embodiment of the present invention. The electronic device 700 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 701 and one or more memories 702. Among them, at least one instruction is stored in the memory 702, and the at least one instruction is loaded and executed by the processor 701 to implement the steps of the above-mentioned fusion positioning method based on a magnetic sensor and vision.

[0127] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above-mentioned fusion positioning method based on a magnetic sensor and vision. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0128] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0129] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A fusion positioning method based on magnetic sensors and vision, characterized in that The method includes: S1. Construct a fused positioning data set, where the data in the data set includes: the three-dimensional point cloud model of the building itself, the magnetic field signals output by the magnetic field sensor array at multiple moments, the surrounding point cloud data scanned by the point cloud scanning device at multiple moments, and the position information of the communication terminal at multiple moments. Divide the data set into a training set and a test set; S2. Input the data in the training set into and train a fused positioning model. The fused positioning model is a time series network that uses magnetic sensing signals and visual sensing signals for dual memory update, and includes a magnetic sensor, a visual sensor, and a multi-source locator; The magnetic sensor is used to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it; The visual sensor is used to learn and extract valuable visual positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the visual positioning information to output the current valuable multi-source positioning information; The multi-source locator is used to learn the valuable positioning information in the positioning memory unit and the multi-source positioning information based on the magnetic field and vision through a convolutional neural network, and finally output the accurate positioning value of the target at the current moment; S3. Use the trained fused positioning model to position the target to be positioned.

2. The method according to claim 1, characterized in that, The construction of the fused positioning data set in S1 specifically includes: Set multiple magnetic positioning base stations in the building to transmit magnetic positioning beacon signals containing the position information of the base stations; Scan the building based on a three-dimensional point cloud device to construct its three-dimensional point cloud model G; Let the communication terminal carrying the magnetic field sensor array and the point cloud scanning device move to different positions in the building along an arbitrary trajectory, and record the magnetic field signals output by the magnetic field sensor array in real time. The data dimension is K×N×M, where K is the number of magnetic field sensors, N is the number of base stations, and M is the number of output parameters of a single sensor. The surrounding point cloud data scanned by the point cloud scanning device has a data dimension of 6×S, where S is the number of point clouds, and 6 represents the three-dimensional spatial coordinates and RGB values of the point cloud. The position data of the communication terminal includes its spatial coordinates (x, y, z) in the building and its angle θ with the due north direction; After obtaining a sufficient amount of data, correspond the magnetic field signal outputs, the surrounding point cloud data, and the position data received by the communication terminal at different moments one by one to form a fused positioning data set.

3. The method according to claim 2, characterized in that The input of the magnetic sensor includes two parts: the magnetic field signal Ct output by the magnetic field sensor array at the current moment; the position information Rt-1 of the communication terminal at the previous moment; In the magnetic sensor, operations of two branches are performed. The two branches respectively learn and extract the relevant positioning information in the magnetic field signal and the real positioning information feature at the previous moment from multiple dimensions, and give the current valuable magnetic field positioning information based on the learned features; Among them, the operation of the first branch: The magnetic field signal Ct is respectively input into two correlation perception layers. In the first correlation perception layer, dimension compression is performed by taking the mean of all values in the base station number dimension N of the magnetic field signal Ct with dimension K×N×M, obtaining a feature with dimension K×1×M. On the sensor output parameter dimension M, the Euclidean distances between each parameter and the other M - 1 parameters are calculated respectively, and the M - 1 Euclidean distances corresponding to each parameter are sequentially concatenated to the dimension where the mean is located, obtaining a feature Ct10 with dimension K×(1 + M - 1)×M, that is, K×M×M; In the second correlation perception layer, dimension compression is performed by taking the mean of all values in the parameter dimension M of the magnetic field signal Ct, obtaining a feature with dimension K×N×1. On the sensor output base station number dimension N, the Euclidean distances between each base station and the other N - 1 base stations are calculated respectively, and the N - 1 Euclidean distances corresponding to each base station are sequentially concatenated to the dimension where the mean is located, obtaining a feature Ct20 with dimension K×N×(1 + N - 1), that is, K×N×N; Subsequently, the outputs Ct10 and Ct20 of the two correlation perception layers are respectively input into two fully - connected layers with a kernel of 1024 for feature extraction, obtaining two outputs Ct11 and Ct21 with dimension K×1024 respectively; Then Ct11 and Ct21 are respectively input into a fully - connected layer with a kernel of S / 2 for further feature extraction and dimension adjustment, obtaining two outputs Ct12 and Ct22 with dimension K×S / 2 respectively; Subsequently, Ct12 and Ct22 are concatenated in terms of features, obtaining an output feature Ct3 with dimension K×S; Operations of the second branch: First, the magnetic field signal Ct with dimension K×N×M is input into a global pooling layer, and mean pooling is performed on the last two - dimensional features, obtaining a feature Ct4 with dimension K×1; Ct4 is multiplied by the communication terminal position information Rt - 1 of the previous moment with dimension 1×4, obtaining a feature Ct5 with dimension K×4; Then Ct5 is input into the global pooling layer, obtaining a feature Ct6 with dimension K×1; Then the feature Ct3 output by the first branch is multiplied by the transpose of the feature Ct6 output by the second branch, obtaining a K×S - dimensional feature, and it is input into the sigmoid activation layer to obtain the current output CL0 of the magnetic sensor.

4. The method according to claim 3, characterized in that The input of the visual sensor includes three parts: the point cloud data Vt output by the point cloud scanning device at the current moment; the own three - dimensional point cloud model G of the building; the current output CL0 of the magnetic sensor; In the visual sensor, first, the current point cloud data Vt is input into the global pooling layer for global pooling operation, obtaining a feature Vt1 with dimension 1×S; Then, k - means clustering operation is performed on the own three - dimensional point cloud model G of the building, so that the internal point clouds form S point cloud clusters, and the center points of the point cloud clusters are taken as the S feature points of G; The feature Vt1 is multiplied by the S feature points of G respectively, forming an S×S - dimensional feature matrix Vt2; The feature Vt1 is multiplied by the K×S - dimensional output CL0 of the magnetic sensor, obtaining an S×K - dimensional feature matrix Vt3; At this time, the feature matrix Vt2 includes global point cloud information, and Vt3 includes magnetic positioning information. Then, Vt2, Vt3, and Vt are concatenated along the channel dimension to obtain a feature Vt4 with a dimension of S×(S + K + 6). Subsequently, the feature Vt4 is input into two consecutive fully connected layers to learn the relative relationship between the current point cloud information and the global point cloud information, and jointly optimize the magnetic positioning information and the visual positioning information, and output a feature Vt6 with a dimension of S×K. Finally, Vt6 is input into the sigmoid activation layer to obtain the current multi-source positioning information VL0.

5. The method according to claim 4, wherein The input of the multi-source locator includes two parts: the current positioning memory unit Lt and the current multi-source positioning information VL0 output by the visual sensor. The dimension of the current positioning memory unit Lt is Q×K, where Q is a preset value, which is obtained by multiplying the feature CLt-1 and the current multi-source positioning information VL0 output by the visual sensor. The dimension of the feature CLt-1 is Q×S, which is obtained by multiplying the positioning memory unit Lt-1 with a dimension of Q×K at the previous moment and the current output CL0 of the magnetic sensor. The processing process of the multi-source locator is as follows: First, a fully connected operation with a kernel of K / 2 is performed on the current positioning memory unit Lt with a dimension of Q×K and the multi-source positioning information VL0 with a dimension of S×K respectively to obtain features Dt0 and Dt1 with a dimension of K×K / 2. The features Dt0 and Dt1 are concatenated to obtain a feature matrix Dt2 with a dimension of K×K. Then, Dt2 is replicated three times along the channel dimension and input into the lightweight backbone network MobileNet-V2 for feature extraction, and a 1×1280-dimensional vector Dt3 is output. Dt3 is input into a local pooling layer with a kernel of 320, that is, the average value is taken for every 320 values, and an accurate positioning value Rt of the target at the current moment is obtained.

6. The method according to claim 5, wherein The processing process of the fusion positioning model is as follows: The positioning memory unit Lt-1 with a dimension of Q×K at the previous moment is multiplied by the current output CL0 of the magnetic sensor to obtain a feature CLt-1 with a dimension of Q×S, so that the positioning memory unit updates its own state information according to the current magnetic sensing signal. Then, the feature CLt-1 is multiplied by the current output VL0 of the visual sensor to obtain the current positioning memory unit Lt with a dimension of Q×K, so that the positioning memory unit updates its own state information according to the current visual sensing signal, which contains valuable positioning information learned by the network so far. Then, Lt and the visual sensor output VL0 are simultaneously input into the multi-source locator to obtain an accurate positioning value Rt of the target at the current moment.

7. The method according to claim 6, wherein The training process of the fusion positioning model is as follows: First, the overall fusion positioning model is trained based on the fusion positioning dataset. After the training converges relatively, other internal parameters are fixed, and only the lightweight backbone network MobileNet-V2 in the multi-source locator is trained. Finally, the parameter fixation is cancelled, and the overall fusion positioning model is trained again based on the fusion positioning dataset. After the training converges, the final fusion positioning model is obtained.

8. A fusion positioning system based on magnetic sensors and vision, characterized in that, The system includes: A building block for constructing a fused positioning dataset, where the data in the dataset includes: the three-dimensional point cloud model of the building itself, the magnetic field signals output by the magnetic field sensor array at multiple times, the surrounding point cloud data scanned by the point cloud scanning device at multiple times, and the location information of the communication terminal at multiple times. The dataset is divided into a training set and a test set; A training module for inputting and training a fused positioning model with the data in the training set. The fused positioning model is a time series network that uses magnetic sensing signals and visual sensing signals for dual memory updates, and includes a magnetic sensor, a visual sensor, and a multi-source locator; The magnetic sensor is used to learn and extract valuable magnetic field positioning information from the magnetic field signals output by the magnetic field sensor array and output it; The visual sensor is used to learn and extract valuable visual positioning information from the surrounding point cloud data scanned by the point cloud scanning device, and fuse the magnetic field positioning information output by the magnetic sensor with the visual positioning information to output the current valuable multi-source positioning information; The multi-source locator is used to learn the valuable positioning information in the positioning memory unit and the multi-source positioning information based on the magnetic field and vision through a convolutional neural network, and finally output the accurate positioning value of the target at the current moment; A positioning module for using the trained fused positioning model to position the target to be positioned.

9. An electronic device, the electronic device includes a processor and a memory, and at least one instruction is stored in the memory, characterized in that, The at least one instruction is loaded and executed by the processor to implement the magnetic sensor and vision-based fused positioning method according to any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that, The at least one instruction is loaded and executed by the processor to implement the magnetic sensor and vision-based fused positioning method according to any one of claims 1-7.