Multi-modal Transform-based UWB multi-sensor fusion positioning method and system
By introducing a multimodal Transformer model and dynamic weight allocation mechanism into positioning technology, combined with an adaptive correction strategy, the problem of multi-sensor data fusion in complex environments is solved, and a high-precision and reliable positioning effect is achieved.
Patent Information
- Application Number
- CN202510296931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing positioning technologies are difficult to achieve efficient and accurate multi-sensor data fusion in complex environments and dynamic scenarios, resulting in reduced positioning accuracy and reduced reliability.
UWB multi-sensor fusion positioning method based on multimodal Transformer is adopted, and the multimodal data set is feature extraction and fusion through the multimodal Transformer model, combining dynamic weight allocation mechanism and adaptive correction strategy to achieve high-precision positioning of the target position.
High-precision and reliable multi-sensor fusion positioning is achieved in complex environments, improving positioning efficiency and adaptability in multi-source heterogeneous data scenarios, and having strong real-time and anti-interference capabilities.
Smart Images

Figure CN120101803A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of positioning and navigation technology, and in particular to a UWB multi-sensor fusion positioning method and system based on a multi-modal Transformer. Background Art
[0002] In the context of the rapid development of modern Internet of Things and intelligent technology, high-precision indoor and outdoor positioning technology plays an important role in various fields, especially in the fields of intelligent transportation, intelligent manufacturing, smart medical care and smart security. Existing positioning technologies mainly include global navigation satellite system (GNSS), inertial navigation system (INS), radio frequency identification (RFID), Bluetooth and ultra-wideband (UWB). Among them, UWB technology has become the preferred solution for high-precision indoor and outdoor positioning due to its advantages of high precision, low power consumption and high security. However, the single UWB positioning method still has some limitations in practical applications, such as multipath effect, occlusion problem and dynamic environmental changes, which will lead to reduced positioning accuracy and reliability.
[0003] In addition, existing multi-sensor fusion positioning methods usually rely on traditional Kalman filtering, particle filtering and other algorithms, which have the disadvantages of complex models, large amount of calculation and poor real-time performance when processing multi-modal data. Especially in complex environments and dynamic scenes, traditional methods are difficult to achieve efficient and accurate multi-sensor data fusion, which affects the overall positioning performance. Therefore, it is of great significance to develop a multi-sensor fusion positioning method that can effectively process multi-modal data, has strong real-time performance and high accuracy.
[0004] The present invention aims to solve the deficiencies of the above-mentioned prior art and proposes a UWB multi-sensor fusion positioning method and system based on multimodal Transformer. By introducing the Transformer model, its advantages in multimodal data processing and feature extraction are fully utilized to achieve high-precision and high-reliability multi-sensor fusion positioning. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a UWB multi-sensor fusion positioning method and system based on a multimodal Transformer, which solves the problems of the prior art positioning technology, such as insufficient accuracy, poor adaptability to complex environments, and inability to effectively fuse multi-source heterogeneous data.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a UWB multi-sensor fusion positioning method based on a multi-modal Transformer, comprising the following steps:
[0007] By deploying multiple UWB sensors and auxiliary sensors, the signal characteristics of the target at different locations are collected to build a multimodal data set;
[0008] Using a multimodal Transformer model to extract and fuse features of the multimodal dataset to generate a global feature representation of the target;
[0009] According to the global feature representation, combined with a dynamic weight allocation mechanism, the location information of the target is predicted, and adaptive correction is performed for abnormal signals.
[0010] Preferably, the multimodal data set includes distance and angle information obtained from the UWB sensor, and acceleration, gyroscope and magnetic field strength information obtained from the auxiliary sensor;
[0011] The multimodal Transformer model includes mapping the data of each sensor into a unified feature space and capturing the correlation between different modalities through a self-attention mechanism;
[0012] The dynamic weight allocation mechanism includes dynamically adjusting the weight of each modality data in the global feature representation according to its confidence.
[0013] Preferably, the feature extraction and fusion includes extracting local features of each modality data through a multi-head self-attention module, and realizing interactive fusion of different modality data through a cross-modality attention module;
[0014] During the fusion process, if the confidence level of a certain modality data is lower than the preset threshold, it is preliminarily judged that the modality data is abnormal.
[0015] Preferably, the feature extraction and fusion also includes analyzing the confidence distribution of each abnormal modal data. When the abnormal modal data is only a single dimension, it is repaired through correlation with other modal data, and the current positioning process is treated as a normal state; when the abnormal modal data is multi-dimensional, the current positioning process is treated as an abnormal state.
[0016] Preferably, the adaptive correction includes constructing trajectory smoothing constraints based on historical trajectory information of the target;
[0017] Assume that there are n historical trajectory points, take each trajectory point as a node, and use graph neural network to model the spatiotemporal relationship between nodes;
[0018] The local trajectory features are extracted through convolutional neural networks, and the time series characteristics of the target movement are captured using recurrent neural networks, and the output is a feature vector;
[0019] The fully connected layer reduces the dimension of the feature vector to a d-dimensional real space and generates a compact trajectory representation vector;
[0020] The trajectory consistency score is defined by combining the projection error and direction consistency. If the score is less than the consistency threshold, all modal data are taken as normal input and positioning continues. If the score is not less than the consistency threshold, it is determined that there is abnormal interference in the modal data and selective correction is performed.
[0021] Preferably, the UWB multi-sensor fusion positioning system based on multi-modal Transformer is characterized by comprising:
[0022] The acquisition unit deploys multiple UWB sensors and auxiliary sensors to collect signal characteristics of the target at different locations and construct a multimodal data set;
[0023] A fusion unit, which uses a multimodal Transformer model to perform feature extraction and fusion on the multimodal data set to generate a global feature representation of the target;
[0024] The positioning unit predicts the location information of the target according to the global feature representation and combines the dynamic weight allocation mechanism, and performs adaptive correction for abnormal signals.
[0025] Preferably, a computer device comprises a memory and a processor,
[0026] The memory stores a computer program, and when the processor executes the computer program, the steps of the UWB multi-sensor fusion positioning method based on multi-modal Transformer are implemented.
[0027] Preferably, a computer-readable storage medium is characterized in that:
[0028] A computer program is stored thereon, and when the computer program is executed by a processor, the steps of a UWB multi-sensor fusion positioning method based on a multi-modal Transformer are implemented.
[0029] The present invention provides a UWB multi-sensor fusion positioning method and system based on multi-modal Transformer, which has the following beneficial effects:
[0030] The present invention achieves high-precision positioning of target positions in complex environments by introducing a multimodal Transformer model, combining a dynamic weight allocation mechanism and an adaptive correction strategy. In response to the problem of abnormal signals, the present invention optimizes the robustness of positioning through weight adjustment and trajectory consistency analysis; by using the deep fusion of multimodal data and spatiotemporal relationship modeling, the positioning efficiency and adaptability in multi-source heterogeneous data scenarios are significantly improved, while having strong real-time and anti-interference capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the overall process of the method of the present invention;
[0032] Figure 2 Schematic diagram of the structure of the multimodal Transformer model in the present invention;
[0033] Figure 3 This is a schematic diagram of the implementation process of the dynamic weight allocation mechanism in the present invention;
[0034] Figure 4 It is a schematic diagram of the process of the adaptive correction strategy in the present invention;
[0035] Figure 5 It is a schematic diagram of the system module structure of the present invention;
[0036] Figure 6 Schematic diagram of trajectory path generation and optimization in the present invention. DETAILED DESCRIPTION
[0037] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] The present invention provides a UWB multi-sensor fusion positioning method and system based on multi-modal Transformer, the core of which is to achieve high-precision positioning of the target position in a complex environment through the deep fusion of multi-modal data and dynamic weight allocation mechanism combined with an adaptive correction strategy. Figure 1 To Attachment Figure 6 Specific embodiments of the present invention are described in detail.
[0039] First, if Figure 1 As shown, the overall process of the present invention includes four main steps: data acquisition, feature extraction and fusion, dynamic weight allocation, and adaptive correction. In the data acquisition stage, by deploying multiple UWB sensors and auxiliary sensors, the signal characteristics of the target at different positions are collected to construct a multimodal data set. Specifically, the UWB sensor is used to obtain the distance and angle information of the target, while the auxiliary sensor is responsible for collecting information such as acceleration, gyroscope, and magnetic field strength. These data together constitute a multimodal data set, which provides a basis for subsequent feature extraction and fusion. In practical applications, such as in indoor positioning scenarios, UWB sensors can be arranged in the four corners of the room, and auxiliary sensors are integrated into the target device to ensure the comprehensiveness and real-time nature of data acquisition.
[0040] Next, in the feature extraction and fusion stage, the multimodal Transformer model is used to process the multimodal dataset. Figure 2 As shown in Figure 1, the core structure of the multimodal Transformer model includes a multi-head self-attention module and a cross-modal attention module. The multi-head self-attention module is responsible for extracting local features of each modal data, while the cross-modal attention module is used to realize the interactive fusion between different modal data. Specifically, the data of each sensor is first mapped to a unified feature space, and then the correlation between different modalities is captured through the self-attention mechanism. For example, when the confidence of a certain modal data is lower than the preset threshold, it is preliminarily judged that the modal data is abnormal. At this time, if the abnormal modal data is only a single dimension, it is repaired through the correlation with other modal data, and the current positioning process is treated as a normal state; if the abnormal modal data is multi-dimensional, the current positioning process is treated as an abnormal state. This process ensures that the system can maintain a high positioning accuracy even when some data is abnormal.
[0041] In the implementation of the dynamic weight allocation mechanism, Figure 3 As shown, according to the confidence distribution of each modal data, its weight in the global feature representation is dynamically adjusted. Specifically, first obtain other modal data directly related to the abnormal modal data as the reference modality. Assume the reference modality A, and determine the weight adjustment range A1 of an abnormal modal data according to the correlation between the reference modality A and the target motion state and the interaction relationship with other modal data. Then, according to the interaction relationship between A1 and other modal data, determine the feasible domain a2 of the weight of the next abnormal modal data. In a2, randomly select a value as the weight A2 of the abnormal modal data, and adjust the weight of the next modal data until the weight adjustment of all abnormal modal data is obtained to form a weight adjustment combination. Repeat the operation of randomly selecting weights in the feasible domain until no new weight adjustment combination is generated. Finally, by dynamically analyzing all weight adjustment combinations, any combination is combined with other normal modal data as a positioning state, and trajectory consistency analysis is performed to find the combination with the highest consistency as the optimal result for output. At the same time, the weight adjustment results of abnormal modal data are color-coded in the visualization interface so that users can intuitively understand the weight changes of each modal data.
[0042] The adaptive correction strategy is another important component of the present invention, and its process is as follows: Figure 4As shown. First, based on the historical trajectory information of the target, a trajectory smoothing constraint is constructed. Assuming that there are n historical trajectory points, each trajectory point is taken as a node, and the spatiotemporal relationship between nodes is modeled through a graph neural network. Specifically, in each trajectory path, the consistency between the current node and the next node and the actual motion state of the target is analyzed respectively. The local trajectory features are extracted through a convolutional neural network, and the time series characteristics of the target motion are captured using a recurrent neural network, and the output is a feature vector. The fully connected layer reduces the feature vector to a d-dimensional real space to generate a compact trajectory representation vector. For each trajectory representation vector, its characteristic subspace is determined by principal component analysis, and the basis matrix of the characteristic subspace is B. The trajectory representation vector is projected onto the characteristic subspace of the trajectory path, and the projection error is calculated to quantify whether the characteristic vector belongs to the trajectory characteristic subspace. At the same time, the consistency of the direction of the projection vector with the target vector is analyzed, and the trajectory consistency score is defined by combining the projection error and direction consistency. The formula is as follows: , where e is the projection error, is the maximum allowable value of the projection error, is the directional consistency between the projection vector and the target vector, and is the weight parameter. If it is less than the consistency threshold, all modal data are taken as normal input and positioning continues; If it is not less than the consistency threshold, it is determined that there is abnormal interference in the modal data, and selective correction is performed to retain only the positioning results of normal modal data. After the correction is completed, the position of the trajectory node where the positioning result is located in the current node and the next node is determined. If the positioning result is the next node, the trajectory path is updated, and the original next node is used as the current node to continue positioning. Among the n trajectory paths, the paths that do not conform to the selected node order are eliminated through the selection of nodes by the target motion, and nm paths that conform to the current motion trajectory are obtained, where m represents the number of eliminated paths, and m < n.
[0043] In addition, if Figure 5 As shown, the present invention also provides a UWB multi-sensor fusion positioning system based on a multimodal Transformer, including an acquisition unit, a fusion unit and a positioning unit. The acquisition unit is responsible for collecting signal features of the target at different positions by deploying multiple UWB sensors and auxiliary sensors to construct a multimodal data set. The fusion unit uses a multimodal Transformer model to extract and fuse features of the multimodal data set to generate a global feature representation of the target. The positioning unit predicts the location information of the target based on the global feature representation and combines the dynamic weight allocation mechanism, and performs adaptive correction for abnormal signals. The modular design of the system enables each functional unit to operate independently, and is also easy to expand and maintain.
[0044] Finally, if Figure 6 As shown, the present invention also includes a process of trajectory path generation and optimization. Specifically, each key point in the target motion trajectory is defined and used as a node in the path, and the nodes are connected in series to obtain a trajectory path after series connection. If the order of nodes in the path is replaceable, different node orders are connected in series according to the replaceable positions to generate all trajectory paths. Construct a trajectory path corresponding to each motion mode, and automatically match the trajectory path after series connection after selecting the motion mode. This process not only improves the flexibility of the trajectory path, but also enhances the system's adaptability to complex motion modes.
[0045] In summary, the present invention achieves high-precision positioning of target positions in complex environments by introducing a multimodal Transformer model, combined with a dynamic weight allocation mechanism and an adaptive correction strategy. In response to the problem of abnormal signals, the present invention optimizes the robustness of positioning through weight adjustment and trajectory consistency analysis; by using the deep fusion of multimodal data and spatiotemporal relationship modeling, the positioning efficiency and adaptability in multi-source heterogeneous data scenarios are significantly improved, while having strong real-time and anti-interference capabilities.
[0046] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. The UWB multi-sensor fusion positioning method based on multimodal Transformer is characterized by: The following steps are involved: By deploying multiple UWB sensors and auxiliary sensors, the signal characteristics of the target at different locations are collected to build a multimodal data set; Using a multimodal Transformer model to extract and fuse features of the multimodal dataset to generate a global feature representation of the target; According to the global feature representation, combined with a dynamic weight allocation mechanism, the location information of the target is predicted, and adaptive correction is performed for abnormal signals.
2. The UWB multi-sensor fusion positioning method based on multimodal Transformer according to claim 1 is characterized in that: The multimodal data set includes distance and angle information obtained from the UWB sensor, and acceleration, gyroscope, and magnetic field strength information obtained from auxiliary sensors; The multimodal Transformer model includes mapping the data of each sensor into a unified feature space and capturing the correlation between different modalities through a self-attention mechanism; The dynamic weight allocation mechanism includes dynamically adjusting the weight of each modality data in the global feature representation according to its confidence.
3. The UWB multi-sensor fusion positioning method based on multi-modal Transformer according to claim 1 is characterized in that: The feature extraction and fusion includes extracting local features of each modality data through a multi-head self-attention module, and realizing interactive fusion of different modality data through a cross-modality attention module; During the fusion process, if the confidence level of a certain modality data is lower than the preset threshold, it is preliminarily judged that the modality data is abnormal.
4. The UWB multi-sensor fusion positioning method based on multi-modal Transformer according to claim 3 is characterized in that: The feature extraction and fusion also includes analyzing the confidence distribution of each abnormal modal data. When the abnormal modal data is only a single dimension, it is repaired through the correlation with other modal data, and the current positioning process is treated as a normal state; when the abnormal modal data is multi-dimensional, the current positioning process is treated as an abnormal state.
5. The UWB multi-sensor fusion positioning method based on multi-modal Transformer according to claim 1, characterized in that: The adaptive correction includes constructing trajectory smoothing constraints based on historical trajectory information of the target; Assume that there are n historical trajectory points, take each trajectory point as a node, and use graph neural network to model the spatiotemporal relationship between nodes; The local trajectory features are extracted through convolutional neural networks, and the time series characteristics of the target movement are captured using recurrent neural networks, and the output is a feature vector; The fully connected layer reduces the dimension of the feature vector to a d-dimensional real space and generates a compact trajectory representation vector; The trajectory consistency score is defined by combining the projection error and direction consistency. If the score is less than the consistency threshold, all modal data are taken as normal input and positioning continues. If the score is not less than the consistency threshold, it is determined that there is abnormal interference in the modal data and selective correction is performed.
6. UWB multi-sensor fusion positioning system based on multi-modal Transformer, characterized by: include: The acquisition unit deploys multiple UWB sensors and auxiliary sensors to collect signal characteristics of the target at different locations and construct a multimodal data set; A fusion unit, which uses a multimodal Transformer model to perform feature extraction and fusion on the multimodal data set to generate a global feature representation of the target; The positioning unit predicts the location information of the target according to the global feature representation and combines the dynamic weight allocation mechanism, and performs adaptive correction for abnormal signals.
7. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Cited By
RFID indoor positioning method and system based on hybrid neural network model
CN120317273A
Pentahedron machining center precision calibration method and system based on multi-sensor fusion
CN120480662A
High-precision AI positioning method and system based on spatial multi-modal data fusion
CN120491127A
Intelligent storage plate accurate positioning management system based on Internet of Things
CN120873392A
Civil aviation passenger-oriented cross-mechanism privacy protection portrait fusion method and system
CN121723411A