RFID tag positioning error correction method based on deep learning
By combining multimodal perception networks and deep learning models, the accuracy and adaptability issues of RFID positioning technology in complex environments are solved, achieving high-precision and fast-response positioning effects.
Patent Information
- Application Number
- CN202510768467.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing RFID positioning technology has low positioning accuracy in complex environments. Traditional methods are difficult to adapt to environmental changes. Multipath fading and electromagnetic noise interference lead to large positioning errors, and the generalization ability of existing deep learning models is insufficient.
A multimodal perception network is used to synchronously collect RFID signals and environmental auxiliary data, and a spatiotemporal multidimensional feature dataset is constructed. A hybrid model of the spatiotemporal Transformer network and the generative adversarial network is combined. Adaptation parameters are generated in real time through a model-independent meta-learning algorithm. Edge computing nodes perform preprocessing and feature extraction, and cloud servers perform global error compensation to build a multi-label collaborative positioning network.
It improves positioning accuracy, enhances the model's robustness to multipath signals and electromagnetic noise, enables rapid adaptation to environmental changes, reduces positioning error by 53%, and improves positioning generalization capabilities in complex environments.
Smart Images

Figure CN120671695A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless radio frequency identification technology, and in particular to a method for correcting RFID tag positioning errors based on deep learning. Background Art
[0002] Radio Frequency Identification (RFID) is a technology that uses radio waves to achieve contactless information exchange and target positioning. Its core principle is to utilize radio frequency signal transmission between a tag and a reader to automatically identify and estimate the location of a target object. RFID positioning technology, due to its advantages such as non-line-of-sight sensing, low cost, and easy deployment, has been widely used in scenarios such as smart warehousing, logistics management, and the Industrial Internet of Things. However, in complex real-world environments, radio wave propagation is susceptible to factors such as multipath (where the signal is reflected or refracted by obstacles, forming multiple paths), non-line-of-sight (NLOS) propagation (where the signal is completely blocked by obstacles, causing propagation path distortion), and dynamic environmental interference (such as human movement and electromagnetic noise from equipment), resulting in significant increases in positioning errors. Traditional positioning methods (such as maximum likelihood estimation based on signal propagation models and weighted centroid algorithms) rely on the assumption of a static nonlinear mapping between signal strength and distance, making them difficult to adapt to real-time changes in environmental parameters. Positioning accuracy plummets in multipath fading scenarios. In recent years, although deep learning-based positioning solutions have improved positioning accuracy through data-driven approaches, certain issues remain.
[0003] First, it relies solely on RFID signal features (such as RSSI sequences) without integrating environmental auxiliary data such as the inertial measurement unit (IMU) and visual images. This leads to positioning failure due to feature loss in scenarios with strong occlusion or signal blind spots.
[0004] Second, the fixed model parameters used in offline training cannot respond to environmental changes in real time (such as shelf layout adjustments and reader failures). During long-term operation, errors accumulate with dynamic environmental disturbances.
[0005] Third, traditional feature extraction networks lack robustness to nonlinear interference such as multipath signal superposition and electromagnetic noise, and their model generalization capabilities are limited, making them difficult to adapt to complex electromagnetic environments such as factories and hospitals.
[0006] To this end, a RFID tag positioning error correction method based on deep learning is proposed. Summary of the Invention
[0007] In view of this, the embodiments of the present invention hope to provide an RFID tag positioning error correction method based on deep learning to solve or alleviate the technical problems existing in the prior art and at least provide a beneficial option.
[0008] To solve the above technical problems, the technical solution adopted in this application is: a method for correcting RFID tag positioning errors based on deep learning, comprising the following steps:
[0009] Step 1: Deploy a multimodal sensing network in the target area to synchronously collect the radio wave signals of RFID tags and environmental auxiliary data to construct a spatiotemporal multidimensional feature dataset;
[0010] Step 2: Build a joint architecture hybrid model consisting of a spatiotemporal Transformer network and a generative adversarial network;
[0011] Step 3: Input the multidimensional feature dataset into the spatiotemporal Transformer network to generate a feature encoding vector including spatiotemporal correlation information;
[0012] Step 4: Based on the feature encoding vector, a training framework based on the generative adversarial network is constructed. The hybrid model parameters are optimized through the adversarial game process, so that the hybrid model learns robust feature expression.
[0013] Step 5: Initialize the hybrid model parameters based on the model-independent meta-learning algorithm and automatically generate adaptive parameters according to environmental changes;
[0014] Step 6: Deploy the adaptation parameters to the edge computing node, configure the edge hybrid model parameters, preprocess and extract features of the real-time collected radio wave signals, and generate local positioning results;
[0015] Step 7: The cloud server receives the local positioning results uploaded by the edge node, performs global hybrid model training and error compensation, outputs the global error compensation amount, and generates the corrected coordinates;
[0016] Step 8: Construct the corrected coordinates and the position information of adjacent tags into a graph structure to form a globally consistent multi-label collaborative localization network.
[0017] As a further preferred embodiment of the present technical solution, in step 2, the spatial attention module of the spatiotemporal Transformer network adopts a multi-head self-attention mechanism, and the calculation formula is:
[0018]
[0019] Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively, d k is the key vector dimension.
[0020] As a further preferred embodiment of the present technical solution, in step 5, the specific steps of automatically generating the adaptation parameters are as follows:
[0021] Step 501: monitor the current environmental parameters in real time and calculate the deviation between the current environmental parameters and the historical reference values;
[0022] Step 502: Control the multimodal perception network to collect label sample data in the new scenario based on the deviation result;
[0023] Step 503: Based on the model-independent meta-learning algorithm, the model parameters initialized by the historical tasks are used to perform gradient update on the newly collected labeled sample data to generate temporary parameters;
[0024] Step 504: Evaluate the importance of temporary parameters to historical tasks and impose knowledge protection constraints on key parameters during parameter update.
[0025] Step 505: Integrate the meta-learning update results and the knowledge protection constraints to generate the final adaptation parameters that adapt to the current environment and deploy them to the edge computing node.
[0026] As a further preferred embodiment of the present technical solution, in step 6, the specific steps of generating the local positioning result are as follows:
[0027] Step 601: pre-processing the radio wave signal collected in real time;
[0028] Step 602: Input the pre-processed radio wave signal into the spatiotemporal Transformer network to extract the spatial correlation features of the multi-reader signals;
[0029] Step 603: Configure the edge hybrid model based on the deployed adaptation parameters, process the spatial correlation features through the lightweight network compressed by knowledge distillation, and generate initial positioning coordinates;
[0030] Step 604: perform spatial constraint matching on the initial positioning coordinates of the current tag and the position information of adjacent known tags, and iteratively optimize the initial positioning coordinates to generate a final local positioning result.
[0031] As a further preferred embodiment of the present technical solution, in step seven, the specific steps of generating the corrected coordinates are as follows:
[0032] Step 701: The cloud server receives local positioning results from multiple edge nodes, aligns the data in time and space dimensions, and integrates global environment map information.
[0033] Step 702: Using historical positioning data and real coordinates, train a global error compensation model that combines traditional statistical methods with deep learning to learn the error distribution patterns under different environments.
[0034] Step 703: For each local positioning result, according to its environmental characteristics and historical error patterns, calculate the error correction amount through the global model and fuse the compensation results of multiple methods;
[0035] Step 704: Apply the correction amount to the local positioning result to generate the corrected coordinates, verify the rationality of the result through the consistency of adjacent label positions and global topological constraints, and handle outliers.
[0036] As a further preferred embodiment of the present technical solution, in step four, the generative adversarial network adopts a conditional generative adversarial network architecture, and the generator input of the conditional generative adversarial network architecture includes a feature encoding vector and an environment category label.
[0037] As a further preferred embodiment of the present technical solution, in step 1, the radio wave signal includes signal strength, carrier phase difference, arrival time, and arrival angle; and the environmental auxiliary data includes three-dimensional acceleration, angular velocity data, and environmental image features;
[0038] The spatiotemporal multidimensional feature dataset is constructed by performing spatiotemporal alignment and standardization preprocessing on the collected radio wave signals and environmental auxiliary data.
[0039] As a further preferred embodiment of the present technical solution, in step eight, the method of constructing the graph structure specifically includes:
[0040] Step 801: Use the corrected coordinates of each label as a graph node and the spatial distance between adjacent labels as edge weights to construct an undirected weighted graph;
[0041] Step 802: Perform feature propagation on the graph structure of the undirected weighted graph based on a graph convolutional network or a graph attention network, and optimize global positioning consistency by constraining the positions of adjacent nodes.
[0042] Step 803: Define the global loss function as the sum of the squares of all adjacent node position errors, and iteratively update the node coordinates in combination with the gradient descent algorithm to form a multi-label collaborative localization network.
[0043] As a further preferred embodiment of the present technical solution, the environmental parameters include at least one of signal strength distribution, multipath fading coefficient, and non-line-of-sight probability.
[0044] As a further preferred embodiment of the present technical solution, the preprocessing includes time domain processing, frequency domain processing, feature normalization, and outlier detection and repair.
[0045] The embodiment of the present invention adopts the above technical solution, which has the following advantages:
[0046] 1. This invention uses a multimodal sensing network to synchronously collect RFID signals and environmental auxiliary data to construct a spatiotemporal multidimensional feature dataset. This addresses the problem of missing features of a single RFID signal in strong occlusion or signal blind areas, enabling the positioning system to maintain effective output in complex environments through multi-source data fusion.
[0047] 2. This invention introduces a model-independent meta-learning algorithm to monitor environmental parameter deviations in real time and dynamically generate adaptive parameters. Combined with knowledge protection constraints, this algorithm prevents the forgetting of historical task knowledge. The parameter update process can be completed within 30 seconds, enabling the model to quickly adapt to dynamic environmental changes such as shelf adjustments and electromagnetic interference, and suppressing the accumulation of errors due to environmental disturbances.
[0048] 3. This invention adopts a joint architecture of spatiotemporal Transformer and generative adversarial network, captures the spatial correlation of multipath signals through a multi-head self-attention mechanism, optimizes feature expression using adversarial game, and enhances the model's robustness to nonlinear interference such as multipath fading and electromagnetic noise. Experiments show that the positioning error of this architecture in metal-dense areas is reduced by 53% compared with traditional methods, and the generalization ability is significantly improved.
[0049] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 This is a flow chart of a method for correcting RFID tag positioning errors based on deep learning according to the present invention;
[0052] Figure 2 A schematic diagram of a process for automatically generating adaptive parameters according to the present invention;
[0053] Figure 3 A schematic diagram of a flow chart of a method for generating local positioning results according to the present invention;
[0054] Figure 4 A schematic diagram of a flow chart of a method for generating corrected coordinates according to the present invention;
[0055] Figure 5 The figure is a flow chart of the method for constructing a graph structure according to the present invention. DETAILED DESCRIPTION
[0056] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0057] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0058] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0059] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0060] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0061] Example 1
[0062] Figure 1 This is a flow chart of a method for correcting RFID tag positioning errors based on deep learning according to an embodiment of the present invention. It should be noted that if there are substantially the same results, the method of this application is not based on Figure 1The process sequence shown is limited. Figure 1-Figure 5 As shown: A method for correcting RFID tag positioning errors based on deep learning, comprising the following steps:
[0063] Step 1: Deploy a multimodal sensing network in the target area to synchronously collect the radio wave signals of RFID tags and environmental auxiliary data to construct a spatiotemporal multidimensional feature dataset;
[0064] Specifically, first, multimodal sensing devices such as RFID readers, IMU modules, and industrial cameras are deployed at preset intervals in the target area, and microsecond-level time synchronization between devices is achieved through the NTP protocol;
[0065] Then, the reader is used to collect radio wave signals such as signal strength and carrier phase difference in real time. At the same time, the IMU module and industrial camera are used to obtain inertial data such as three-dimensional acceleration and angular velocity, as well as environmental image features.
[0066] Then, the collected multi-source data is subjected to spatiotemporal alignment to ensure that the data of each modality accurately matches in terms of timestamps and spatial coordinates;
[0067] Finally, the processed signals and data are subjected to standardized preprocessing, including time domain filtering, frequency domain transformation, feature normalization, and outlier repair operations, and finally a multidimensional feature data set containing spatiotemporal correlation information is constructed.
[0068] Step 2: Build a joint architecture hybrid model consisting of a spatiotemporal Transformer network and a generative adversarial network;
[0069] Specifically, we first designed a spatiotemporal Transformer network structure, inputting a multi-dimensional feature dataset into the temporal embedding layer and the position encoding layer. A multi-head self-attention mechanism was used to capture the dependencies between different time steps and spatial positions. Meanwhile, a gated recurrent unit was introduced to enhance the temporal modeling capability.
[0070] Then, a generative adversarial network architecture is constructed, in which the generator uses residual blocks and upsampling layers to generate simulated signal features from the features extracted by the spatiotemporal Transformer, while the discriminator uses a deep convolutional structure to distinguish between real and generated signals.
[0071] Next, we design a joint training mechanism that feeds the output features of the spatiotemporal Transformer and the output of the generator into the discriminator. We then implement end-to-end training of the hybrid model by minimizing the adversarial loss function and the reconstruction loss function.
[0072] Finally, an attention fusion module is introduced to adaptively adjust the weight ratio of spatiotemporal Transformer features and generator features, thereby improving the model's ability to represent signal features in complex environments.
[0073] Step 3: Input the multidimensional feature dataset into the spatiotemporal Transformer network to generate a feature encoding vector including spatiotemporal correlation information;
[0074] Specifically, first, the preprocessed multidimensional feature dataset is divided into time windows to form an input sequence containing historical time series and spatial position information. The input of each time step contains RFID signal features, IMU inertial data and image feature vectors;
[0075] Then, the time index is mapped into a low-dimensional vector through the time embedding layer and added element-by-element to the spatial position encoding vector to form a position representation that contains both temporal and spatial information.
[0076] Next, the spatiotemporal encoded input sequence is fed into the encoder part of the spatiotemporal Transformer network. The multi-head self-attention mechanism is used to calculate the association weights between different time steps and spatial positions to generate a spatiotemporal dependency matrix.
[0077] Subsequently, a feedforward neural network is used to perform nonlinear transformation on the attention output, and residual connections and layer normalization operations are used to enhance the feature transfer capability.
[0078] Finally, a feature encoding vector containing global spatiotemporal correlation information is extracted from the output of the encoder. This vector will serve as the input of the subsequent generative adversarial network to simulate the propagation characteristics of RFID signals in complex environments.
[0079] Step 4: Based on the feature encoding vector, a training framework based on the generative adversarial network is constructed. The hybrid model parameters are optimized through the adversarial game process, so that the hybrid model learns robust feature expression.
[0080] Specifically, we first designed a generative adversarial network architecture. The feature encoding vector output by the spatiotemporal transformer is input into the generator. The generator uses a deconvolutional network structure to gradually map low-dimensional features into high-dimensional simulated RFID signal features. The discriminator uses a convolutional network structure to receive the real RFID signal features and the simulated signal features output by the generator, and outputs the discrimination result through multi-layer nonlinear transformation.
[0081] Next, we build an adversarial training framework, defining the generator’s objective function as maximizing the probability of the discriminator’s misclassification, and the discriminator’s objective function as minimizing the classification error. We implement adversarial game by alternately optimizing the parameters of the generator and discriminator.
[0082] Then, a gradient penalty mechanism is introduced to ensure that the discriminator satisfies the Lipschitz constraint and enhance training stability. At the same time, a feature matching loss function is designed to measure the distance between the generator output and the true signal in the feature space, which helps the generator learn richer signal features.
[0083] Finally, the historical buffer technology is introduced during the training process to store the historical output of the generator, increase the training difficulty of the discriminator, and continuously optimize the parameters of the hybrid model through multiple rounds of adversarial games, so that it can learn robust feature expressions and effectively deal with environmental noise and signal interference.
[0084] Step 5: Initialize the hybrid model parameters based on the model-independent meta-learning algorithm and automatically generate adaptive parameters according to environmental changes;
[0085] Specifically, the model-agnostic meta-learning algorithm (MAML) is first used to initialize the parameters of the hybrid model. By performing rapid gradient updates in multiple historical environment tasks, the initial parameters with generalization capabilities are obtained.
[0086] Next, the system monitors the current environmental parameters (such as signal strength distribution and multipath fading coefficient) in real time and calculates their deviation from the historical baseline values. When the deviation exceeds the preset threshold, the parameter adaptation process is triggered.
[0087] Then, the multimodal perception network is controlled to collect labeled sample data in the new environment. Based on the MAML algorithm, the initialization parameters are used to perform gradient updates on the new data to generate temporary parameters. During this period, the importance of the temporary parameters to the historical task is evaluated through the Fisher information matrix, and L2 regularization or mask constraints are applied to key parameters to prevent important knowledge from being forgotten.
[0088] Finally, the meta-learning update results and knowledge protection constraints are integrated to generate the final parameters that adapt to the current environment. The parameters are deployed through edge computing nodes to enable the model to respond quickly to dynamic environments (the entire process is usually completed within 30 seconds).
[0089] Step 6: Deploy the adaptation parameters to the edge computing node, configure the edge hybrid model parameters, preprocess and extract features of the real-time collected radio wave signals, and generate local positioning results;
[0090] Specifically, first, the generated adaptation parameters are transmitted to the edge computing node (such as NVIDIA Jetson Nano) through the network to complete the parameter configuration and initialization of the edge hybrid model;
[0091] Next, the real-time collected radio wave signals (including signal strength, carrier phase difference, etc.) are sequentially processed in the time domain (using sliding window filtering to remove burst noise and combining Kalman filtering to smooth signal fluctuations), frequency domain processing (using fast Fourier transform to analyze spectral characteristics and wavelet transform to extract specific frequency band features), feature normalization (mapping signal features to the [0,1] interval), and outlier detection and repair (identifying and repairing abnormal signal values based on historical data models).
[0092] Subsequently, the preprocessed signal is fed into a spatiotemporal Transformer network, where its spatial attention module is used to extract the spatial correlation features of the multi-reader signals. Based on the deployed adaptation parameters, a lightweight network (such as MobileNetV2) compressed by knowledge distillation is used to reduce the dimensionality and classify the features to generate initial positioning coordinates.
[0093] Finally, the initial coordinates of the current tag are matched with the location information of adjacent known tags under spatial constraints (such as setting a distance threshold of 2 meters). The Levenberg-Marquardt algorithm or gradient descent method is used to iteratively optimize the initial coordinates and output the final local positioning result (single-tag inference delay is usually <50ms).
[0094] Step 7: The cloud server receives the local positioning results uploaded by the edge node, performs global hybrid model training and error compensation, outputs the global error compensation amount, and generates the corrected coordinates;
[0095] Specifically, first, the cloud server receives the local positioning results uploaded by each edge node through a high-speed network, aligns and calibrates the multi-source data according to the timestamp and spatial coordinate dimensions, and integrates the global environment map information (such as building floor plans, obstacle distribution, etc.) to build a unified spatiotemporal data view;
[0096] Next, using historically accumulated positioning data and the corresponding real-world coordinates, a global error compensation model is trained that combines traditional statistical methods (such as Kalman filtering) with deep learning models (such as convolutional neural networks). By analyzing the error distribution patterns under different environmental characteristics (such as metal obstruction level and multipath fading intensity), an error mapping relationship is established.
[0097] Then, for the local positioning results uploaded by each edge node, the corresponding error correction amount is calculated through a global error compensation model based on the environmental characteristics (such as signal strength distribution and non-line-of-sight probability) and historical error patterns. The compensation results of multiple methods such as weighted averaging and Gaussian process regression are integrated to improve the correction accuracy.
[0098] Finally, the calculated error correction amount is applied to the original local positioning result to generate preliminary corrected coordinates. The rationality of the result is then verified through adjacent tag position consistency check (setting a spatial distance threshold, such as 30 cm) and global topological structure constraints (such as shelf row and column alignment requirements). Median filtering or isolation forest algorithm is used to eliminate outliers and output the final corrected coordinates.
[0099] Step 8: Construct the corrected coordinates and the position information of adjacent tags into a graph structure to form a globally consistent multi-tag collaborative localization network;
[0100] Specifically, first, the corrected coordinates of each label are used as nodes in the graph structure, the Euclidean distance between adjacent labels is calculated as the edge weight, and an undirected weighted graph is constructed (a threshold for determining adjacent nodes is set, such as within 2 meters is considered adjacent);
[0101] Next, we propagate features of the graph structure based on a graph convolutional network (GCN) or a graph attention network (GAT). By aggregating the position information of adjacent nodes and edge weight constraints, we optimize the spatial consistency between nodes (for example, we use the Laplace smoothing property of GCN to propagate position features).
[0102] Then, define the global loss function as the sum of the squares of the position errors of all adjacent nodes, and combine the gradient descent algorithm (such as Adam optimizer) to iteratively update the node coordinates to minimize the global positioning error;
[0103] Finally, the optimization process is terminated by setting the number of iterations or the loss convergence threshold to form a global collaborative localization network that includes spatial dependencies between labels. This network can continuously adapt to label position changes in a dynamic environment by updating graph nodes and edge weights in real time.
[0104] In one embodiment, specifically, in step 2, the spatial attention module of the spatiotemporal Transformer network adopts a multi-head self-attention mechanism, and the calculation formula is:
[0105]
[0106] Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively, d k is the key vector dimension;
[0107] Specifically, the self-attention mechanism is a core component of the Transformer architecture. It allows the model to assign attention weights based on the degree of correlation between elements at different positions when processing sequence data, thereby capturing long-range dependencies in the sequence. In the spatial attention module of the spatiotemporal Transformer network, the self-attention mechanism is used to mine the correlation between signal features at different spatial positions.
[0108] The specific calculation process is as follows:
[0109] First, define the matrix:
[0110] The query matrix is Q: It is used to "query" the elements in the sequence and find relevant information. It determines the focus of the model;
[0111] The key matrix is K: it can be regarded as an "index" for matching queries, providing comparable information for queries;
[0112] The value matrix is V: It contains the information actually used to calculate the output. According to the result of matching the query with the key, the corresponding information is extracted from the value matrix;
[0113] Then, the similarity calculation is performed as follows:
[0114] Calculate the transpose K of the query matrix Q and the key matrix K T The product QK T , we get an attention score matrix, which measures the degree of association between each position in the sequence; in order to prevent the score from being too large and causing the softmax function gradient to disappear, we divide it by d k is the dimension of the key vector, which plays a normalization role;
[0115] Next, the weight calculation is performed as follows:
[0116] The normalized attention score matrix is processed by the softmax function to convert the score into attention weights in the form of probability distribution. That is, different attention weights are assigned to each position, indicating the degree of attention paid to the position in the current calculation.
[0117] Then, a weighted summation is performed as follows:
[0118] The obtained attention weights are used to perform weighted summation on the value matrix V to obtain the final attention output, which comprehensively considers the information of each position in the sequence and focuses on the position information with high relevance to the query;
[0119] The multi-head self-attention mechanism executes the above self-attention mechanism in parallel h times (h is the number of heads), using different linear projections each time to obtain different Q, K, and V matrices, thereby capturing feature information from different subspaces. Finally, the outputs of h heads are spliced and linearly transformed to obtain the final output; this allows the model to capture the relationship between sequence features from multiple angles and enhance the model's expressive power.
[0120] In one embodiment, specifically, in step five, the specific steps of automatically generating the adaptation parameters are as follows:
[0121] Step 501: monitor the current environmental parameters in real time and calculate the deviation between the current environmental parameters and the historical reference values;
[0122] Specifically, first, determine the environmental parameters that need to be monitored, such as RFID signal strength, ambient temperature, humidity, electromagnetic interference intensity, etc., and set a reasonable collection frequency based on scenario requirements (such as collecting data every 5 seconds). At the same time, use the NTP protocol to ensure the time synchronization of each sensor data;
[0123] Then, in the initial stage of the system, environmental data is continuously collected for a certain period of time (such as one week). The statistical characteristics such as the mean and standard deviation of each parameter are calculated according to the time window to construct a historical baseline value. During this period, an algorithm is used to eliminate abnormal data, and the historical data is regularly updated through a sliding window mechanism.
[0124] Next, the environmental parameters collected in real time are preprocessed by filtering and normalization, and the point deviation, trend deviation, distribution deviation, etc. are calculated respectively. The comprehensive deviation score is calculated based on the weights determined in advance through cross-validation.
[0125] Finally, the deviation is judged based on the set three-level warning threshold. If the red warning threshold is reached, the parameter adaptation process is triggered. If it is a yellow warning, the adaptive preparation is started. The green warning indicates that the environment is stable.
[0126] Step 502: Control the multimodal perception network to collect label sample data in the new scenario based on the deviation result;
[0127] Specifically, first, when the environmental deviation score reaches the yellow warning threshold, the edge computing node automatically activates the enhanced acquisition mode of the multimodal perception network, increasing the RFID reader sampling rate to 200Hz, adjusting the IMU sensor frequency to 100Hz, and increasing the industrial camera frame rate to 30fps;
[0128] Then, the sensor configuration parameters are dynamically adjusted according to the deviation type. For example, when increased interference from metal objects is detected, the UHF RFID reader transmit power is increased to 30dBm, and the multi-band joint acquisition mode (860-960MHz segmented scanning) is enabled at the same time.
[0129] Next, an intelligent sampling strategy was initiated, using a spatio-temporal interest points (SPIP) algorithm to identify key sampling areas. Intensive sampling (increasing sampling density by three times) was implemented in areas with drastic deviation changes, while regular sampling was maintained in other areas.
[0130] Subsequently, the microsecond-level timestamp alignment of multi-sensor data is achieved through a data synchronization bus, and the raw data is pre-processed using a Kalman fusion filter to eliminate time delay errors between sensors.
[0131] Finally, the collected raw data is packaged into 5-second time windows and transmitted to the edge server through the MQTT protocol. At the same time, the environmental deviation feature vector (including 12-dimensional features such as temperature change rate and RSSI fluctuation variance) is attached to form a labeled sample dataset with environmental context awareness capabilities.
[0132] Step 503: Based on the model-independent meta-learning algorithm, the model parameters initialized by the historical tasks are used to perform gradient update on the newly collected labeled sample data to generate temporary parameters;
[0133] Specifically, first, the model parameters θ0, which are pre-trained based on historical environmental tasks (such as warehouse shelf layout changes and logistics vehicle entry and exit scenarios), are loaded from the parameter server. This parameter is obtained through rapid adaptive training on multiple environmental tasks using the MAML algorithm and has good generalization capabilities.
[0134] Then, the newly collected labeled sample data is divided into a support set and a query set in a ratio of 1:4. The support set is used to calculate the gradient update direction, and the query set is used to evaluate the performance of the updated model.
[0135] Next, for each task (such as the positioning task in a specific metal interference environment), the support set data is used to calculate the loss function Gradient with respect to parameter θ0 And generate temporary parameters through one-step gradient update Among them, the learning rate α adopts an adaptive adjustment strategy (such as dynamic scaling based on the gradient norm);
[0136] Then, the query set data is used to calculate the temporary parameters θ′ i The loss of yuan And calculate the meta-gradient through the second-order derivative Used to evaluate the effectiveness of parameter updates;
[0137] Finally, the meta-gradients of all tasks are aggregated (e.g., weighted average) to generate the final temporary parameter θ temp This parameter not only retains the knowledge of historical tasks but also adapts to the feature distribution of the new environment. The entire process uses model quantization technology (such as 8-bit integer inference) to reduce the computational complexity by 75%, ensuring that the processing delay on the edge device is <100ms.
[0138] Step 504: Evaluate the importance of temporary parameters to historical tasks and impose knowledge protection constraints on key parameters during parameter update.
[0139] Specifically, we first use a parameter importance evaluation method based on the Fisher information matrix to calculate the Fisher information of each parameter on the historical task, quantify its contribution to historical knowledge, and identify key parameters that have a significant impact on the performance of historical tasks (such as the parameters of the self-attention module of the spatiotemporal Transformer and the parameters of the discriminator convolution kernel of the generative adversarial network).
[0140] Then, construct the parameter importance mask matrix (D is the total number of parameters);
[0141] Among them, the mask value corresponding to the key parameter is set to m key ∈(0.1, 0.5) (indicates strong constraint);
[0142] Non-critical parameters are set to m non-key =1 (indicates no constraint);
[0143] Then, in the parameter update process, the temporary parameter θ temp Imposing intellectual property protection constraints:
[0144] θ protected =θ0+M☉(θ temp –θ0);
[0145] Among them, ⊙ represents the element-by-element product, through which the update amplitude of the key parameters is forced to not exceed a certain proportion of the historical parameters (such as limiting the update to ≤20% of the original parameter norm), avoiding excessive forgetting of historical task knowledge due to environmental changes;
[0146] Subsequently, a contrastive learning mechanism is introduced to calculate the cosine similarity of the parameters before and after the update in the historical task feature space. If the similarity is lower than a threshold (such as 0.7), the parameter rollback mechanism is triggered to restore to the most recent parameter version that meets the constraint conditions.
[0147] Finally, the hyperparameters of the mask matrix M (such as m key ), maximizing the parameter adaptation speed in new environments while ensuring that the performance loss of historical tasks is ≤5%. For example, in smart warehousing scenarios, when shelf layout adjustments lead to an increase in multipath effects, this mechanism can reduce the parameter update amplitude of the Transformer layer responsible for spatial feature extraction by 40%, while simultaneously restoring the positioning accuracy in the new environment to 92% of the historical level within 30 seconds.
[0148] Step 505: Integrate the meta-learning update results and the knowledge protection constraints to generate final adaptation parameters that adapt to the current environment and deploy them to the edge computing node;
[0149] Specifically, first, construct the weighted fusion function:
[0150] F(θ0,θ temp ,M)=(1-λ)·θ0+in·(θ0+M☉(θ temp –θ0));
[0151] Where λ∈[0,1] is the adaptive fusion coefficient, which is dynamically determined by the environment complexity evaluator (e.g., adjusted according to the normalized value of the deviation score);
[0152] Then, the validation set data (including historical task samples and new environment samples) is used to calculate the model performance indicators (such as positioning error, feature reconstruction loss, etc.) under different λ values, and the λ that makes the overall performance optimal is selected. * ; Then, based on λ * Generate the final adaptation parameters θ final =F(θ0,θ temp , M), and perform quantization compression (such as INT8 quantization) on it, reducing the parameter size by 75%;
[0153] Subsequently, the integrity of the parameters is verified through an edge-cloud collaborative signature mechanism to ensure that the transmission process has not been tampered with;
[0154] Finally, the incremental update strategy is used to deploy the parameters to the edge computing nodes, as follows:
[0155] Push parameter update instructions to edge nodes via the MQTT protocol, including parameter hash values and version numbers;
[0156] After receiving the data, the edge node compares the new parameters with the historical parameters cached locally and only downloads the changed parts (the average transmission volume is reduced by 90%).
[0157] Perform parameter hot update and load new parameters into the model memory space without interrupting the real-time positioning service;
[0158] Start the parameter verification process to verify the inference accuracy of the new parameters using local test data. If the error exceeds a threshold (such as 15%), a rollback mechanism is triggered.
[0159] The entire deployment process takes less than 2 seconds in a 5G network environment and less than 5 seconds in a Wi-Fi environment. After deployment, the edge node automatically restarts the positioning service and uses the new parameters to process real-time signals to achieve rapid response to dynamic environments. For example, in a factory production line adjustment scenario, the system's positioning accuracy for RFID tags is improved from 1.2 meters to 0.35 meters within 10 seconds after deployment, meeting industrial-grade positioning requirements.
[0160] In one embodiment, specifically, in step six, the specific steps of generating the local positioning result are as follows:
[0161] Step 601: pre-processing the radio wave signal collected in real time;
[0162] Specifically, the real-time radio wave signals (including signal strength RSSI, carrier phase difference, time of arrival (ToA), angle of arrival (AoA), etc.) are first processed in the time domain. Sliding window filtering (with a window size of 100ms) is used to remove burst noise. Then, a Kalman filter or particle filter algorithm is used to smooth signal fluctuations. For example, a state-space model is established for the ToA signal, and a prediction-update step is used to suppress random errors caused by multipath effects.
[0163] Next, frequency domain processing is performed: the time-domain filtered signal is converted to the frequency domain using a fast Fourier transform (FFT). The signal spectrum characteristics are analyzed to identify interference frequency bands (such as 50 Hz power frequency noise). Specific frequency band features (such as high-frequency multipath components) are then extracted using a wavelet transform (such as the db4 wavelet basis) and the noise frequency band is filtered out.
[0164] Then, feature normalization is performed: the feature values of different dimensions, such as RSSI and phase, are mapped to the [0, 1] interval through the Min-Max or Z-Score normalization method to eliminate the signal strength differences between different readers or tags. For example, the RSSI range of [-100dBm, -20dBm] is linearly scaled to [0, 1].
[0165] Finally, outlier detection and repair are performed: a signal distribution model (such as a Gaussian mixture model (GMM)) is constructed based on historical data. Anomalous signal values are identified using thresholding or the isolation forest algorithm. Detected outliers are repaired using interpolation methods (such as linear interpolation or K-nearest neighbor interpolation) to ensure the continuity and reliability of the signal features input to the model. The entire preprocessing process has a processing latency of less than 20ms at the edge computing node, effectively improving the accuracy of subsequent feature extraction and positioning.
[0166] Step 602: Input the pre-processed radio wave signal into the spatiotemporal Transformer network to extract the spatial correlation features of the multi-reader signals;
[0167] Specifically, first, the preprocessed signal is split into sequence inputs according to time windows (e.g., 50ms). Each time step contains the signal feature vectors (e.g., RSSI, phase, etc.) of multiple readers, and position encoding is added to preserve spatial information.
[0168] Then, the multi-head self-attention mechanism is used to calculate the association weights between different reader signals to capture the spatial correlation of the signals. For example, the attention weight of the i-th reader to the j-th reader is calculated as:
[0169]
[0170] Among them, Q, K, and V are query, key, and value matrices respectively, and d k is the key vector dimension;
[0171] Next, the gated recurrent unit (GRU) is introduced to enhance the temporal modeling capability and process the temporal variation characteristics of the signal;
[0172] Subsequently, feature transfer is optimized through residual connections and layer normalization to alleviate the gradient vanishing problem of deep networks;
[0173] Finally, a feature encoding vector containing spatial correlation is extracted from the network output. This vector integrates the dependencies between multiple reader signals and is used for subsequent positioning calculations. For example, in smart warehousing scenarios, the network can effectively identify the impact of occlusion relationships between shelves on signal propagation, improving the tag positioning accuracy to within 0.3 meters.
[0174] Step 603: Configure the edge hybrid model based on the deployed adaptation parameters, process the spatial correlation features through the lightweight network compressed by knowledge distillation, and generate initial positioning coordinates;
[0175] Specifically, first, the adaptation parameters generated in step 505 are loaded into the edge computing node, and a hybrid model of the spatiotemporal Transformer and the generative adversarial network is configured. To address the resource limitations of edge devices (such as the 4GB memory of the NVIDIA Jetson Nano), knowledge distillation technology is used to compress the original model into a lightweight network:
[0176] Teacher-student architecture: Using the full hybrid model as the teacher network, a student network with 80% fewer parameters is constructed (e.g., MobileViT-S)
[0177] Distillation loss function:
[0178] in, is the classification cross entropy loss, For the teacher prediction distribution p T With the Student predictive distribution p S KL divergence, α = 0.3 is the balance coefficient;
[0179] Quantization optimization: Quantize model parameters from FP32 to INT8 to further reduce computational complexity;
[0180] Next, the spatial correlation features extracted in step 602 are input into the compressed lightweight network, and the initial positioning coordinates are generated through the following processing:
[0181] Feature mapping layer: Projects high-dimensional feature vectors into 3D coordinate space:
[0182] Coord x ,Coord y Coord z=MLP(f spatial );
[0183] Uncertainty estimation: Output the confidence score of positioning at the same time:
[0184] Confidence = σ(MLP conf (f spatial ));
[0185] Coordinate refinement: Apply spatial constraints (such as shelf row and column spacing priors) to modify the initial coordinates;
[0186] The latency of the entire inference process on the edge device is less than 15ms, and the positioning error is less than 0.2m in open environments and less than 0.5m in complex environments, meeting industrial-grade real-time positioning requirements. For example, in smart warehousing scenarios, this lightweight network can achieve centimeter-level positioning of tags in areas with dense shelves while maintaining a model update frequency of >30Hz.
[0187] Step 604: perform spatial constraint matching on the initial positioning coordinates of the current tag and the position information of adjacent known tags, and iteratively optimize the initial positioning coordinates to generate the final local positioning result:
[0188] Specifically, first, based on the initial positioning coordinates of the current tag, a set radius (e.g., 2 meters) is used to search for tags with adjacent known locations to construct a local spatial constraint set, where the location information of adjacent tags comes from historical calibration results or manual calibration.
[0189] Next, the Euclidean distance between the current tag and each adjacent tag is calculated and compared with the theoretical distance in the actual physical space (based on the map or prior layout) to generate a distance error vector Δd = [d 实测 -d 理论 ];
[0190] Subsequently, a spatial constraint optimization model is constructed with the objective function of minimizing the sum of squared distance errors. Combined with the physical reachability constraints between tags (such as whether there are obstacles blocking them), the Levenberg-Marquardt algorithm or particle swarm optimization algorithm is used for iterative optimization. The calculation formula is:
[0191]
[0192] Among them, X is the coordinate of the current label to be optimized, X i are the coordinates of adjacent labels, c(x) is the obstacle blocking constraint (the penalty coefficient λ increases 10 times when there is an obstacle);
[0193] Finally, when the iterative error converges (e.g., the error change rate is <1% after 5 consecutive iterations) or the maximum number of iterations is reached (e.g., 50 times), the optimization is terminated and the final local positioning result is output. For example, in the smart warehouse shelf scenario, this method can further reduce the positioning error of densely arranged tags from the initial 0.5 meters to 0.15 meters, effectively improving the positioning consistency of local areas.
[0194] In one embodiment, specifically, in step seven, the specific steps of generating the corrected coordinates are as follows:
[0195] Step 701: The cloud server receives local positioning results from multiple edge nodes, aligns the data in time and space dimensions, and integrates global environment map information.
[0196] Specifically, first, the cloud server receives the local positioning results uploaded by each edge node through a message queue (such as Kafka). Each message contains the tag ID, three-dimensional coordinates, timestamp, confidence score, and environmental context features (such as RSSI distribution histogram);
[0197] Next, a two-way time synchronization algorithm (such as NTP combined with Kalman filtering) is used to eliminate the clock offset of edge nodes and unify all data to the UTC time reference, with the time alignment error controlled within ±50μs.
[0198] Subsequently, a spatial index (such as a quadtree or octree) is constructed to spatially partition the positioning results, mapping the geographic coordinates to the grid cells of the global environment map (the grid size is set according to the scene accuracy requirements, such as 0.5m×0.5m). A map matching algorithm (such as HMM-Viterbi) is then used to constrain the alignment of the label trajectory with the feasible areas in the environment map (such as shelf aisles and operating areas) to correct for position jumps caused by signal multipath.
[0199] Finally, the system fuses global environmental semantic information (such as the location of metal shelves and obstacle distribution) and adds spatial semantic labels to each local positioning point (such as "Shelf A - Level 3"), forming a temporally and spatially consistent and semantically rich global positioning dataset. For example, in a large logistics warehouse scenario, this step can align and fuse positioning data from 20 edge nodes within 300ms, supporting real-time global tracking of tens of thousands of tags.
[0200] Step 702: Using historical positioning data and real coordinates, train a global error compensation model that combines traditional statistical methods with deep learning to learn the error distribution patterns under different environments.
[0201] First, extract at least 1,000 hours of positioning data (covering different time periods and environmental conditions) from the historical database and divide it into a test set and a training set at a ratio of 1:9. Each data set contains:
[0202] Original positioning coordinates (x raw ,y raw ,z raw );
[0203] The corresponding real coordinates (x gt ,y gt , z gt )
[0204] Environmental characteristic vector e = [RSSI mean ,phase std ,temp;humidity,metal_density,...];
[0205] Next, build the hybrid model architecture:
[0206] Traditional statistical layer: Kalman filtering is used to perform preliminary smoothing on the original coordinates and estimate the state transfer moment A and the observation matrix H;
[0207] Deep Learning Layers:
[0208] Input layer: connects the original coordinates and the environment feature vector (3+n in total env dimension);
[0209] Hidden layer: 3 layers of residual blocks, each layer contains 256 neurons, and the activation function uses LeakyReLU (α=0.1);
[0210] Output layer: 3D error compensation vector (Δx, Δy, Δz);
[0211] Fusion layer: weighted fusion of traditional estimation results and deep learning output;
[0212] x compensated =x kalman +w·XNN; weight w is adaptively learned through training data;
[0213] Then, design the loss function:
[0214]
[0215] Among them, p pred and p gt are the distributions of prediction error and true error, respectively, obtained by kernel density estimation;
[0216] Finally, the Adam optimizer (learning rate 10 -4 The model was trained with a batch size of 128 using an early stopping strategy (stopping if the validation set loss does not decrease for 10 consecutive epochs) and a learning rate decay strategy (decreasing by a factor of 0.7 every 5 epochs). After training, the model achieved the following performance on the test set:
[0217] Average positioning error: reduced from the original 0.85 meters to 0.21 meters;
[0218] 95% percentile error: reduced from 2.1 meters to 0.45 meters;
[0219] Environmental adaptability: error increase is less than 30% in metal-intensive areas;
[0220] The model has been deployed to a cloud server and can process the positioning data uploaded by edge nodes in real time to compensate for system errors caused by environmental changes.
[0221] Step 703: For each local positioning result, according to its environmental characteristics and historical error patterns, calculate the error correction amount through the global model and fuse the compensation results of multiple methods;
[0222] Specifically, for each local positioning result uploaded by the edge node, the corresponding real-time environmental information is extracted from the environmental feature database, such as signal strength distribution, the number of metal objects in the current area, temperature and humidity, and other data. At the same time, the error patterns in similar environments (such as past positioning deviations in the same metal-dense area) are searched in the historical positioning records.
[0223] Next, the error correction is calculated using a trained global error compensation model. This model combines two capabilities: first, deep learning automatically analyzes the complex relationship between environmental characteristics and positioning error (for example, identifying the specific impact of metal occlusion on the signal); second, it uses traditional statistical methods (such as linear regression) to quickly estimate the error range in common environments. For example, in areas with large signal fluctuations, the deep learning model will output a more accurate correction value based on patterns learned from historical data. In areas with stable environments, statistical methods directly provide empirical error compensation.
[0224] Then, multiple compensation methods are introduced for cross-validation and fusion. In addition to the deep learning and statistical methods mentioned above, Kalman filtering is also used to smooth positioning errors in time series to avoid drastic jumps in correction results at adjacent moments. For example, when the positioning results of a tag show abnormalities multiple times in a row, the Kalman filter will combine the data trends of the previous and next moments to provide a more reasonable error correction direction.
[0225] Finally, the weight of each method is automatically adjusted according to the complexity of the current environment. For example, in simple and open environments, statistical methods are dominant (because the error pattern is stable); in complex environments with dense metal and frequent personnel flow, the weight of the deep learning model will be significantly improved (because it can capture more subtle environmental changes). The error correction amount after fusion will undergo a confidence check. If the credibility of the corrected result is found to be insufficient (such as inconsistency with the position of adjacent tags), the system will automatically trigger data re-collection or use historical reliable values for correction to ensure that the final output positioning coordinates are accurate and reliable.
[0226] Step 704: Apply the correction amount to the local positioning result to generate the corrected coordinates, verify the rationality of the result through the consistency of adjacent label positions and global topological constraints, and handle outliers;
[0227] Specifically, first, the calculated error correction amount is directly superimposed on the original local positioning coordinates uploaded by the edge node to obtain the preliminary corrected coordinates; for example, if the original coordinates are (5.2 meters, 3.8 meters) and the error correction amount is (-0.15 meters, +0.08 meters), the corrected coordinates become (5.05 meters, 3.88 meters);
[0228] Next, the consistency of adjacent tag positions is verified: with the corrected coordinates as the center, other known tag positions within a preset distance (such as 3 meters) are searched, and the difference between the actual distance between the two and the theoretical distance on the map is calculated; for example, if the theoretical distance between two tags on the map is 2 meters, and the measured distance after correction is 3.5 meters, then it is determined that the positioning result of at least one of the tags is abnormal. At this time, the system will prioritize retaining coordinates that are consistent with the majority of adjacent tag positions and eliminate outliers that are significantly deviated from the group;
[0229] Then, global topological constraints are introduced for secondary verification: combining the fixed structures in the environment map (such as shelf rows and wall positions), the system checks whether the corrected coordinates conform to the preset spatial logic. For example, the coordinates of a label in the warehouse should be located within the shelf aisle, not inside the wall or shelf entity. If the coordinates are found to fall into a prohibited area (such as a wall position), the system automatically projects them to the nearest feasible area (such as the aisle centerline) and recalculates the positional relationship with adjacent labels.
[0230] When processing outliers, the median filter algorithm is first used to smooth the positioning results of the continuous time series and eliminate sudden jump points. For periodic outliers, the isolation forest algorithm is used to identify outliers, and interpolation methods (such as linear interpolation of reliable coordinates of the previous and next moments) are used to generate replacement values. For example, if a tag's positioning results deviate from the adjacent tags by more than 2 meters three times in 10 seconds, the system will identify it as an anomaly and estimate a reasonable value based on the tag's stable coordinates in the past 5 seconds and the next 5 seconds.
[0231] Finally, coordinates that have undergone consistency verification and exception handling are marked as "reliable positioning results" and stored in a global database, simultaneously updating the tag's real-time status (e.g., "stationary" or "moving"). The entire verification process is processed on the cloud server with a latency of less than 100 milliseconds, ensuring real-time and reliable positioning results for large-scale tags (e.g., tens of thousands). For example, in smart warehousing scenarios, the positioning error of tags in the shelf area can be controlled to within 0.2 meters, meeting the accuracy requirements of automated sorting systems.
[0232] In one embodiment, specifically, in step four, the generative adversarial network adopts a conditional generative adversarial network architecture, and the generator input of the conditional generative adversarial network architecture includes a feature encoding vector and an environment category label.
[0233] In one embodiment, specifically, in step 1, the radio wave signal includes signal strength, carrier phase difference, arrival time, and arrival angle; the environmental auxiliary data includes three-dimensional acceleration, angular velocity data, and environmental image features;
[0234] Signal strength (RSSI): The power strength of the received signal, reflecting the distance between the tag and the reader and the signal attenuation;
[0235] Carrier phase difference: The phase difference between multiple signals is used for high-precision ranging and positioning;
[0236] Time of arrival (TOA): The time it takes for a signal to travel from a tag to a reader, which can be used to calculate distance;
[0237] Angle of Arrival (AOA): The angle of the signal's incident direction, used to determine the tag's orientation;
[0238] Three-dimensional acceleration and angular velocity data: usually comes from IMU (inertial measurement unit) and is used to capture the tag's motion state (such as stillness, movement, and vibration);
[0239] Environmental image features: Visual features (such as object edges, textures, and colors) extracted from scene images captured by cameras are used to identify the surrounding environment (such as shelves and obstacles).
[0240] The spatiotemporal multidimensional feature dataset is constructed by performing spatiotemporal alignment and standardized preprocessing on the collected radio wave signals and environmental auxiliary data.
[0241] In one embodiment, specifically, in step eight, the method of constructing a graph structure specifically includes:
[0242] Step 801: Use the corrected coordinates of each label as a graph node and the spatial distance between adjacent labels as edge weights to construct an undirected weighted graph;
[0243] Specifically, first, all the calibrated label coordinates are extracted from the global database, and a graph node is created for each label. The node attributes contain information such as 3D coordinates, label ID, and confidence score.
[0244] Next, calculate the Euclidean distance between any two tags and set a distance threshold (e.g., 5 meters). Only when the spatial distance between two tags is less than the threshold, create an undirected edge between them, and the weight of the edge is set to the actual distance between the two tags.
[0245] Then, spatial index is used to accelerate neighborhood search. For example, an octree is constructed to divide the three-dimensional space into grid cells. Each node only needs to search other nodes in the adjacent grid, which reduces the time complexity of edge construction from O(n 2 ) is reduced to O(n logn);
[0246] Finally, the constructed graph is thinned out, removing edges with excessive weights (e.g., exceeding three times a threshold). Edge pruning algorithms (such as the minimum spanning tree algorithm) are then used to retain key connections, ensuring that the graph structure fully reflects the spatial relationships between labels while also being non-redundant. For example, in a smart warehousing scenario, this method can reduce the number of edges in a graph constructed for 1,000 labels from approximately 500,000 to approximately 20,000, while retaining over 95% of the spatial information.
[0247] Step 802: Perform feature propagation on the graph structure of the undirected weighted graph based on a graph convolutional network or a graph attention network, and optimize global positioning consistency by constraining the positions of adjacent nodes.
[0248] Specifically, first, the constructed undirected weighted graph is input into the graph neural network to create a feature vector containing coordinates and environment information for each label node;
[0249] Next, feature propagation is performed through a graph convolutional network or a graph attention network: each node is allowed to "absorb" information from neighboring nodes, just like learning new information through friends in a social network. For graph convolutional networks, the features of neighboring nodes are weighted and aggregated according to the weights of the edges (i.e., spatial distance). For graph attention networks, they are more "intelligent" and focus on neighboring nodes that are more closely related to themselves.
[0250] Then, a loss function is designed to optimize global positioning consistency. On the one hand, it penalizes pairs of adjacent nodes that deviate significantly from their actual physical distances, ensuring that the relative positions between nodes conform to real-world spatial relationships. On the other hand, it requires that the positioning errors of adjacent nodes do not differ too much, making the entire positioning system smoother and more consistent. At the same time, it gives greater influence to nodes with high confidence, allowing reliable positioning results to have a greater impact on overall optimization.
[0251] Finally, an optimization algorithm is used to continuously adjust network parameters, and the loss function value is made smaller and smaller through multiple iterations, making the label positioning results on the entire graph structure more consistent and accurate globally. For example, in smart warehousing scenarios, the originally scattered label positioning points will gradually "gather" to a more reasonable position through this optimization, greatly improving the overall positioning accuracy and reliability.
[0252] Step 803: Define the global loss function as the sum of the squares of all adjacent node position errors, and iteratively update the node coordinates using the gradient descent algorithm to form a multi-label collaborative localization network.
[0253] Specifically, first, the global loss function is defined as the sum of the squared position errors of all pairs of adjacent nodes. This error refers to the difference between the corrected calculated distance and the actual physical distance between two label nodes that are directly connected in the graph structure. For example, if two labels are 2 meters apart in the real environment, but the distance calculated by the positioning system is 2.3 meters, then their position error is 0.3 meters, which is squared and included in the loss function.
[0254] Next, the gradient descent algorithm is used to iteratively optimize the loss function. Gradient descent is like a person walking down a mountain, taking a step in the steepest direction each time, gradually reaching the foot of the mountain (i.e., the minimum value of the loss function). Specifically, the algorithm calculates the gradient (i.e., the rate of change) of the loss function with respect to each node coordinate, and then adjusts the node coordinates in the opposite direction of the gradient, with the step size of each adjustment controlled by the learning rate. For example, if fine-tuning the coordinates of a node significantly reduces the loss function value, the algorithm will make a larger adjustment in that direction.
[0255] During the iteration process, the coordinates of all nodes are updated simultaneously, forming a collaborative localization network. This means that the position of each tag is no longer determined independently, but its spatial relationship with neighboring tags is taken into account. For example, when the position of a tag changes, its neighboring nodes will also adjust accordingly, just like a group of people holding hands to maintain formation, ensuring the global consistency of the entire localization system.
[0256] To prevent overfitting and control the convergence speed, some optimization strategies are introduced: first, a reasonable learning rate decay mechanism is set to gradually reduce the step size as the number of iterations increases, so that the algorithm can more accurately approach the optimal solution; second, a regularization term is added to penalize excessive coordinate adjustments to prevent the model from being overly sensitive to noise; third, an early stopping strategy is adopted to stop iterations when the loss function on the validation set no longer decreases significantly to prevent overtraining;
[0257] In actual deployment, the algorithm can iterate at a speed of up to 200 times per second on the NVIDIA Jetson AGX Xavier. For a warehouse scenario containing 1,000 tags, it usually converges after 50-100 iterations, keeping the positioning error of more than 90% of the tags within 0.2 meters, meeting the needs of most industrial applications.
[0258] In one embodiment, specifically, the environmental parameter includes at least one of signal strength distribution, multipath fading coefficient, and non-line-of-sight probability;
[0259] Signal strength distribution refers to the statistical characteristics of signal strength (RSSI) within a certain time or space range, such as mean, variance, extreme value, probability density distribution, etc.
[0260] The multipath fading coefficient describes the amplitude and phase changes of radio waves when they reach the receiver through multiple paths (direct, reflected, and diffracted). It is usually expressed by the channel impulse response h(t) or fading factors (such as Rayleigh fading and Ricean fading parameters).
[0261] The non-line-of-sight probability refers to the probability that there are obstacles blocking the signal propagation path (non-line-of-sight, NLOS), which is usually estimated through statistical methods or machine learning models.
[0262] In one embodiment, specifically, the preprocessing includes time domain processing, frequency domain processing, feature normalization, and outlier detection and repair;
[0263] Time domain processing: Use sliding window filtering to eliminate burst noise, and use Kalman filtering or particle filtering to smooth signal fluctuations;
[0264] Frequency domain processing: Apply fast Fourier transform (FFT) to analyze the signal spectrum characteristics and extract specific frequency band features through wavelet transform;
[0265] Feature normalization: Mapping characteristic values such as signal strength and phase to the [0, 1] interval to eliminate signal differences between different readers or tags;
[0266] Outlier detection and repair: Build a signal distribution model based on historical data, and identify and repair abnormal signal values using the threshold method or isolation forest algorithm.
[0267] In summary, the embodiment of the present invention provides a deep learning-based RFID tag positioning error correction method, which deploys a multimodal perception network in the target area, synchronously collects the radio wave signal and environmental auxiliary data of the RFID tag and constructs a spatiotemporal multidimensional feature data set; constructs a joint architecture hybrid model consisting of a spatiotemporal Transformer network and a generative adversarial network, and optimizes parameters through adversarial game to learn robust feature expression; generates adaptive parameters based on a model-independent meta-learning algorithm combined with environmental changes and deploys them to edge nodes to achieve preprocessing, feature extraction and local positioning of real-time signals; the cloud server integrates global data to train an error compensation model, generates corrected coordinates, and then constructs a graph structure, and uses a graph neural network to optimize global positioning consistency; this method effectively integrates multi-source data, adapts to dynamic environments, and improves the accuracy and robustness of RFID tag positioning in complex scenarios.
[0268] Example 2
[0269] This embodiment provides a specific application of a deep learning-based RFID tag positioning error correction method in a smart warehousing scenario:
[0270] In the actual deployment of the smart warehouse center, a multimodal perception network is first constructed. ThingMagic Mercury6 UHF RFID readers are installed at the intersection of shelf aisles (every 10 meters). These readers utilize a cross-polarized antenna design with a coverage radius of 8 meters, support simultaneous reading of over 200 tags, and collect real-time signal strength (RSSI), carrier phase difference (PDOA), time of arrival (ToA), and angle of arrival (AoA). An MPU-6050 IMU module monitors shelf vibration at a 200Hz frequency, while a Hikvision industrial camera captures images at 5 frames per second, extracting visual features of the product placement using ResNet50. All devices achieve microsecond-level time synchronization using the NTP protocol, ensuring data temporal and spatial alignment accuracy of <10μs.
[0271] RFID signal preprocessing uses a three-layer discrete wavelet transform (db4 wavelet basis) to suppress multipath interference, and Kalman filtering is performed on PDOA, ToA, and AoA. The state transfer matrix is set as:
[0272]
[0273] The process noise covariance Q = diag([0.01, 0.001]), and the measurement noise covariance R = 0.05. The IMU data was subjected to FFT to extract vibration energy features. The visual features were reduced to 128 dimensions using PCA. These features were then concatenated with the RFID signal features to form a 138-dimensional spatiotemporal multidimensional feature vector. The dataset was then constructed after Z-score normalization.
[0274] The hybrid model uses a joint architecture of the spatiotemporal Transformer and the conditional generative adversarial network (cGAN). The spatiotemporal Transformer contains a 3-layer encoder and a multi-head self-attention mechanism with 8 heads per layer (key vector dimension 64). The formula is:
[0275]
[0276] Capturing spatial correlations, followed by a bidirectional LSTM to process time series. The cGAN generator inputs a 256-dimensional feature vector and an environment category label, generating robust features through four layers of transposed convolution. The discriminator uses a four-layer convolutional network and performs adversarial training using the WGAN-GP loss function with a gradient penalty coefficient of λ = 10.
[0277] In the meta-learning-driven parameter adaptation process, environmental parameters such as the standard deviation of the signal strength distribution and the multipath fading coefficient are monitored in real time. Parameter updates are triggered when the Euclidean distance from the historical baseline exceeds 0.3σ. The multimodal perception network is controlled to collect 100 new scene samples. Five gradient updates are performed using the MAML algorithm (learning rate 0.01) to generate temporary parameters θ'. Parameter importance is evaluated using the Fisher information matrix. L2 regularization (weight decay coefficient 0.001) is applied to the first two layers of the spatiotemporal Transformer. The generated adaptive parameters are then deployed to the edge node (NVIDIA Jetson Nano). The entire process is completed within 30 seconds.
[0278] The real-time positioning process at the edge includes: time-domain Butterworth low-pass filtering (cutoff frequency 100Hz), FFT frequency domain transformation, and IQR outlier repair. The preprocessed signal is input into the knowledge distilled and compressed MobileNetV2, and the Levenberg-Marquardt algorithm is used for iterative optimization in combination with the adjacent tag positions. The single-label inference time is <50ms. The cloud server integrates the global environment map information, trains the Kalman filter and CNN hybrid model, inputs local coordinates, environmental features and historical error data, and outputs the error correction amount. For each local result, the compensation results of multiple methods are fused according to the environmental characteristics (such as the occlusion level of metal shelves), and verified by the consistency of adjacent tag positions (threshold 30cm) and global topological constraints.
[0279] In the graph neural network collaborative optimization phase, the corrected coordinates are used as graph nodes, and the spatial distance between adjacent labels (threshold 2 meters) is used as the edge weight to construct an undirected weighted graph. A two-layer graph convolutional network (64 neurons per layer) is used for feature propagation. The global loss function is defined as the sum of squared errors between adjacent node positions. Node coordinates are iteratively updated using the Adam optimizer (learning rate 0.001). Global topology verification is performed after every five updates to ensure that the row and column alignment error of the shelves is less than 0.1 meter.
[0280] In a 2,000-square-meter smart warehouse test, this solution achieved an average positioning error of 0.82 meters in the metal shelving area, a 53% improvement over traditional RSSI fingerprint matching (1.75 meters). Error fluctuations in dynamic scenarios were reduced by 70%. After adjusting the shelf layout, the model recovered 92% accuracy within 30 seconds through meta-learning. Edge node inference latency was <50ms, and cloud-based compensation latency was <200ms, meeting real-time requirements. Through multimodal data fusion, dynamic adaptive models, edge-cloud collaborative architecture, and graph neural network optimization, the accuracy and robustness of RFID tag positioning in complex environments were significantly improved.
[0281] In summary, this embodiment provides an application of a deep learning-based RFID tag positioning error correction method in smart warehousing scenarios. It uses microsecond-level spatiotemporal alignment technology (accuracy <10ms) to fuse RFID signals (including RSSI, PDOA, etc.), IMU vibration data and visual image features, effectively solving the problem of single signal failure in complex scenarios such as strong occlusion, and significantly improving positioning robustness; based on the model-independent meta-learning (MAML) algorithm combined with knowledge protection constraints, it realizes rapid environmental adaptation of model parameters (update completed within 30 seconds) and suppresses error accumulation in dynamic environments; constructs an edge-cloud collaborative architecture, and the edge uses a lightweight model to realize real-time signal preprocessing and fast inference (delay <50ms). The cloud uses global data to train a hybrid error compensation model and generate corrections, balancing computational efficiency and positioning accuracy; introduces a graph convolutional network (GCN) to construct a spatial relationship graph structure between labels, and achieves multi-label positioning consistency through feature propagation and global loss optimization, which improves positioning accuracy by 15% compared with traditional geometric constraint methods, forming a complete technical system covering data collection, model adaptation, edge-cloud collaboration and global optimization.
[0282] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0283] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0284] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0285] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0286] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.
[0287] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0288] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for correcting RFID tag positioning errors based on deep learning, characterized in that: The following steps are involved: Deploy a multimodal perception network in the target area to simultaneously collect radio wave signals from RFID tags and environmental auxiliary data to construct a spatiotemporal multidimensional feature dataset; Build a joint architecture hybrid model consisting of a spatiotemporal Transformer network and a generative adversarial network; Input the multidimensional feature dataset into the spatiotemporal Transformer network to generate a feature encoding vector including spatiotemporal correlation information; Based on the feature encoding vector, a training framework based on the generative adversarial network is constructed. The hybrid model parameters are optimized through the adversarial game process, so that the hybrid model can learn robust feature expressions. Initialize the hybrid model parameters based on the model-independent meta-learning algorithm and automatically generate adaptive parameters according to environmental changes; Deploy the adaptation parameters to the edge computing nodes, configure the edge hybrid model parameters, preprocess and extract features of the real-time collected radio wave signals, and generate local positioning results; The cloud server receives the local positioning results uploaded by the edge node, performs global hybrid model training and error compensation, outputs the global error compensation amount, and generates the corrected coordinates; The corrected coordinates and the position information of adjacent labels are constructed into a graph structure to form a globally consistent multi-label collaborative localization network.
2. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, characterized in that: The spatial attention module of the spatiotemporal Transformer network adopts a multi-head self-attention mechanism, and the calculation formula is: Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively, d k is the key vector dimension.
3. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, characterized in that: The specific steps of automatically generating the adaptation parameters are as follows: Monitor current environmental parameters in real time and calculate the deviation between current environmental parameters and historical benchmark values; According to the deviation results, the multimodal perception network is controlled to collect labeled sample data in the new scenario; Based on the model-independent meta-learning algorithm, the model parameters initialized by historical tasks are used to perform gradient updates on the newly collected labeled sample data to generate temporary parameters; Evaluate the importance of temporary parameters to historical tasks and impose knowledge protection constraints on key parameters during parameter updating; The meta-learning update results and knowledge protection constraints are integrated to generate the final adaptation parameters that adapt to the current environment and deploy them to the edge computing nodes.
4. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, characterized in that: The specific steps of generating the local positioning result are as follows: Preprocessing the radio wave signals collected in real time; The preprocessed radio wave signals are input into the spatiotemporal Transformer network to extract the spatial correlation features of multi-reader signals; Based on the deployed adaptive parameters, the edge hybrid model is configured, and the spatial correlation features are processed through a lightweight network compressed by knowledge distillation to generate initial positioning coordinates. The initial positioning coordinates of the current tag are spatially constrained and matched with the position information of adjacent known tags, and the initial positioning coordinates are iteratively optimized to generate the final local positioning result.
5. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, wherein: The specific steps of generating the corrected coordinates are as follows: The cloud server receives the local positioning results of multiple edge nodes, aligns the data in time and space dimensions, and integrates the global environment map information; Using historical positioning data and real coordinates, we train a global error compensation model that combines traditional statistical methods with deep learning to learn the error distribution patterns under different environments. For each local positioning result, the error correction amount is calculated through the global model according to its environmental characteristics and historical error patterns, and the compensation results of multiple methods are integrated; The correction amount is applied to the local positioning results to generate corrected coordinates. The rationality of the results is verified by the consistency of adjacent label positions and global topological constraints, and outliers are handled.
6. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, characterized in that: The generative adversarial network adopts a conditional generative adversarial network architecture, and the generator input of the conditional generative adversarial network architecture includes a feature encoding vector and an environment category label.
7. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, characterized in that: The radio wave signal includes signal strength, carrier phase difference, arrival time and arrival angle; the environmental auxiliary data includes three-dimensional acceleration, angular velocity data and environmental image features; The spatiotemporal multidimensional feature dataset is constructed by performing spatiotemporal alignment and standardization preprocessing on the collected radio wave signals and environmental auxiliary data.
8. The method for correcting RFID tag positioning errors based on deep learning according to claim 1, wherein: The method for constructing a graph structure specifically includes: The corrected coordinates of each label are used as graph nodes, and the spatial distance between adjacent labels is used as the edge weight to construct an undirected weighted graph; Based on graph convolutional networks or graph attention networks, feature propagation is performed on the graph structure of undirected weighted graphs, and global positioning consistency is optimized by constraining the positions of adjacent nodes. The global loss function is defined as the sum of the squares of the position errors of all adjacent nodes. The node coordinates are iteratively updated using the gradient descent algorithm to form a multi-label collaborative localization network.
9. The method for correcting RFID tag positioning errors based on deep learning according to claim 3, wherein: The environmental parameters include at least one of signal strength distribution, multipath fading coefficient, and non-line-of-sight probability.
10. The method for correcting RFID tag positioning errors based on deep learning according to claim 4, characterized in that: The preprocessing includes time domain processing, frequency domain processing, feature normalization, and outlier detection and repair.
Citation Information
Cited By
Spectral overlapping peak decomposition method, device, equipment and medium
CN120971470A
Metabonomics data batch effect correction method based on generative adversarial network
CN121281627A
Label positioning method and device, electronic equipment and storage medium
CN121486752A
Wearable electroencephalogram signal processing method and device based on feature fusion
CN122310449A