Intelligent archive checking method and system based on thermodynamic diagram

By using a heatmap-based intelligent inventory method, combined with radio frequency-vision bimodal mapping tensor and manifold constraint denoising technology, the problem of insufficient accuracy of traditional RFID inventory in complex environments is solved, and efficient file inventory and tag anomaly handling are achieved.

CN121903526AInactive Publication Date: 2026-04-21SUZHOU DUNYI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU DUNYI TECHNOLOGY CO LTD
Filing Date
2026-01-14
Publication Date
2026-04-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional RFID inventory technology is susceptible to interference when dealing with large-scale, high-concurrency readings and environmental obstructions, resulting in insufficient accuracy and completeness of inventory records. It is also difficult to handle abnormal situations such as tag detachment or signal failure, increasing the cost of manual verification.

Method used

A heatmap-based intelligent inventory method is adopted. Through a closed-loop feedback mechanism of manifold constraint denoising and dual fault detection, combined with radio frequency-vision bimodal mapping tensor, manifold constraint boundaries are constructed using the file arrangement rules to generate real-time heatmaps. Faults are determined by difference matrix and logical operations, and the global mapping is dynamically corrected.

Benefits of technology

It significantly improves the accuracy and robustness of archive inventory, enhances the system's automation and fault tolerance, effectively eliminates physical topological distortion, and achieves precise consistency constraints between real-time heat maps and the spatial distribution of archive entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903526A_ABST
    Figure CN121903526A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent archive checking method and system based on a thermodynamic diagram, and relates to the technical field of information management, and the system comprises the following modules: a bimodal reference construction module which is used for generating a radio frequency-vision bimodal mapping tensor; the wide-area scanning and exception locking module is used for generating a visual wake-up queue; the high-resolution sampling flow generation module is used for generating an original local image flow; the manifold constraint denoising generation module is used for improving the MaxViT network based on manifold geometric constraints and predicting and generating a real-time thermodynamic diagram; the dual fault discrimination module is used for constructing a differential matrix and logic and dual fault discrimination mechanism; the virtual identification and index increment updating module is used for updating index increment; and the tensor iteration updating module is used for executing iteration updating. According to the method, the limitation that a traditional inventory method is easily interfered, lacks visual verification and is weak in physical topology constraint is overcome, and an efficient solution is provided for intelligent archive inventory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information management technology, and in particular to a method and system for intelligent inventory of archives based on heat maps. Background Technology

[0002] With the continuous expansion of archives and the dynamic complexity of the storage environment, traditional single-RFID inventory technology faces severe challenges in dealing with large-scale high-concurrency readings and environmental obstructions. Existing inventory solutions mainly rely on RFID signals for tag positioning and quantity verification. Although the non-line-of-sight characteristic of radio frequency communication improves the reading range, it is inherently limited by the physical propagation characteristics of electromagnetic waves. This method, which is solely based on radio frequency signals, is susceptible to interference from metal mobile shelving, signal conflicts between adjacent tags, and multipath effects, leading to missed or misreads when building inventory indexes, thus limiting the accuracy and completeness of archive inventory. In addition, classic inventory methods often struggle to incorporate visual semantic information for secondary verification and repair of abnormal signals. This makes it impossible to locate targets through the geometric constraints of physical space when dealing with anomalies such as tag detachment or signal failure, significantly increasing the cost of manual verification.

[0003] Therefore, how to provide a heatmap-based intelligent inventory method for archives is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] This invention proposes a heatmap-based intelligent archive inventory method. Through a closed-loop feedback mechanism of manifold constraint denoising and dual fault detection, it constructs manifold constraint boundaries based on the archive arrangement patterns using an improved MaxViT network. During denoising, latent variables are corrected to generate a real-time heatmap conforming to the physical topology. Faults are determined by combining differential matrix and logical operations of radio frequency anomaly results, and the global mapping is dynamically corrected through virtual identifier mounting and tensor iterative updates. This invention overcomes the limitations of traditional inventory methods, such as susceptibility to interference, lack of visual verification, and weak physical topology constraints, providing an efficient solution for intelligent archive inventory.

[0005] According to an embodiment of the present invention, a method for intelligent inventory of archives based on heatmaps includes the following steps:

[0006] S1. Collect panoramic images of the standard warehouse and extract spine features to construct a baseline heat map. Simultaneously read RFID to establish an index and bind the ID with spatial coordinate hash to generate an RFID-visual bimodal mapping tensor.

[0007] S2. Control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual dual-modal mapping tensor, and if data abnormality is detected, lock the target area and generate a visual wake-up queue.

[0008] S3. Respond to the visual wake-up queue command, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream;

[0009] S4. Based on manifold geometric constraints, the MaxViT network is improved by mapping the original local image stream to the latent space and injecting Gaussian noise. The manifold constraint boundary is constructed by utilizing the file arrangement rules, and iterative denoising prediction is performed according to the time step to generate a real-time heat map.

[0010] S5. Construct a dual fault discrimination mechanism of differential matrix and logical AND, align the real-time heat map with the reference heat map in time and space and calculate the differential matrix, extract differential features and perform logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost.

[0011] S6. If the label is determined to be invalid, create a virtual pending identifier, extract the coordinates of high-probability anchor points in the real-time heatmap and mount them, and perform incremental updates on the index.

[0012] S7. Perform iterative updates based on the updated index, use real-time heatmap data to overlay and map historical data, and complete the iteration of the radio frequency-visual dual-modal mapping tensor.

[0013] Optionally, S1 specifically includes:

[0014] S11. Collect panoramic image sequences of the standard warehouse, construct a multidimensional spatiotemporal probability density field through spine edge detection and texture feature extraction algorithms, and calibrate the multidimensional spatiotemporal probability density field as a benchmark heat map.

[0015] S12. Analyze the warehouse radio frequency signal to establish an electronic tag index, determine the pixel-level physical coordinates based on the baseline heat map, and perform a hash operation on the unique ID of the tag and the physical coordinates to generate a spatial hash key-value pair.

[0016] S13. By fusing spatial hash key-value pairs, electronic tag indexes, and multidimensional spatiotemporal probability density fields, tensor mapping operations are performed to generate a radio frequency-visual bimodal mapping tensor.

[0017] Optionally, S2 specifically includes:

[0018] S21. Control the inventory equipment to perform wide-area traversal scanning along the preset operation path, capture radio frequency signals in real time and parse them to generate a real-time tag set;

[0019] S22. Perform hash mapping verification on the real-time tag set and the radio frequency-vision bimodal mapping tensor, and calculate the tag existence deviation and spatial topology consistency index.

[0020] S23. Construct an abnormal state discrimination function based on the existence deviation and spatial topology consistency index. If the function value exceeds the preset information interval, it is determined that the radio frequency data is abnormal and the area locking mechanism is triggered.

[0021] S24. Analyze the index nodes in the abnormal state discrimination function and reverse map them to the global physical coordinate system of the baseline heatmap to accurately locate the abnormal physical coordinate region;

[0022] S25. Encapsulate the state characteristics of the abnormal physical coordinate region into a structured visual wake-up request, sort it according to the time sequence of the abnormality level, and generate a visual wake-up queue.

[0023] Optionally, S3 specifically includes:

[0024] S31. Analyze the structured request parameters in the visual wake-up queue, solve the mechanical kinematics model based on the abnormal physical coordinate region, and generate spatial pose adjustment parameters.

[0025] S32. Respond to the spatial pose adjustment parameters, control the imaging device to point to the target abnormal area, perform fixed-point high-resolution optical sampling, and acquire multiple frames of high-dimensional image data;

[0026] S33. Perform temporal registration and streaming reconstruction on multi-frame high-dimensional image data to generate the original local image stream of the target anomaly region.

[0027] Optionally, the improved MaxViT network includes a multi-scale block attention extraction layer, a multi-scale channel attention extraction layer, a cascaded downsampling module, a latent space Gaussian injection layer, and a manifold denoising iterative generation layer:

[0028] The multi-scale block attention extraction layer is used to divide and flatten the input original local image stream into image block sequence tensors. The image block sequence tensors are then mapped to query key-value pair tensors through a multilayer perceptron. Local self-attention calculation is performed on the query key-value pair tensors based on a window partitioning strategy, and global extended self-attention calculation is performed on the query key-value pair tensors based on a grid partitioning strategy. Local detail features and global semantic features are fused to output a multi-scale block feature tensor.

[0029] The multi-scale channel attention extraction layer is used to perform a transpose operation on the multi-scale block feature tensor, dividing it into multi-head sub-feature tensors along the channel dimension; it generates channel query key-value pair tensors through linear transformation, and performs channel self-attention operation to capture the feature dependencies between channels. After aggregation and transpose processing, it outputs the enhanced channel feature tensor.

[0030] The cascaded downsampling module is used to flatten and linearly project the enhanced channel feature tensor onto a latent space of a preset dimension, and output the latent space feature tensor.

[0031] The latent space Gaussian injection layer is used to calculate the statistical moments of the latent space feature tensor and generate the mean vector and variance vector. Based on the preset time step, the standard Gaussian noise tensor is sampled, and the mean vector and variance vector are used to perform affine transformation and reparameterization operations on the standard Gaussian noise tensor to obtain the noisy latent variable tensor.

[0032] The manifold denoising iterative generation layer is used to input the noisy latent variable tensor into the denoising U-Net network and perform iterative noise residual prediction along the reverse time step direction. The archive arrangement pattern constructs the manifold constraint boundary, and during the iterative denoising process, manifold projection clipping and range constraints are performed on the intermediate latent variable tensor to correct the feature distribution to conform to the physical topology of the archive. Based on the corrected latent variables and the noise residual at the current time step, combined with the preset noise scheduler parameters, the denoised latent variable tensor is recursively calculated. The final denoised latent variable tensor is mapped back to the pixel space to generate a real-time heatmap.

[0033] Optionally, the difference matrix and logical AND dual fault detection mechanism specifically includes:

[0034] Based on Fourier Merlin transform, rigid transformation parameters are estimated between real-time heatmaps and reference heatmaps. The real-time heatmaps are then resampled according to the transformation parameters to generate a registered target image sequence.

[0035] A pixel-wise Euclidean distance metric is performed on the registered target image sequence and the baseline heatmap, and the difference metric results are mapped to the Gaussian probability space to generate a confidence distribution matrix.

[0036] The proportion of low-confidence regions below a preset threshold in the statistical confidence distribution matrix is ​​converted into the first binary discriminant variable, and the radio frequency anomaly detection result is mapped into the second binary discriminant variable.

[0037] The first binary discriminant variable and the second binary discriminant variable are input to an AND gate for processing, and the joint judgment state is output. Based on the information entropy feature of the confidence distribution matrix, the classification label of label failure or true loss is output.

[0038] Optionally, S6 specifically includes:

[0039] S61. Initialize the virtual pending identifier based on the fault determination signal, and define an attribute structure containing a unique identifier, a timestamp, and a status bit vector;

[0040] S62. Perform morphological gradient operation on the real-time heatmap to extract local extreme value regions, apply non-maximum suppression algorithm to filter neighborhood redundant responses, use the mapping matrix from image coordinate system to global coordinate system to perform spatial transformation on extreme point coordinates, and output a set of high-probability anchor point positions.

[0041] S63. Calculate the Euclidean distance metric between the predicted location of the virtual undetermined identifier and the set of high-probability anchor points. Based on the principle of minimizing the distance cost function, establish the nearest neighbor matching relationship and perform the coordinate mounting operation of the virtual identifier.

[0042] S64. Concatenate the attribute structure of the virtual pending identifier with the high-probability anchor position using a tensor to generate an incremental update data packet. Perform an update operation on the global hash table based on the time-series index key value to complete the index update.

[0043] Optionally, S7 specifically includes:

[0044] S71. Based on the hash collision resolution strategy, read the updated index key-value pairs, and parse the corresponding radio frequency identity feature vector and visual geometric coordinate matrix through key-value inversion to calculate the source addressing pointer and target storage base address of the data stream.

[0045] S72. Perform a traversal query in the historical radio frequency-visual dual-modal mapping tensor library according to the target storage base address, and extract historical tensor data slices whose timestamp index is within the preset sliding window interval;

[0046] S73. Calculate the cosine similarity measure between real-time heatmap data and historical tensor data slices, introduce an exponentially decaying time forgetting factor to calculate dynamic fusion weights, and generate the corresponding dynamically updated mask matrix.

[0047] S74. Perform pixel-level weighting on the real-time heatmap data using a dynamically updated mask matrix, and write the weighted data into the corresponding storage unit of the historical tensor data slice through tensor assignment operation to update the numerical field of the mapped tensor.

[0048] S75. Calculate the Frobenius norm error of the mapping tensor before and after the update, and determine whether the error has converged. If the error is greater than the preset threshold, correct the index parameters based on the updated tensor value and repeat the update steps until the convergence condition is met, and output the final RF-vision bimodal mapping tensor.

[0049] According to an embodiment of the present invention, an intelligent archive inventory system based on a heat map includes the following modules:

[0050] The dual-modal benchmark construction module is used to acquire panoramic images of standard warehouses and extract spine features to construct benchmark heat maps. Simultaneously, it reads RFID to establish indexes and binds IDs with spatial coordinate hashes to generate radio frequency-visual dual-modal mapping tensors.

[0051] The wide-area scanning and anomaly locking module is used to control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual bimodal mapping tensor, and lock the target area and generate a visual wake-up queue if an anomaly is detected.

[0052] The high-resolution sampling stream generation module is used to respond to visual wake-up queue instructions, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream.

[0053] The manifold constraint denoising generation module is used to improve the MaxViT network based on manifold geometric constraints. It maps the original local image stream to the latent space and injects Gaussian noise. It constructs the manifold constraint boundary by using the file arrangement rules and performs iterative denoising prediction to generate a real-time heat map according to the time step.

[0054] The dual fault discrimination module is used to construct a dual fault discrimination mechanism of differential matrix and logical AND. It aligns the real-time heat map with the reference heat map in time and space and calculates the differential matrix. It extracts differential features and performs logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost.

[0055] The virtual identifier and index incremental update module is used to create a virtual pending identifier if the tag is determined to be invalid, extract the coordinates of high-probability anchor points in the real-time heatmap for attachment, and perform incremental updates on the index.

[0056] The tensor iteration update module is used to perform iterative updates based on the updated index, and to use real-time heatmap data to overwrite historical data to complete the iteration of the radio frequency-visual dual-modal mapping tensor.

[0057] The beneficial effects of this invention are:

[0058] This invention significantly improves the accuracy of anomaly identification and system robustness during archive inventory in complex warehouse environments by introducing a closed-loop feedback optimization mechanism that combines manifold-constrained denoising generation with dual fault discrimination. By innovatively constructing manifold constraints with the archive arrangement patterns as boundaries, manifold projection pruning and range constraints are applied to latent space features during the iterative denoising process of the improved MaxViT network. This effectively eliminates physical topological distortion in pure visual generation, achieving precise consistency constraints between real-time heatmaps and the spatial distribution of archive entities. A cross-modal data association is constructed using a radio frequency-visual dual-modal mapping tensor, deeply exploring the potential logical connections between radio frequency signals and visual semantics. In the fault discrimination stage, difference matrix and logic gate operations are integrated to enhance the system's multi-dimensional perception capabilities. Furthermore, iterative self-correction of the index is achieved using dynamically updated masks and Frobenius norm convergence verification, thereby significantly improving the efficiency, automation level, and fault tolerance for tag failures and environmental interference during archive inventory. Attached Figure Description

[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0060] Figure 1 This is an overall flowchart of a heatmap-based intelligent archive inventory method proposed in this invention.

[0061] Figure 2 This is a flowchart illustrating the working principle of the improved MaxViT network for an intelligent archive inventory method based on heatmaps proposed in this invention.

[0062] Figure 3 This is a flowchart illustrating the working principle of the differential matrix and logic AND dual fault discrimination mechanism of the intelligent archive inventory method based on heatmap proposed in this invention.

[0063] Figure 4 This is a schematic diagram of the structure of an intelligent archive inventory system based on heatmap proposed in this invention. Detailed Implementation

[0064] The invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0065] refer to Figure 1-3 A method for intelligent inventory of archives based on heatmaps includes the following steps:

[0066] S1. Collect panoramic images of the standard warehouse and extract spine features to construct a baseline heat map. Simultaneously read RFID to establish an index and bind the ID with spatial coordinate hash to generate an RFID-visual bimodal mapping tensor.

[0067] S2. Control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual dual-modal mapping tensor, and if data abnormality is detected, lock the target area and generate a visual wake-up queue.

[0068] S3. Respond to the visual wake-up queue command, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream;

[0069] S4. Based on manifold geometric constraints, the MaxViT network is improved by mapping the original local image stream to the latent space and injecting Gaussian noise. The manifold constraint boundary is constructed by utilizing the file arrangement rules, and iterative denoising prediction is performed according to the time step to generate a real-time heat map.

[0070] S5. Construct a dual fault discrimination mechanism of differential matrix and logical AND, align the real-time heat map with the reference heat map in time and space and calculate the differential matrix, extract differential features and perform logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost.

[0071] S6. If the label is determined to be invalid, create a virtual pending identifier, extract the coordinates of high-probability anchor points in the real-time heatmap and mount them, and perform incremental updates on the index.

[0072] S7. Perform iterative updates based on the updated index, use real-time heatmap data to overlay and map historical data, and complete the iteration of the radio frequency-visual dual-modal mapping tensor.

[0073] In this embodiment, S1 specifically includes:

[0074] S11. Deploy high-resolution panoramic imaging equipment to perform multi-angle surround shooting of the standard warehouse, acquiring a panoramic image sequence covering the compact shelving area; use the Canny edge detection operator combined with the Sobel filter to perform convolution operation on the image sequence, extract the vertical gradient features of the spine contour, and apply the local binary mode algorithm to statistically analyze the micro-texture distribution pattern of the spine region of the archives, constructing an initial spatiotemporal probability density field based on the occurrence frequency of spine feature points in different grid cells; divide the image sequence into grid cells with a side length of 10 pixels, statistically analyze the occurrence frequency of spine feature points in each grid cell to generate a local probability density, and combine the spatial adjacency relationship of adjacent grid cells to utilize Markov... A random field model is constructed to establish a spatial smoothing constraint term. The weighted average of the local probability density and the spatial smoothing constraint term is calculated to obtain the joint probability distribution. A Gaussian kernel with a standard deviation of 1.5 is used to perform convolutional smoothing on the joint probability distribution to eliminate noise interference. The minimum-maximum normalization algorithm is used to map the values ​​to a grayscale range of 0 to 255. The normalized numerical matrix is ​​mapped to the global physical coordinate system of the warehouse through homography matrix transformation to obtain a pixel-level aligned initial heatmap matrix. The empty regions in the initial heatmap matrix are filled using a bilinear interpolation algorithm to generate a continuously distributed baseline heatmap matrix, which is used as the visual baseline map for subsequent radio frequency-vision bimodal mapping.

[0075] S12. The RFID module is activated to perform a full-power scan of the warehouse area at a center frequency of 840 MHz, receiving signal data returned by the electronic tags and establishing an electronic tag index based on the 96-bit electronic product code of the tags; the pixel position corresponding to each tag is extracted using the reference heat map, and converted into physical coordinates in centimeters using the global scale of the warehouse, with the accuracy retained to two decimal places; the unique identifier of the tag is concatenated with the physical coordinate value into a string, which is then input into the cyclic redundancy check algorithm to perform a hash operation, generating a spatial hash key-value pair composed of 16 hexadecimal characters, realizing a strong binding between the tag ID and the physical location;

[0076] S13. Associate the generated spatial hash key-value pairs with the identity information in the electronic tag index, and fill the corresponding index nodes with the probability values ​​of the gray range from 0 to 255 in the multidimensional spatiotemporal probability density field; construct a three-dimensional data matrix containing radio frequency identity, physical coordinates and visual probability, and convert the three-dimensional data matrix into a radio frequency-visual bimodal mapping tensor with a dimension of 3 through a tensor mapping function, so as to store radio frequency identity information and visual distribution features in a unified data structure at the same time.

[0077] In this embodiment, S2 specifically includes:

[0078] S21. Control the inventory robot equipped with an ultra-high frequency radio frequency reader to move along the preset S-shaped operation path in the warehouse aisle at a speed of 1.5 meters per second. Use the reader to scan the coverage area in real time at a sampling frequency of 200 times per second, capture tag signals and extract 96-bit electronic product codes, and perform signal analysis based on the received signal strength indication value to generate a real-time tag set containing all the tag identifications read.

[0079] S22. Perform hash mapping verification between the identity identifiers in the real-time tag set and the indexes stored in the radio frequency-vision bimodal mapping tensor. By comparison, filter out the tags that exist in the mapping tensor but are missing in the real-time tag set. Divide the number of missing tags by the total number of tags in the mapping tensor to obtain the tag existence deviation. At the same time, calculate the Euclidean distance between the physical coordinates of the tags in the real-time tag set and the reference physical coordinates stored in the mapping tensor. The proportion of tags with a distance value exceeding 3 cm to the total number of tags in the real-time tag set is recorded as the spatial topology consistency index.

[0080] S23. Multiply the tag existence deviation and spatial topology consistency index by weight coefficients of 0.6 and 0.4 respectively, and sum them to construct an abnormal state discrimination function. When the calculated function value exceeds the confidence interval threshold of 0.75, it is determined that the radio frequency data is abnormal and the area locking mechanism is immediately triggered.

[0081] S24. Analyze the index node of the triggering abnormal state discrimination function, and use the coordinate mapping relationship of the baseline heat map to reverse map the index node to the global physical coordinate system of the warehouse, locate the abnormal area with a physical coordinate deviation greater than 5 cm, and lock the boundary range of the area as a rectangular area with a length of 50 cm and a width of 30 cm.

[0082] S25. Encapsulate the label identity, coordinate range and anomaly type of the abnormal physical coordinate region into a structured visual wake-up request containing a timestamp, sort it in descending order of the anomaly value, and generate a visual wake-up queue containing priority information.

[0083] In this embodiment, S3 specifically includes:

[0084] S31. Analyze the structured request parameters in the visual wake-up queue, extract the center point coordinates and boundary range of the abnormal physical coordinate region, use the mechanical kinematic model established by the DH parameter method to perform inverse kinematics solution, calculate the rotation angle and displacement of each joint axis, and generate spatial pose adjustment parameters containing six degrees of freedom spatial position and attitude information.

[0085] S32, responding to spatial pose adjustment parameters, controlling the servo driver to adjust the joint angle of the robotic arm, guiding the high-resolution imaging device to point to the target abnormal area, performing fixed-point optical sampling at a rate of 25 frames per second, continuously acquiring 20 million pixel images, and obtaining multi-frame high-dimensional image data;

[0086] S33. The scale-invariant feature transformation algorithm is used to extract feature points from the acquired multi-frame high-dimensional image data and perform feature matching. The transformation matrix between images is calculated based on the feature matching results. Temporal registration is performed using the transformation matrix. The registered image data is stacked between frames and streamed to generate the original local image stream of the target abnormal region.

[0087] In this embodiment, the improved MaxViT network includes a multi-scale block attention extraction layer, a multi-scale channel attention extraction layer, a cascaded downsampling module, a latent space Gaussian injection layer, and a manifold denoising iterative generation layer:

[0088] The multi-scale block attention extraction layer is used to divide the input original local image stream into image blocks of 7 pixels in both width and height, flatten the image blocks into image block sequence tensors, and use a multilayer perceptron containing two neural networks to map the image block sequence tensors into query key-value pair tensors. Based on a window partitioning strategy, the query key-value pair tensors are divided into local windows of 7 pixels in both width and height to perform local self-attention calculation. At the same time, based on a grid partitioning strategy, the query key-value pair tensors are divided into grids of 7 pixels in both width and height to perform global extended self-attention calculation. The calculated local detail features and global semantic features are fused to output a multi-scale block feature tensor.

[0089] The multi-scale channel attention extraction layer is used to perform a transpose operation on the multi-scale block feature tensor, dividing the channel dimension into multi-head sub-feature tensors with 4 heads; it generates channel query key-value pair tensors through linear transformation, and performs channel self-attention operation to capture the feature dependencies between channels. The calculation results are aggregated and transposed again to output the enhanced channel feature tensor.

[0090] The cascaded downsampling module is used to flatten the enhanced channel feature tensor, and then use a linear projection layer to map it to a latent space of dimension 512, outputting the latent space feature tensor.

[0091] The latent space Gaussian injection layer is used to calculate the statistical moments of the latent space feature tensor and generate the mean vector and variance vector. Based on a specific step size in a preset 1000 time steps, the standard Gaussian noise tensor is sampled, and the mean vector and variance vector are used to perform affine transformation and reparameterization operations on the standard Gaussian noise tensor to obtain the noisy latent variable tensor.

[0092] The manifold denoising iterative generation layer is used to input the noisy latent variable tensor into the denoising U-Net network and perform iterative noise residual prediction along the reverse time step direction. A manifold constraint boundary is constructed based on the archive arrangement pattern. During the iterative denoising process, a manifold projection pruning operation is performed on the intermediate latent variable tensor to restrict the values ​​to a range of -3 to +3, correcting the feature distribution to conform to the physical topology of the archive. Based on the corrected latent variables and the noise residual at the current time step, combined with preset noise scheduler parameters, the denoised latent variable tensor is recursively calculated. Finally, the denoised latent variable tensor is mapped back to the pixel space to generate a real-time heatmap.

[0093] In this embodiment, the difference matrix and logical AND dual fault detection mechanism specifically includes:

[0094] Rigid transformation parameter estimation is performed on real-time and reference heatmaps based on Fourier-Melin transform. Fast Fourier transform is applied to the input 640 x 480 pixel real-time and reference heatmaps, and the transformation results are converted to logarithmic polar coordinates. Bilinear interpolation is used to divide the radial coordinates into 60 sampling intervals and the angular coordinates into 360 sampling intervals, constructing a logarithmic polar coordinate transformation matrix. In the transformed frequency domain amplitude spectrum, the rotation angle and scaling ratio are calculated using the phase correlation method, calibrating the detection accuracy of the rotation angle to 0.1 degrees and the detection accuracy of the scaling ratio to 0.01 times. Based on the calculated rotation angle and scaling ratio parameters, bicubic interpolation resampling is performed on the real-time heatmap, and pixel displacement compensation is applied in the horizontal and vertical directions to generate a registered target image sequence.

[0095] A pixel-by-pixel Euclidean distance metric is performed on the registered target image sequence and the reference heatmap. For an image region of size 640 x 480 pixels, the absolute value of the difference between the gray values ​​of each pixel in the two heatmaps is calculated, and the absolute difference is normalized by dividing by 255. The normalized difference result is input into a Gaussian distribution model with a probability density function of mean 0 and standard deviation 1 to calculate the cumulative distribution probability value corresponding to each pixel difference. The calculated probability value is multiplied by 255 and converted into an 8-bit unsigned integer to generate a confidence distribution matrix with values ​​ranging from 0 to 255, where 255 represents complete consistency and 0 represents the maximum difference.

[0096] The proportion of low-confidence regions with values ​​below a preset threshold of 0.3 in the statistical confidence distribution matrix is ​​calculated. The proportion of low-confidence pixels to the total number of pixels is calculated. When the proportion exceeds 15%, the first binary discriminant variable is set to 1; otherwise, it is set to 0. At the same time, the radio frequency anomaly detection result is mapped to the second binary discriminant variable.

[0097] The first binary discriminant variable and the second binary discriminant variable are input to an AND gate for processing. When both variables are 1, the joint judgment state is output as abnormal; otherwise, it is normal. Based on the information entropy feature of the confidence distribution matrix, when the information entropy value is lower than 1.5, the classification label of the label failure is output; when the information entropy value is higher than 1.5, the classification label of the true loss is output.

[0098] In this embodiment, S6 specifically includes:

[0099] S61. Receive the fault determination signal, initialize the virtual pending identifier according to the abnormality level in the signal, and define an attribute structure containing a 64-bit unique identifier, a timestamp accurate to milliseconds, and a status bit vector composed of 3 Boolean values ​​to record the identity and status of the virtual identifier.

[0100] S62. Perform morphological dilation and erosion operations on the real-time heatmap. Use a rectangular structuring element with a width and height of 3 pixels to calculate the pixel grayscale difference between the dilated and eroded maps. Extract pixels with a grayscale change rate exceeding 50 as local extremum regions. Apply a non-maximum suppression algorithm to establish a circular neighborhood window with a radius of 5 pixels centered on each extremum point. If the grayscale value of the window center point is not the maximum, set it to zero to filter redundant responses in the neighborhood. Use a preset 3x3 image coordinate system to global coordinate system homography mapping matrix to perform spatial transformation on the pixel coordinates of the retained extremum points after filtering. Convert the pixel coordinates to physical coordinates through matrix multiplication. Sort according to confidence scores from high to low and output a set of high-probability anchor point positions containing the top 10 candidate points.

[0101] S63. Calculate the Euclidean distance between the predicted position of the virtual undetermined identifier and each point in the set of high-probability anchor points, measure the straight-line interval in physical space, and select anchor points with a distance value of less than 15 cm to establish the nearest neighbor matching relationship based on the principle of minimizing the distance cost function. Then, attach the coordinates of the virtual identifier to the successfully matched anchor point position.

[0102] S64. Perform a tensor concatenation operation on the attribute structure data of the virtual pending identifier and the high-probability anchor position data that has been successfully matched to generate an incremental update data packet containing complete spatial information. Use the unique identifier in the virtual identifier as the time-series index key value to perform a key-value pair overwrite operation on the global hash table to complete the index update.

[0103] In this embodiment, S7 specifically includes:

[0104] S71. Based on the chaining hash collision resolution strategy, read the updated index key-value pairs and construct a hash table structure containing 1024 buckets. When a hash collision occurs, find the target key-value by traversing the linked list nodes attached to the first address of each bucket. Through key-value inversion parsing, separate the 128-dimensional radio frequency identification feature vector corresponding to the high 32 bits and the 4x4-dimensional visual geometric coordinate matrix corresponding to the low 64 bits from the key-value pair. Calculate the source addressing pointer and the target storage base address of the data stream using the base address plus offset addressing method.

[0105] S72. Perform a linear traversal query in the historical radio frequency-visual dual-modal mapping tensor library according to the target storage base address, and extract historical tensor data slices whose timestamp index is within the preset sliding window interval from 30 seconds before the current time to the current time.

[0106] S73. Flatten the real-time heatmap data into a vector form, perform a dot product operation with the historical tensor data slices and divide by the product of the vector magnitude, calculate the cosine similarity measure, introduce an exponentially decaying time forgetting factor, set the weight of data within 1 second from the current time to 0.9, set the weight of data within 30 seconds to 0.1, calculate the dynamic fusion weight, and generate the corresponding dynamic update mask matrix.

[0107] S74. Perform pixel-level weighting on the real-time heatmap data using a dynamically updated mask matrix, set pixels with weight values ​​less than 0.3 to zero, and write the weighted data into the corresponding storage unit of the historical tensor data slice through tensor assignment operation to update the numerical field of the mapped tensor.

[0108] S75. Calculate the Frobenius norm error of the mapping tensor before and after the update. Obtain the error value by subtracting the matrices, summing the squares, and then taking the square root. Determine whether the error has converged. If the error value is greater than the preset threshold of 0.001, then correct the index parameters based on the updated tensor value and repeat the update step until the convergence condition is met, and output the final RF-vision bimodal mapping tensor.

[0109] refer to Figure 4 A heatmap-based intelligent archive inventory system includes the following modules:

[0110] The dual-modal benchmark construction module is used to acquire panoramic images of standard warehouses and extract spine features to construct benchmark heat maps. Simultaneously, it reads RFID to establish indexes and binds IDs with spatial coordinate hashes to generate radio frequency-visual dual-modal mapping tensors.

[0111] The wide-area scanning and anomaly locking module is used to control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual bimodal mapping tensor, and lock the target area and generate a visual wake-up queue if an anomaly is detected.

[0112] The high-resolution sampling stream generation module is used to respond to visual wake-up queue instructions, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream.

[0113] The manifold constraint denoising generation module is used to improve the MaxViT network based on manifold geometric constraints. It maps the original local image stream to the latent space and injects Gaussian noise. It constructs the manifold constraint boundary by using the file arrangement rules and performs iterative denoising prediction to generate a real-time heat map according to the time step.

[0114] The dual fault discrimination module is used to construct a dual fault discrimination mechanism of differential matrix and logical AND. It aligns the real-time heat map with the reference heat map in time and space and calculates the differential matrix. It extracts differential features and performs logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost.

[0115] The virtual identifier and index incremental update module is used to create a virtual pending identifier if the tag is determined to be invalid, extract the coordinates of high-probability anchor points in the real-time heatmap for attachment, and perform incremental updates on the index.

[0116] The tensor iteration update module is used to perform iterative updates based on the updated index, and to use real-time heatmap data to overwrite historical data to complete the iteration of the radio frequency-visual dual-modal mapping tensor.

[0117] Example 1:

[0118] To verify the feasibility of this invention in intelligent archival inventory and misplacement detection, the method was applied to the intelligent archival inventory system of a large archival management center (hereinafter referred to as "Center A"). Traditional archival inventory systems typically employ RFID-based scanning or manual visual counting. These methods are not only limited by signal obstruction from dense metal shelving, leading to high miss rates, but also cannot effectively handle physical anomalies such as detached tags or misplaced archives, easily resulting in discrepancies between records and actual archives and difficulties in retrieval. To address these problems, Center A decided to adopt the heatmap-based intelligent archival inventory method proposed in this invention.

[0119] During implementation, Center A first uses its deployed panoramic image acquisition equipment to acquire a global image of the warehouse area, extracts the spine features of the archives to construct a baseline heat map, and simultaneously reads RFID tags to create an index and binds the IDs with spatial coordinate hashes, generating a radio frequency-visual bimodal mapping tensor. Subsequently, it operates mobile devices to perform wide-area scanning to acquire a real-time tag set. If data anomalies are detected, the target area is locked and a raw local image stream is generated using high-resolution imaging equipment.

[0120] Center A utilizes multi-scale block and channel attention extraction layer techniques to deeply fuse local ridge detail features and global arrangement features from the original local image stream, efficiently capturing the spatial distribution of archives at the pixel level. On the other hand, through a latent space Gaussian injection layer and a manifold denoising iterative generation layer, combined with a manifold constraint boundary constructed from the archive arrangement rules, affine transformations and reparameterization operations are performed on noisy latent variables in the latent space. During iterative denoising, projection clipping is used to correct feature distribution, effectively restoring a high-fidelity real-time heatmap that conforms to the physical topology of the archives. Subsequently, the generated real-time heatmap and the baseline heatmap are registered using Fourier Merlin transform and difference matrix operations. Based on joint judgment state and information entropy features, accurate classification of label failure (missed reading) and loss of archive authenticity (misplaced / missing) is achieved.

[0121] During implementation, the technical team at Center A discovered that, compared to traditional RFID-based inventory methods, the method of this invention significantly improves the accuracy and robustness of inventory checks. Traditional methods cannot precisely characterize the physical arrangement features of archives, while the method of this invention, through a multi-scale attention mechanism and a manifold-constrained denoising model, effectively achieves visual completion of archives in signal-obstructed areas and precise location of misplaced archives.

[0122] To further verify the actual performance of the method of the present invention, Center A conducted a detailed comparative test between the method of the present invention and the traditional method. The specific performance data is shown in Table 1:

[0123] Table 1 Performance Comparison of Center A's Intelligent Archives Inventory Method

[0124] index Traditional methods Method of the present invention Increase Inventory accuracy rate (%) 88.5 96.2 +8.7% Misidentification rate (%) 65.4 93.8 +43.4% Tag expiration identification accuracy (%) 65.0 94.0 +44.6% Abnormal positioning error (cm) 45.0 8.5 -81.1% Image reconstruction PSNR (dB) 22.4 32.6 +45.5% Time taken for a single inventory count (milliseconds) 150 42 -72.0% Real-time processing frame rate (FPS) 6 24 +300.0% Radio frequency-vision fusion matching success rate (%) 82.0 98.5 +20.1% Data update index query time (milliseconds) 20 3 -85.0% Overall system resource utilization rate (%) 75.0 45.0 -40.0%

[0125] As shown in Table 1, the performance of the intelligent archive inventory system was comprehensively improved after applying the method of this invention. The inventory accuracy increased from 88.5% of the traditional method to 96.2%, and the label failure identification accuracy reached 94.0%, significantly improving the reliability of inventory and effectively solving the problem of missed reporting caused by metal obstruction. The anomaly positioning error decreased from 45.0 cm to 8.5 cm, and the image reconstruction quality (PSNR) improved by 45.5%, ensuring accurate positioning of misplaced archives. The processing time for a single inventory count decreased from 150 milliseconds to 42 milliseconds, and the real-time processing frame rate increased from 6 FPS to 24 FPS, meeting the real-time inventory requirements of large-scale warehouses.

[0126] Through the method of this invention, Center A has successfully achieved intelligent and accurate inventory of the archive storage environment, effectively improving the consistency rate of records and physical records and the retrieval efficiency of archive management, ensuring the storage security of physical archives and digital assets, significantly improving the level of intelligence and automation of archive management, significantly reducing the cost of manual inventory, enhancing the robustness of the system in complex environments, and providing strong technical support for the construction of smart archives.

[0127] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent inventory of archives based on heatmaps, characterized in that, Includes the following steps: S1. Collect panoramic images of the standard warehouse and extract spine features to construct a baseline heat map. Simultaneously read RFID to establish an index and bind the ID with spatial coordinate hash to generate an RFID-visual bimodal mapping tensor. S2. Control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual dual-modal mapping tensor, and if data abnormality is detected, lock the target area and generate a visual wake-up queue. S3. Respond to the visual wake-up queue command, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream; S4. Based on manifold geometric constraints, the MaxViT network is improved by mapping the original local image stream to the latent space and injecting Gaussian noise. The manifold constraint boundary is constructed by utilizing the file arrangement rules, and iterative denoising prediction is performed according to the time step to generate a real-time heat map. S5. Construct a dual fault discrimination mechanism of differential matrix and logical AND, align the real-time heat map with the reference heat map in time and space and calculate the differential matrix, extract differential features and perform logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost. S6. If the label is determined to be invalid, create a virtual pending identifier, extract the coordinates of high-probability anchor points in the real-time heatmap and mount them, and perform incremental updates on the index. S7. Perform iterative updates based on the updated index, use real-time heatmap data to overlay and map historical data, and complete the iteration of the radio frequency-visual dual-modal mapping tensor.

2. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, S1 includes: S11. Collect panoramic image sequences of the standard warehouse, construct a multidimensional spatiotemporal probability density field through spine edge detection and texture feature extraction algorithms, and calibrate the multidimensional spatiotemporal probability density field as a benchmark heat map. S12. Analyze the warehouse radio frequency signal to establish an electronic tag index, determine the pixel-level physical coordinates based on the baseline heat map, and perform a hash operation on the unique ID of the tag and the physical coordinates to generate a spatial hash key-value pair. S13. By fusing spatial hash key-value pairs, electronic tag indexes, and multidimensional spatiotemporal probability density fields, tensor mapping operations are performed to generate a radio frequency-visual bimodal mapping tensor.

3. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, S2 includes: S21. Control the inventory equipment to perform wide-area traversal scanning along the preset operation path, capture radio frequency signals in real time and parse them to generate a real-time tag set; S22. Perform hash mapping verification on the real-time tag set and the radio frequency-vision bimodal mapping tensor, and calculate the tag existence deviation and spatial topology consistency index. S23. Construct an abnormal state discrimination function based on the existence deviation and spatial topology consistency index. If the function value exceeds the preset information interval, it is determined that the radio frequency data is abnormal and the area locking mechanism is triggered. S24. Analyze the index nodes in the abnormal state discrimination function and reverse map them to the global physical coordinate system of the baseline heatmap to accurately locate the abnormal physical coordinate region; S25. Encapsulate the state characteristics of the abnormal physical coordinate region into a structured visual wake-up request, sort it according to the time sequence of the abnormality level, and generate a visual wake-up queue.

4. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, S3 specifically includes: S31. Analyze the structured request parameters in the visual wake-up queue, solve the mechanical kinematics model based on the abnormal physical coordinate region, and generate spatial pose adjustment parameters; S32. Respond to the spatial pose adjustment parameters, control the imaging device to point to the target abnormal area, perform fixed-point high-resolution optical sampling, and acquire multiple frames of high-dimensional image data; S33. Perform temporal registration and streaming reconstruction on multi-frame high-dimensional image data to generate the original local image stream of the target anomaly region.

5. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, The improved MaxViT network includes a multi-scale block attention extraction layer, a multi-scale channel attention extraction layer, a cascaded downsampling module, a latent space Gaussian injection layer, and a manifold denoising iterative generation layer. The multi-scale block attention extraction layer is used to divide and flatten the input original local image stream into image block sequence tensors. The image block sequence tensors are then mapped to query key-value pair tensors through a multilayer perceptron. Local self-attention calculation is performed on the query key-value pair tensors based on a window partitioning strategy, and global extended self-attention calculation is performed on the query key-value pair tensors based on a grid partitioning strategy. Local detail features and global semantic features are fused to output a multi-scale block feature tensor. The multi-scale channel attention extraction layer is used to perform a transpose operation on the multi-scale block feature tensor, dividing it into multi-head sub-feature tensors along the channel dimension; it generates channel query key-value pair tensors through linear transformation, and performs channel self-attention operation to capture the feature dependencies between channels. After aggregation and transpose processing, it outputs the enhanced channel feature tensor. The cascaded downsampling module is used to flatten and linearly project the enhanced channel feature tensor onto a latent space of a preset dimension, and output the latent space feature tensor. The latent space Gaussian injection layer is used to calculate the statistical moments of the latent space feature tensor and generate the mean vector and variance vector. Based on the preset time step, the standard Gaussian noise tensor is sampled, and the mean vector and variance vector are used to perform affine transformation and reparameterization operations on the standard Gaussian noise tensor to obtain the noisy latent variable tensor. The manifold denoising iterative generation layer is used to input the noisy latent variable tensor into the denoising U-Net network and perform iterative noise residual prediction along the reverse time step direction. The arrangement of the archives constructs a manifold constraint boundary. During the iterative denoising process, manifold projection pruning and range constraints are performed on the intermediate latent variable tensors to correct the feature distribution to conform to the physical topology of the archives. Based on the corrected latent variables and the noise residual at the current time step, combined with the preset noise scheduler parameters, the denoised latent variable tensor is recursively calculated. The final denoised latent variable tensor is mapped back to the pixel space to generate a real-time heatmap.

6. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, The difference matrix and logical AND dual fault detection mechanism specifically includes: Based on Fourier Merlin transform, rigid transformation parameters are estimated between real-time heatmaps and reference heatmaps. The real-time heatmaps are then resampled according to the transformation parameters to generate a registered target image sequence. A pixel-wise Euclidean distance metric is performed on the registered target image sequence and the baseline heatmap, and the difference metric results are mapped to the Gaussian probability space to generate a confidence distribution matrix. The proportion of low-confidence regions below a preset threshold in the statistical confidence distribution matrix is ​​converted into the first binary discriminant variable, and the radio frequency anomaly detection result is mapped into the second binary discriminant variable. The first binary discriminant variable and the second binary discriminant variable are processed by an AND gate, and the joint judgment state is output. Based on the information entropy feature of the confidence distribution matrix, the classification label of label failure or true loss is output.

7. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, S6 specifically includes: S61. Initialize the virtual pending identifier based on the fault determination signal, and define an attribute structure containing a unique identifier, a timestamp, and a status bit vector; S62. Perform morphological gradient operation on the real-time heatmap to extract local extreme value regions, apply non-maximum suppression algorithm to filter neighborhood redundant responses, use the mapping matrix from image coordinate system to global coordinate system to perform spatial transformation on extreme point coordinates, and output a set of high-probability anchor point positions. S63. Calculate the Euclidean distance metric between the predicted location of the virtual undetermined identifier and the set of high-probability anchor points. Based on the principle of minimizing the distance cost function, establish the nearest neighbor matching relationship and perform the coordinate mounting operation of the virtual identifier. S64. Concatenate the attribute structure of the virtual pending identifier with the high-probability anchor position using a tensor to generate an incremental update data packet. Perform an update operation on the global hash table based on the time-series index key value to complete the index update.

8. The method for intelligent inventory of archives based on heatmaps according to claim 1, characterized in that, Specifically, S7 includes: S71. Based on the hash collision resolution strategy, read the updated index key-value pairs, and parse the corresponding radio frequency identity feature vector and visual geometric coordinate matrix through key-value inversion to calculate the source addressing pointer and target storage base address of the data stream. S72. Perform a traversal query in the historical radio frequency-visual dual-modal mapping tensor library according to the target storage base address, and extract historical tensor data slices whose timestamp index is within the preset sliding window interval; S73. Calculate the cosine similarity measure between real-time heatmap data and historical tensor data slices, introduce an exponentially decaying time forgetting factor to calculate dynamic fusion weights, and generate the corresponding dynamically updated mask matrix. S74. Perform pixel-level weighting on the real-time heatmap data using a dynamically updated mask matrix, and write the weighted data into the corresponding storage unit of the historical tensor data slice through tensor assignment operation to update the numerical field of the mapped tensor. S75. Calculate the Frobenius norm error of the mapping tensor before and after the update, and determine whether the error has converged. If the error is greater than the preset threshold, correct the index parameters based on the updated tensor value and repeat the update steps until the convergence condition is met, and output the final RF-vision bimodal mapping tensor.

9. A heatmap-based intelligent archive inventory system, comprising executing the heatmap-based intelligent archive inventory method according to any one of claims 1 to 8, characterized in that, Includes the following modules: The dual-modal benchmark construction module is used to acquire panoramic images of standard warehouses and extract spine features to construct benchmark heat maps. Simultaneously, it reads RFID to establish indexes and binds IDs with spatial coordinate hashes to generate radio frequency-visual dual-modal mapping tensors. The wide-area scanning and anomaly locking module is used to control the mobile device to perform wide-area scanning to obtain a real-time tag set, perform hash comparison with the radio frequency-visual bimodal mapping tensor, and lock the target area and generate a visual wake-up queue if an anomaly is detected. The high-resolution sampling stream generation module is used to respond to visual wake-up queue instructions, adjust the mechanical pose to perform fixed-point high-resolution sampling on the abnormal target area, and generate the original local image stream. The manifold constraint denoising generation module is used to improve the MaxViT network based on manifold geometric constraints. It maps the original local image stream to the latent space and injects Gaussian noise. It constructs the manifold constraint boundary by using the file arrangement rules and performs iterative denoising prediction to generate a real-time heat map according to the time step. The dual fault discrimination module is used to construct a dual fault discrimination mechanism of differential matrix and logical AND. It aligns the real-time heat map with the reference heat map in time and space and calculates the differential matrix. It extracts differential features and performs logical AND operation with the radio frequency anomaly results to determine whether the tag is invalid or truly lost. The virtual identifier and index incremental update module is used to create a virtual pending identifier if the tag is determined to be invalid, extract the coordinates of high-probability anchor points in the real-time heatmap for attachment, and perform incremental updates on the index. The tensor iterative update module is used to perform iterative updates based on the updated index, and to use real-time heatmap data to overwrite historical data to complete the iteration of the radio frequency-visual dual-modal mapping tensor.