An asset identification method and device based on a multi-mode fusion visual map
By generating multi-modal fusion visualization maps and combining panoramic images and inertial data with radio frequency signals, the spatial intuitiveness and efficiency problems of asset inventory in existing technologies are solved. This enables the visualization and efficient identification of assets in images, and is suitable for refined asset environments such as office spaces, computer rooms, and hospitals.
Patent Information
- Application Number
- CN202511339757.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing indoor asset inventory methods lack spatial intuitiveness, are disconnected from images and identities, are inefficient, and have weak spatial distribution perception. They cannot accurately present asset locations on spatial maps and bind asset identification information, and require a lot of manual intervention.
By generating a multi-modal fusion visualization map, using panoramic images and inertial data combined with radio frequency signals, the location and type of assets in the panoramic images are marked. The orientation and signal strength of the assets are determined by using a panoramic camera, inertial measurement unit and radio frequency receiver, so as to realize the visualization and spatial navigation of assets in the images.
It enables assets to be visible, selectable, and annotated in images, improving the efficiency and accuracy of asset recognition. It supports lightweight indoor deployment, is suitable for dense asset scenarios, and provides human-computer interactive correction and multi-image fusion visualization map management.
Smart Images

Figure CN120823050B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of indoor asset digital acquisition and spatial perception technology, specifically an asset identification method, device, equipment, medium, and product program based on a multi-modal fusion visualization map. Background Technology
[0002] In existing technologies, indoor asset inventory and management methods generally rely on manual barcode scanning, handheld terminals, text ledgers, or separate image management systems. These methods have the following drawbacks:
[0003] 1) Lack of spatial intuitiveness: The location of assets cannot be accurately represented on spatial maps;
[0004] 2) Image and identity disconnect: Images are only used for archiving and are difficult to link to asset identification information;
[0005] 3) Inefficiency: It requires scanning each asset one by one, which involves a lot of manual intervention;
[0006] 4) Weak spatial distribution perception: It cannot perform visual inventory analysis based on asset type, location relationship, etc. Summary of the Invention
[0007] The purpose of this application is to provide at least one method and device for asset identification based on multi-modal fusion visualization maps, in order to solve at least a part of the above-mentioned technical problems.
[0008] To address the aforementioned technical problems, at least one embodiment of this application provides an asset identification method based on a multi-modal fusion visualization map, comprising:
[0009] A panoramic image of the space is generated based on multiple images of the space where the asset is located and the inertial data when the image data was acquired; wherein, the multiple images are multiple frames of image data that are consecutive in time.
[0010] The location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength are determined by the radio frequency receiver; wherein the asset is equipped with an radio frequency tag;
[0011] The location of the corresponding asset is marked in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0012] In some embodiments, marking the location of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity includes:
[0013] The asset type is labeled according to the radio frequency signal and the pre-generated mapping relationship between the radio frequency signal and the asset;
[0014] The type of the asset is labeled in the panoramic image based on the acquisition path, the orientation, and the strength of the radio frequency signal.
[0015] In some embodiments, labeling the type of the asset in the panoramic image based on the acquisition path, the orientation, and the strength of the radio frequency signal includes:
[0016] The initial location of the asset type is determined based on the acquisition path, the orientation, and the strength of the radio frequency signal;
[0017] The initial position is corrected based on the marked position on the acquisition path and the strength of the corresponding radio frequency signal to determine the final position of the asset type;
[0018] The type of the asset is labeled in the panoramic image based on the final location.
[0019] In some embodiments, the step of determining the acquisition path includes:
[0020] Generate an image frame sequence based on the plurality of images;
[0021] The acquisition path is determined based on the image frame sequence and the inertial data.
[0022] In some embodiments, generating a panoramic image of the space based on multiple images of the space where the asset is located and inertial data from when the image data was acquired includes:
[0023] Extract feature points from frame image data;
[0024] Multiple feature points are matched according to the time order in the corresponding image frame sequence to generate feature point matching results;
[0025] The relative displacement and rotation angle between consecutive frames in the image frame sequence are determined based on the inertial data.
[0026] The panoramic image is generated by stitching together multiple images based on the feature point matching results, the relative displacement, and the projection of the rotation angle.
[0027] In some embodiments, after marking the location of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity, the method further includes:
[0028] Spatial navigation of the asset is performed based on the panoramic image after the asset is marked.
[0029] At least one embodiment of this application also provides an asset identification device based on a multi-modal fusion visualization map, comprising:
[0030] A panoramic image generation module is used to generate a panoramic image of the space based on multiple images of the space where the asset is located and the inertial data when the image data was acquired; wherein, the multiple images are multiple frames of image data that are consecutive in time;
[0031] The radio frequency signal determination module is used to determine the location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength through the radio frequency receiver; wherein the asset is equipped with a radio frequency tag;
[0032] The asset location marking module is used to mark the location of the corresponding asset in the panoramic image according to the pre-determined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0033] In some embodiments, the asset location marking module includes:
[0034] An asset type labeling unit is used to label the type of an asset based on the radio frequency signal and the pre-generated mapping relationship between the radio frequency signal and the asset.
[0035] An asset location marking unit is used to mark the type of the asset in the panoramic image based on the acquisition path, the orientation, and the strength of the radio frequency signal.
[0036] In some embodiments, the asset location marking unit includes:
[0037] An initial location determination unit is used to determine the initial location of the asset type based on the acquisition path, the orientation, and the strength of the radio frequency signal.
[0038] The final location determination unit is used to correct the initial location based on the marked location on the acquisition path and the strength of the radio frequency signal corresponding to that location, so as to determine the final location of the asset type;
[0039] An asset location marking subunit is used to mark the type of the asset in the panoramic image based on the final location.
[0040] In some embodiments, an asset identification device based on a multi-modal fusion visualization map further includes:
[0041] Acquisition path determination module, used to determine the acquisition path; the acquisition path determination module includes:
[0042] An image frame sequence generation unit is used to generate an image frame sequence based on the plurality of images;
[0043] The acquisition path determination unit is used to determine the acquisition path based on the image frame sequence and the inertial data.
[0044] In some embodiments, the panoramic image generation module includes:
[0045] The feature point extraction unit is used to extract feature points from frame image data;
[0046] The feature point matching result generation unit is used to match multiple feature points according to the time order in the image frame sequence to generate feature point matching results;
[0047] A relative displacement determination unit is used to determine the relative displacement and rotation angle between consecutive frames in the image frame sequence based on the inertial data.
[0048] A panoramic image generation unit is used to stitch together multiple images after projection based on the feature point matching results, the relative displacement, and the rotation angle to generate the panoramic image.
[0049] In some embodiments, after marking the location of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity, an asset identification device based on a multi-modal fusion visualization map further includes:
[0050] The asset spatial navigation module is used to perform spatial navigation on the assets based on the panoramic image after the assets are marked.
[0051] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described asset identification method based on a multi-modal fusion visualization map.
[0052] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described asset identification method based on a multi-modal fusion visualization map.
[0053] The embodiments of this application provide an asset identification method and apparatus based on a multi-modal fusion visualization map. First, a panoramic image of the space is generated based on multiple images of the space where the asset is located and inertial data from the acquisition of the image data; wherein, the multiple images are multiple frames of image data that are consecutive in time; next, the orientation of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength are determined by a radio frequency receiver; wherein, the asset is equipped with a radio frequency tag; finally, the position of the corresponding asset is marked in the panoramic image according to the pre-determined acquisition path, orientation, corresponding radio frequency signal and its strength.
[0054] The method provided in this application can quickly acquire asset images and their IDs. In addition, the method can also directly annotate assets in spatial images, allowing users to interactively select or update asset information in the images. Attached Figure Description
[0055] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0056] Figure 1 This is a schematic flowchart of an asset identification method based on a multi-modal fusion visualization map provided in one embodiment of this application;
[0057] Figure 2 This is a flowchart illustrating step 300 provided in one embodiment of this application;
[0058] Figure 3 This is a flowchart illustrating step 302 provided in one embodiment of this application;
[0059] Figure 4 This is a flowchart illustrating step 100 provided in one embodiment of this application;
[0060] Figure 5 This is another flowchart illustrating an asset identification method based on a multi-modal fusion visualization map, provided in one embodiment of this application.
[0061] Figure 6 This is a flowchart illustrating an asset identification method based on a multi-modal fusion visualization map, provided in a specific embodiment of this application.
[0062] Figure 7 This is a block diagram of an asset identification device based on a multi-modal fusion visualization map provided in one embodiment of this application;
[0063] Figure 8 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0066] It should be noted that the terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. Without conflict, the embodiments and features in the embodiments of this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0067] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0068] Provide users with corresponding operation entry points, allowing them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0069] The acquisition, storage, use, and processing of data in this application all comply with relevant laws and regulations. Specifically:
[0070] First, the information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0071] Second, provide users with corresponding operation entry points for them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0072] Example 1:
[0073] This embodiment presents an asset identification method based on a multi-modal fusion visualization map, which can be applied to electronic devices with communication, computing, and data storage capabilities. The specific process can be as follows: Figure 1 As shown, it includes:
[0074] Step 100: Generate a panoramic image of the space based on multiple images of the space where the asset is located and the inertial data when the image data was acquired; wherein, the multiple images are multiple frames of image data that are consecutive in time;
[0075] Step 200: Determine the location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength; wherein, the asset is equipped with an RFID tag;
[0076] Step 300: Mark the location of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0077] An embodiment of this application provides an asset identification method based on a multi-modal fusion visualization map. First, a panoramic image of the space is generated based on multiple images of the space where the asset is located and the inertial data during image data acquisition. The multiple images are multi-frame image data that are consecutive in time. Next, the orientation of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength are determined by a radio frequency receiver. The asset is equipped with a radio frequency tag. Finally, the position of the corresponding asset is marked in the panoramic image according to the pre-determined acquisition path, orientation, corresponding radio frequency signal and its strength.
[0078] The method provided in this application can quickly acquire asset images and their IDs. In addition, the method can also directly annotate assets in spatial images. Finally, users can interactively select or update asset information in the images.
[0079] Example 2:
[0080] In some examples, multiple images of the space where the asset is located in step 100 can be acquired by a high-definition panoramic camera, as well as inertial data when the image data is acquired by an inertial measurement unit.
[0081] Preferably, the aforementioned high-definition panoramic camera incorporates an inertial measurement unit (IMU). An IMU is an electronic device that integrates multi-axis sensors to measure an object's angular velocity, linear acceleration, and attitude changes. It does not rely on external signals (such as GPS) and achieves autonomous motion sensing through its own sensors.
[0082] In some examples, the radio frequency identification (RFID) receiver in step 200 is used to receive and decode radio frequency signals from electronic tags on the asset, converting the information stored in the tag into processable data. It is typically integrated with an RFID reader (called a reader module), but can also be deployed independently for specific scenarios (such as signal monitoring).
[0083] Specifically, the system captures the weak radio frequency signals backscattered by the tag (for passive tags) or the signals actively transmitted (for active tags) through an antenna. Then, the modulated signals on the radio frequency carrier are separated (such as ASK and PSK) to extract the tag ID and stored data. Finally, the data is transmitted to a host computer or cloud system through an interface (such as RS232, Ethernet, or WiFi).
[0084] In some examples, specifically for step 300, the tags sensed by the RFID signals are "position probability delineated and annotated" on the corresponding panoramic frame.
[0085] In some examples, see Figure 2 Step 300 includes:
[0086] Step 301: Label the asset type according to the radio frequency signal and the pre-generated mapping relationship between the radio frequency signal and the asset;
[0087] Step 302: Label the type of the asset in the panoramic image according to the acquisition path, the orientation, and the strength of the radio frequency signal.
[0088] In some examples, see Figure 3 Step 302 includes:
[0089] Step 3021: Determine the initial location of the asset type based on the acquisition path, the orientation, and the strength of the radio frequency signal;
[0090] Specifically, the input data is first determined as follows: the 3D coordinate sequence of the reader's movement trajectory (such as the AGV's travel path), the real-time pointing angle of the directional antenna (azimuth angle φ, elevation angle θ), and the tag RSSI (signal strength) and phase difference received at each path point; then, the RSSI attenuation values of the same tag at multiple path points are substituted into the path loss model (preferably the Log-distance model), and the distance probability distribution is calculated in combination with the antenna directional gain; beam cones are constructed using the antenna pointing angles at different positions, and the initial position area is determined by the intersection of the cones; the radial velocity is calculated by the phase change rate of adjacent points, static reflection interference is eliminated, and finally, the initial 3D coordinates (x0, y0, z0) and confidence score of the asset are output.
[0091] Step 3022: Correct the initial position based on the marked position on the acquisition path and the strength of the radio frequency signal corresponding to that position, so as to determine the final position of the asset type;
[0092] First, deploy reference tags with known coordinates (such as ground positioning QR codes + RFID composite tags) at key locations along the data acquisition path; then, perform signal strength comparisons.
[0093] Measurement: Measured RSSI of the reference tag scanned by the reader;
[0094] Theoretical value: The expected RSSI at this location (calculated based on the environmental attenuation model).
[0095] Next, dynamic compensation is performed: if the measured RSSI is 10dB lower than the theoretical value, it is determined that there is metal occlusion in the area, and a position offset vector is generated. Finally, correction is performed: the initial position (x0, y0, z0) is superimposed with the offset vector: the final position P_final = the initial position P_initial + ΔP; the final coordinates (x, y, z) and error range (e.g., ±15cm) are output.
[0096] Preferably, for an asset's RFID (corresponding to a physical location), it can be collected from multiple collection points. Multiple collection points can obtain multiple location values, and through relevant algorithms, the calculated location values can be made closer to the actual location.
[0097] Step 3023: Label the type of the asset in the panoramic image according to the final location.
[0098] Specifically, panoramic image annotation (spatial mapping and visualization) is performed by first transforming the coordinates:
[0099] Panoramic camera coordinate system alignment: Transform the final position (x, y, z) to the spherical coordinate system of the panoramic camera (radius r, horizontal angle α, pitch angle β):
[0100] ;
[0101] ;
[0102] Next, image pixel mapping: horizontal angle α → horizontal coordinate u of panoramic image (u = image width × α / 360°); pitch angle β → vertical coordinate v of panoramic image (v = image height × (0.5 + β / 180°)).
[0103] Finally, dynamic annotation is performed to draw an asset type icon (such as a medical equipment icon) at the coordinates (u,v); and text labels are overlaid: [Asset Type]@[Error Range] (such as "Ultrasound Instrument @±15cm").
[0104] In some examples, the steps of determining the acquisition path include:
[0105] Generate an image frame sequence based on the plurality of images; and determine the acquisition path based on the image frame sequence and the inertial data.
[0106] Specifically, the generated continuous image frame sequence undergoes distortion correction and other preprocessing, such as denoising and equalization. Acceleration and angular velocity data are acquired from the inertial measurement unit (IMU). The IMU data is calibrated to remove bias and noise.
[0107] Extract feature points (such as ORB, SIFT, SURF) from each frame of the image. Perform feature matching between adjacent image frames to find common feature point pairs.
[0108] Using the feature matching results, the relative pose (translation and rotation) between adjacent frames is estimated using geometric methods (such as the five-point method or the eight-point method). Preferably, the RANSAC algorithm can be used during the feature matching process to eliminate incorrect matches and improve the robustness of the estimation.
[0109] Using IMU acceleration and angular velocity data, the device's pose change is calculated through numerical integration. Double integration is required to derive the position change from acceleration, while also considering the effect of angular velocity. During this process, it is essential to ensure time alignment between the image frames and the IMU data. Interpolation may be performed if necessary to acquire both image and IMU data at the same timestamp.
[0110] The results from visual odometry and IMU can be fused together using fusion algorithms such as Extended Kalman Filter (EKF) or Sliding Window Optimization.
[0111] Specifically, pose prediction is performed based on IMU data, and the predicted state is corrected based on visual odometry results. Pose prediction is optimized simultaneously at multiple time points within a fixed-size window. All poses within the window are re-estimated by minimizing the errors between visual and IMU data.
[0112] Finally, the fused pose results are connected in chronological order to form a complete motion path.
[0113] In some examples, see Figure 4 Step 100 includes:
[0114] Step 101: Extract feature points from the frame image data;
[0115] Unique texture information is captured from each frame of the image as stitching anchors. Feature detection algorithms are used to identify key points in the image (such as corners and edge intersections). A feature descriptor (mathematical vector) is generated for each key point to describe the gradient distribution pattern of the surrounding pixels. Finally, hundreds to thousands of feature points and their descriptors are obtained in each frame of the image.
[0116] Step 102: Match multiple feature points according to the time order in the corresponding image frame sequence to generate feature point matching results;
[0117] Specifically, a correspondence between features of adjacent frames is established, and matching is performed only on frames with consecutive timestamps (e.g., frame 1 ↔ frame 2, frame 2 ↔ frame 3). Bidirectional matching verification is performed: forward: finding the nearest neighbor match for feature points in frame 1 in frame 2; backward: verifying consistency by returning the matching point from frame 2 to frame 1. The fundamental matrix is then calculated using the RANSAC algorithm, and matching point pairs that do not satisfy epipolar geometry are eliminated. Finally, a high-confidence cross-frame feature point correspondence is output (e.g., point A in frame 1 ↔ point A' in frame 2).
[0118] Step 103: Determine the relative displacement and rotation angle between consecutive frames in the image frame sequence based on the inertial data;
[0119] Specifically, the camera pose change between consecutive frames is calculated by inertial data. The relative rotation angle (Δθ) is obtained by integrating the gyroscope angular velocity within the time interval between two adjacent frames. The relative displacement (Δp) is obtained by double integration of the accelerometer data. Based on this technique, the IMU data is converted to the camera coordinate system using calibration parameters.
[0120] Step 104: Stitch together the multiple images after projection based on the feature point matching results, the relative displacement, and the rotation angle to generate the panoramic image.
[0121] An adaptive hybrid stitching method is employed here: in overlapping areas, pixels are dynamically weighted and fused based on feature point matching accuracy; in high-matching areas, pixels from the current frame are prioritized (to avoid ghosting); in low-matching areas, IMU motion prediction results are superimposed (to ensure continuity). The final output is a seamless 360° panoramic image (each frame is precisely aligned in spherical space).
[0122] In some examples, see Figure 5 An asset identification method based on a multi-modal fusion visualization map, after step 300, further includes:
[0123] Step 400: Perform spatial navigation on the asset based on the panoramic image after the asset is marked.
[0124] Specifically, a visual map building system is established: multi-frame panoramic images are stitched together along paths to build an interactive map that supports ID query, spatial navigation, and historical comparison.
[0125] This application provides an asset identification method based on a multi-modal fusion visualization map. First, a panoramic image of the space is generated based on multiple images of the asset's location and inertial data from the image acquisition process. The multiple images are sequential, multi-frame image data. Next, the asset's orientation relative to the RF receiver, the corresponding RF signal, and its strength are determined using an RF receiver. The asset is equipped with an RF tag. Finally, the location of the corresponding asset is marked in the panoramic image based on the predetermined acquisition path, orientation, corresponding RF signal, and its strength. Compared with existing technologies, this application has the following advantages:
[0126] 1) Visible, selectable, and annotable in the image: Unlike traditional 3D map point cloud annotation, the system directly completes asset positioning and identification in high-fidelity panoramic images, enabling image-level asset auditing and recognition.
[0127] 2) Lightweight indoor deployment: Only requires a panoramic camera + RFID module + IMU, without the need for depth cameras, LiDAR, or robot platforms.
[0128] 3) Regional RFID positioning strategy: Instead of precise coordinate RFID positioning, it generates "possible areas" on the image to assist in identification by using radio frequency coverage direction and intensity.
[0129] 4) Supports interactive correction: Provides an image click and drag interface to correct asset positions, improving final accuracy, suitable for dense asset scenarios in space.
[0130] 5) Multi-image fusion visual map: An interactive asset map is built using images as units, allowing users to manage spatial assets like viewing a "street view map".
[0131] Example 3:
[0132] To further illustrate the solution, this application also provides a specific implementation of an asset identification method based on a multi-modal fusion visualization map, which includes the following:
[0133] The asset identification method based on multi-modal fusion visualization map provided in this application is an asset visualization identification and management method that combines panoramic images, inertial navigation, RFID area perception and semantic map interaction technology. It is applicable to refined asset environments such as office spaces, computer rooms, hospitals, and laboratories.
[0134] First, this application provides an asset identification system based on a multi-modal fusion visualization map, the system comprising:
[0135] 1) Panoramic acquisition equipment: including a high-definition panoramic camera and an inertial measurement unit, used to acquire panoramic images and spatial trajectories;
[0136] 2) Area-aware RFID receiver: It uses a directional antenna or array structure to "define" the relative location of the asset and the receiving range;
[0137] First, RFID tags need to be set on the asset; then, a directional antenna is used to collect RFID tags in that direction; depending on the antenna's transmission power, RFID tags within a certain distance range in that direction can be collected; the antenna's position can be obtained based on "image frame sequence and IMU fusion technology", and finally the area range of the asset is delineated.
[0138] 3) Motion trajectory estimation module: Based on the fusion of image frame sequences and IMU, the acquisition path is calculated without the need to build a 3D map;
[0139] 4) Image spatial projection module: "Position probability delineation and annotation" of the tags sensed by RFID signals on the corresponding panoramic frame;
[0140] 5) Asset Information Binding Platform: Supports image click interaction for binding / correcting asset IDs, enabling image-level asset maintenance;
[0141] 6) Visual map building system: stitches together multiple panoramic images along a path to build an interactive map, supporting ID query, spatial navigation, and historical comparison.
[0142] Taking an office inspection scenario as an example: the equipment used is a binocular panoramic camera with an IMU and a handheld RFID antenna array; the data collector walks along the area inspection route and collects data; the RFID reading range is 1.5~3 meters, and a "perception sector" is generated by combining the receiving angle and the IMU direction; the image shows tagged circles and ID numbers, and the system supports clicking to correct erroneous labels; after the panoramic image is synthesized into a map, it supports classification and filtering such as "printer, server, router", and asset location.
[0143] See Figure 6 Based on the above-mentioned asset identification system based on multi-modal fusion visualization map, a specific implementation method for asset identification based on multi-modal fusion visualization map includes the following steps:
[0144] S1: The operator moves the handheld device along a designated route indoors, and the system collects data synchronously.
[0145] Panoramic image sequences and MU record the attitude angle and displacement of each frame; the RFID module records the timestamp of the received signal, RSSI (Received Signal Strength Indication, used to estimate the distance to the asset), and antenna direction.
[0146] S2: Based on the IMU and inter-frame pose changes, calculate the relative position of each frame image in the indoor space.
[0147] S3: The system uses an RFID receiver to "locate" the signal from the tag and draws a probability area (such as a sector, circle, or cone) in the image to represent the possible location of the asset.
[0148] Specifically, assuming a tag A is read in frame t, its spatial orientation angle is θ, and its RSSI is s, the sensing distance is estimated as follows:
[0149]
[0150] The spherical projection angle φ of the label in the image is: φ=θ+γ; where γ is the camera yaw angle of the current frame.
[0151] The perceptual probability region in step S3 is labeled within a single frame of panoramic image, specifically:
[0152] 1. The annotations occur within a single frame of the panoramic image.
[0153] When a handheld device receives an RFID signal for an asset at location 𝑃𝑡 (attitude + position recorded by the IMU), the system will mark the "possible area" corresponding to this signal in the panoramic image of the current frame.
[0154] The region shape (sector, circle, cone) is drawn on the image projection of the current frame based on the RFID antenna directivity and the distance estimated by RSSI.
[0155] For example, if the antenna is unidirectional, the label "sector area" indicates that the RFID may be located within a certain angle range of the current camera's field of view; if the RSSI is strong, the area radius is small, and vice versa.
[0156] 2. Multi-frame fusion annotation (optional).
[0157] Although annotation occurs first in a single frame, multiple perceptual data points for the same label can be merged:
[0158] If a tag is read in multiple adjacent frames, such as 𝐼𝑡1, 𝐼𝑡2, ..., the system can narrow down the possible location range by the intersection of the perception areas of different frames; ultimately, the tag can be highlighted in the "clearest frame" or the "fused spatial view".
[0159] S4: Bind the RFID tag ID to the image area and automatically load asset information onto the image interaction layer. Draw a sector or circular area with radius D in the φ direction as the "possible tag area".
[0160] S5: Perform "point selection + correction + input" on the asset locations in the image to improve the accuracy of annotation and synchronize it to the asset database.
[0161] Specifically, here is an interactive method for step S5: the system automatically generates an RFID sensing area (semi-transparent fan-shaped / circular) on the panoramic image frame. The user performs the following operations through a simple panoramic icon annotation interface (web or mobile):
[0162] Click: Click on the automatically marked area, and the system will pop up an asset information card;
[0163] Drag and drop correction: Users can drag the area to the actual location of the asset;
[0164] Enter information: You can directly bind existing asset information or add new asset records (such as name and number).
[0165] After user confirmation, the system writes the annotation results (image coordinates + asset ID + RFID tag EPC) into the database. This database must include at least:
[0166] Asset Master: Records basic asset information and RFID EPC; Pano Frame: Records frame ID, time, and pose; Annotation: Records image coordinates and asset binding.
[0167] The above-mentioned asset labeling process includes:
[0168] Load Frame: Displays the current panoramic image and automatic annotation results.
[0169] User correction: Click and adjust the position or size of the annotation.
[0170] Asset entry: Select existing assets or add new asset information.
[0171] Submit and save: Store the results in association with the asset ledger and RFID reading events.
[0172] It is understandable that the following beneficial effects can be achieved through step S5:
[0173] Quick confirmation: Users can simply drag the label to avoid scanning each code one by one.
[0174] Image visualization: The location of assets can be viewed intuitively on a panoramic view.
[0175] The system combines automation and manual intervention: the system first marks the data, and then the human staff makes fine adjustments, resulting in higher accuracy.
[0176] S6: Construct a spatial image map.
[0177] It should be noted that the spatial image map in step S6 is not a 3D point cloud map, but it supports rapid asset image retrieval, spatial navigation, and inspection backtracking.
[0178] Step S6 is implemented by sorting and associating the acquired multi-frame panoramic images (including asset annotations) according to the handheld device pose information (position + orientation) to form a "browsable panoramic image path" rather than a complex 3D point cloud map.
[0179] Each panoramic image frame records the pose (position coordinates + yaw angle) as an "image node"; the frames are connected into a topology map according to the acquisition order or spatial distance; asset labels are bound to the corresponding frames to form an index relationship of "asset → image frame → coordinates".
[0180] Users can quickly locate the best panoramic frame containing the asset using the asset ID; the topology map can be used for "frame skipping along the path" to achieve indoor space navigation; it supports image comparison at different acquisition times to achieve inspection and backtracking.
[0181] Understandably, the spatial image map in step S6 does not require global pixel-level stitching. It only uses "pose + panoramic frame" to build a lightweight spatial map, and the construction process is fast. The map can be updated by adding new frames incrementally.
[0182] This application provides a specific application example of an asset identification method based on a multi-modal fusion visualization map. First, a panoramic image of the space is generated based on multiple images of the asset's location and inertial data from the image acquisition process. These multiple images are sequential, multi-frame image data. Next, the asset's orientation relative to the RF receiver, the corresponding RF signal, and its strength are determined using an RF receiver. The asset is equipped with an RF tag. Finally, the location of the corresponding asset is marked in the panoramic image based on the pre-determined acquisition path, orientation, corresponding RF signal, and its strength. Compared with existing technologies, the method provided in this application has the following advantages:
[0183] 1) Visible, selectable, and annotable in the image: Unlike traditional 3D map point cloud annotation, the system directly completes asset positioning and identification in high-fidelity panoramic images, enabling image-level asset auditing and recognition.
[0184] 2) Lightweight indoor deployment: Only requires a panoramic camera + RFID module + IMU, without the need for depth cameras, LiDAR, or robot platforms.
[0185] 3) Regional RFID positioning strategy: Instead of precise coordinate RFID positioning, it generates "possible areas" on the image to assist in identification by using radio frequency coverage direction and intensity.
[0186] 4) Supports interactive correction: Provides an image click and drag interface to correct asset positions, improving final accuracy, suitable for dense asset scenarios in space.
[0187] 5) Multi-image fusion visual map: An interactive asset map is built using images as units, allowing users to manage spatial assets like viewing a "street view map".
[0188] Example 4:
[0189] Another embodiment of this application relates to an asset identification device based on a multi-modal fusion visualization map. The implementation details of this asset identification device based on a multi-modal fusion visualization map are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution. A schematic diagram of this asset identification device based on a multi-modal fusion visualization map can be seen as follows: Figure 7 As shown, the device includes: a panoramic image generation module 801, a radio frequency signal determination module 802, and an asset location marking module 803.
[0190] The panoramic image generation module 801 is used to generate a panoramic image of the space based on multiple images of the space where the asset is located and the inertial data when the image data was collected; wherein, the multiple images are multiple frames of image data that are consecutive in time.
[0191] The radio frequency signal determination module 802 is used to determine the location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength through the radio frequency receiver; wherein the asset is equipped with a radio frequency tag;
[0192] The asset location marking module 803 is used to mark the location of the corresponding asset in the panoramic image according to the pre-determined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0193] In some embodiments, the asset location marking module includes:
[0194] An asset type labeling unit is used to label the type of an asset based on the radio frequency signal and the pre-generated mapping relationship between the radio frequency signal and the asset.
[0195] An asset location marking unit is used to mark the type of the asset in the panoramic image based on the acquisition path, the orientation, and the strength of the radio frequency signal.
[0196] In some embodiments, the asset location marking unit includes:
[0197] An initial location determination unit is used to determine the initial location of the asset type based on the acquisition path, the orientation, and the strength of the radio frequency signal.
[0198] The final location determination unit is used to correct the initial location based on the marked location on the acquisition path and the strength of the radio frequency signal corresponding to that location, so as to determine the final location of the asset type;
[0199] An asset location marking subunit is used to mark the type of the asset in the panoramic image based on the final location.
[0200] In some embodiments, an asset identification device based on a multi-modal fusion visualization map further includes:
[0201] Acquisition path determination module, used to determine the acquisition path; the acquisition path determination module includes:
[0202] An image frame sequence generation unit is used to generate an image frame sequence based on the plurality of images;
[0203] The acquisition path determination unit is used to determine the acquisition path based on the image frame sequence and the inertial data.
[0204] In some embodiments, the panoramic image generation module includes:
[0205] The feature point extraction unit is used to extract feature points from frame image data;
[0206] The feature point matching result generation unit is used to match multiple feature points according to the time order in the image frame sequence to generate feature point matching results;
[0207] A relative displacement determination unit is used to determine the relative displacement and rotation angle between consecutive frames in the image frame sequence based on the inertial data.
[0208] A panoramic image generation unit is used to stitch together multiple images after projection based on the feature point matching results, the relative displacement, and the rotation angle to generate the panoramic image.
[0209] In some embodiments, after marking the location of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity, an asset identification device based on a multi-modal fusion visualization map further includes:
[0210] The asset spatial navigation module is used to perform spatial navigation on the assets based on the panoramic image after the assets are marked.
[0211] A specific application example of this application provides an asset identification device based on a multi-modal fusion visualization map, comprising: a panoramic image generation module, used to generate a panoramic image of the space based on multiple images of the space where the asset is located and inertial data during the acquisition of the image data; wherein the multiple images are multiple frames of image data that are consecutive in time; a radio frequency signal determination module, used to determine the orientation of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its intensity through a radio frequency receiver; wherein the asset is equipped with a radio frequency tag; and an asset location marking module, used to mark the position of the corresponding asset in the panoramic image according to a pre-determined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0212] Compared with the prior art, the device provided in this application has the following beneficial effects:
[0213] 1) Visible, selectable, and annotable in the image: Unlike traditional 3D map point cloud annotation, the system directly completes asset positioning and identification in high-fidelity panoramic images, enabling image-level asset auditing and recognition.
[0214] 2) Lightweight indoor deployment: Only requires a panoramic camera + RFID module + IMU, without the need for depth cameras, LiDAR, or robot platforms.
[0215] 3) Regional RFID positioning strategy: Instead of precise coordinate RFID positioning, it generates "possible areas" on the image to assist in identification by using radio frequency coverage direction and intensity.
[0216] 4) Supports interactive correction: Provides an image click and drag interface to correct asset positions, improving final accuracy, suitable for dense asset scenarios in space.
[0217] 5) Multi-image fusion visual map: An interactive asset map is built using images as units, allowing users to manage spatial assets like viewing a "street view map".
[0218] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.
[0219] Example 5:
[0220] Another embodiment of this application relates to an electronic device, such as... Figure 8 As shown, the electronic device specifically includes the following:
[0221] Processor 1201, memory 1202, communications interface 1203, and bus 1204;
[0222] The processor 1201, memory 1202, and communication interface 1203 communicate with each other via bus 1204; the communication interface 1203 is used to realize information transmission between server-side devices and user-side devices and other related devices.
[0223] The processor 1201 is used to call the computer program in the memory 1202. When the processor executes the computer program, it implements all the steps in the asset identification method based on multi-modal fusion visualization map in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:
[0224] A panoramic image of the space is generated based on multiple images of the space where the asset is located and the inertial data when the image data was acquired; wherein, the multiple images are multiple frames of image data that are consecutive in time.
[0225] The location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength are determined by the radio frequency receiver; wherein the asset is equipped with an radio frequency tag;
[0226] The location of the corresponding asset is marked in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0227] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0228] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0229] Example 6:
[0230] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps in the above-described embodiment of the asset identification method based on a multi-modal fusion visualization map, the steps including:
[0231] A panoramic image of the space is generated based on multiple images of the space where the asset is located and the inertial data when the image data was acquired; wherein, the multiple images are multiple frames of image data that are consecutive in time.
[0232] The location of the asset relative to the radio frequency receiver, the corresponding radio frequency signal and its strength are determined by the radio frequency receiver; wherein the asset is equipped with an radio frequency tag;
[0233] The location of the corresponding asset is marked in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and its intensity.
[0234] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, hardware + program embodiments are relatively simple in description because they are fundamentally similar to method embodiments; relevant parts can be referred to the descriptions in the method embodiments.
[0235] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0236] While this application provides method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the method can be executed in the order shown in the embodiments or drawings or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0237] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer programs can be provided.
[0238] To produce a machine by means of a processor in a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device, such that instructions executable by the processor of the computer or other programmable data processing device generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0239] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0240] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0241] This application uses specific embodiments to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for asset identification based on multi-modal fusion visual map, characterized in that, The method comprises: generating a panoramic image of a space in which an asset is located according to a plurality of images of the space and inertial data when the image data is collected; wherein the plurality of images are a plurality of frames of image data that are continuous in time; determining the position of the asset relative to a radio frequency receiver, the corresponding radio frequency signal and its intensity by means of the radio frequency receiver; wherein the asset is provided with a radio frequency tag; marking the position of the corresponding asset in the panoramic image according to a predetermined collection path of the panoramic image, the position, the corresponding radio frequency signal and its intensity; the step of marking the position of the corresponding asset in the panoramic image according to the predetermined collection path of the panoramic image, the position, the corresponding radio frequency signal and its intensity comprises: annotating the type of asset according to the radio frequency signal and a pre-generated mapping relationship between the radio frequency signal and the asset; annotating the type of asset in the panoramic image according to the collection path, the position and the intensity of the radio frequency signal.
2. The asset identification method of claim 1, wherein, annotating the type of asset in the panoramic image according to the collection path, the position and the intensity of the radio frequency signal comprises: determining the initial position of the asset type according to the collection path, the position and the intensity of the radio frequency signal; correcting the initial position according to the marked position on the collection path and the intensity of the radio frequency signal corresponding to the position to determine the final position of the asset type; annotating the type of asset in the panoramic image according to the final position.
3. The asset identification method of claim 1, wherein, The step of determining the collection path comprises: generating a sequence of image frames according to the plurality of images; determining the collection path according to the sequence of image frames and the inertial data.
4. The asset identification method of claim 3, wherein, The step of generating a panoramic image of a space in which an asset is located according to a plurality of images of the space and inertial data when the image data is collected comprises: extracting feature points in the frame image data; matching a plurality of feature points in the sequence of image frames in time order to generate a feature point matching result; determining the relative displacement and rotation angle between consecutive frame images in the sequence of image frames according to the inertial data; stitching the plurality of images after the projection according to the feature point matching result, the relative displacement and the rotation angle to generate the panoramic image.
5. The asset identification method of claim 1, wherein, After the step of marking the position of the corresponding asset in the panoramic image according to the predetermined collection path of the panoramic image, the position, the corresponding radio frequency signal and its intensity, the method further comprises: spatially navigating the asset according to the panoramic image after the asset is marked.
6. An asset identification method based on multi-modal fusion visual map, characterized in that, The method comprises: a panoramic image generation module for generating a panoramic image of a space in which an asset is located according to a plurality of images of the space and inertial data when the image data is collected; wherein the plurality of images are a plurality of frames of image data that are continuous in time; a radio frequency signal determination module for determining the position of the asset relative to a radio frequency receiver, the corresponding radio frequency signal and its intensity by means of the radio frequency receiver; wherein the asset is provided with a radio frequency tag; An asset position marking module is configured to mark the position of the corresponding asset in the panoramic image according to a predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and the intensity thereof. The marking of the position of the corresponding asset in the panoramic image according to the predetermined acquisition path of the panoramic image, the orientation, the corresponding radio frequency signal and the intensity thereof comprises: annotating the type of the asset according to the radio frequency signal and a pre-generated mapping relationship between the radio frequency signal and the asset; annotating the type of the asset in the panoramic image according to the acquisition path, the orientation and the intensity of the radio frequency signal.
7. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the asset identification method based on the multi-mode fusion visual map according to any one of claims 1 to 5.
8. An electronic device, comprising: comprise: at least one processor; and a memory in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the asset identification method based on the multi-mode fusion visual map according to any one of claims 1 to 5.
9. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by a processor to implement the asset identification method based on the multi-mode fusion visual map according to any one of claims 1 to 5.
Citation Information
Patent Citations
Asset AI identification and positioning method and system based on digital twinning
CN114565849A
Park asset management system based on artificial intelligence
CN117495297A