A park asset intelligent identification automatic inventory method fusing multi-modal data
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING QILIANG TECH CO LTD
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-04
AI Technical Summary
[0003]不过,现有园区资产自动盘点方法多侧重某一类数据的采集或识别,较少处理不同模态之间的异步性、空间误差和可信度差异
[0016]The beneficial effects of this invention are as follows: The intelligent identification and automatic inventory method for park assets, which integrates multimodal data, transforms image, RFID, positioning, point cloud, and ledger data into multimodal asset events with time, source, candidate assets, spatial location, confidence parameters, and quality parameters. This allows incomplete on-site evidence to be preserved and used in judgment with controlled weights, reducing missed or incorrect inventorying caused by missing candidate identities or spatial location supplementation. Through reliable spatiotemporal alignment and sliding window offset estimation, stable acquisition delays are distinguished from abnormal temporal offsets, allowing normal offsets such as camera caching, batch RFID uploads, and positioning refresh lags to be compensated, while combinations of abrupt changes, continuous boundary crossings, or physically unreachable events do not contaminate historical offset benchmarks. Through modal confidence scoring and supplementary acquisition verification and write-back, the original identification confidence, data quality, cross-modal consistency, source stability, and anomaly penalties are combined in the fusion output, making the inventory conclusion no longer dependent on fixed modal weights. It can output asset status with a confidence level under conditions of local modal anomalies, ledger lags, or identity conflicts, and continuously feeds back the verification results to subsequent inventory processes.
Smart Images

Figure CN122508431A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal data fusion and recognition technology, specifically to a method for intelligent identification and automatic inventory of park assets that integrates multimodal data. Background Technology
[0002] With the continuous advancement of digital management in smart parks, industrial parks, and large public buildings, asset inventory has gradually evolved from manual checks and barcode scanning to methods such as RFID, visual recognition, indoor positioning, point cloud sensing, and asset ledger linkage. Existing asset management platforms can typically digitally record asset numbers, locations, statuses, and responsibility information. Some systems also combine cameras, RFID readers, UWB / BLE positioning devices, or inspection robots to collect on-site data, thereby improving asset discovery, location, and verification capabilities. In recent years, the development of multimodal sensing, edge computing, digital twins, and intelligent inspection technologies has enabled asset inventory to shift from single-point identification to multi-source data fusion and judgment.
[0003] However, existing automated asset inventory methods in industrial parks often focus on the collection or identification of a single type of data, with limited handling of the asynchronicity, spatial errors, and reliability differences between different modalities. Image recognition is susceptible to occlusion, lighting, angle, and similar appearances; RFID can provide identification information but struggles to directly reflect the true spatial state of assets; location data suffers from refresh delays and drift; point cloud data can describe spatial occupancy but typically lacks asset identification; and ledger data may be inconsistent with the actual site due to delays caused by allocation, maintenance, or manual entry. Existing fusion methods often employ fixed time windows, fixed weights, or simple rule matching, making it difficult to distinguish between normal collection delays and abnormal time-series offsets, and also difficult to determine whether multiple modal events can be mutually explained in terms of time, space, and asset movement logic. When a modality experiences delayed uploads, misreading, drift, or low-quality identification, the fusion result may still output erroneous inventory conclusions with high confidence, leading to asset misalignment, duplicate identification, false alarms of suspected missing assets, or misjudgments of discrepancies between ledger and physical assets. Furthermore, low-reliability and conflicting data often lack closed-loop re-collection and verification mechanisms, causing subsequent inventory counts to repeatedly rely on the same source of anomalies. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by this invention is that existing automatic asset inventory methods for industrial parks suffer from the following problems: the time asynchrony of multi-source data leads to erroneous event association; there is a lack of a reliable spatiotemporal verification mechanism between image, radio frequency identification, positioning, point cloud and ledger data; fixed weight fusion is difficult to suppress high confidence erroneous inventory results caused by abnormal modes; and there is the problem of how to output reliable asset inventory conclusions and trigger supplementary collection or review when there are time offsets, spatial errors and ledger lags.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a method for intelligent identification and automatic inventory of park assets integrating multimodal data, comprising: collecting images, radio frequency identification, positioning, point cloud, and ledger data of park assets to generate multimodal asset events; constructing cross-modal candidate association relationships between different modalities based on candidate assets, event time, and spatial location of the multimodal asset events; inputting the cross-modal candidate association relationships into a reliable spatiotemporal alignment process to obtain temporal consistency, spatial consistency, and causal consistency between asset events; estimating intermodal temporal offsets and identifying abnormal temporal offsets based on temporal consistency, spatial consistency, and causal consistency; calculating the modal reliability score of each multimodal asset event by combining abnormal temporal offsets, quality parameters, and cross-modal consistency; fusing multimodal asset events according to the modal reliability score, outputting asset inventory conclusions, and generating supplementary collection or review tasks.
[0007] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets integrating multimodal data described in this invention, the generation of multimodal asset events includes: unifying the collection time, binding the source identifier, and mapping the spatial coordinates of different modal data respectively; converting each piece of data that can participate in the inventory judgment into an asset event containing event time, source identifier, candidate asset, spatial location, confidence parameter, and quality parameter; candidate assets are formed by at least one of image recognition results, radio frequency identification results, or ledger records; spatial locations are formed by at least one of positioning results, point cloud mapping results, collection source deployment location, or ledger registration location; when a candidate asset is missing but a spatial location exists, the corresponding asset event is retained and entered into the candidate association; when a spatial location is missing but a candidate asset exists, the spatial mapping result is supplemented according to the source identifier or ledger registration location, and the supplemented result is marked as a low-confidence spatial location.
[0008] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets integrating multimodal data described in this invention, the method for constructing cross-modal candidate associations between different modalities includes: within a preset inventory time window, selecting multimodal asset events with different source identifiers for matching; during matching, first comparing the consistency between candidate assets, and increasing the association priority when candidate assets correspond to the same asset record; when candidate assets cannot be directly confirmed, comparing the proximity between event times and the proximity between spatial locations; the proximity of event times is determined based on the difference between the time difference of two asset events and the historical offset benchmark between modalities; the proximity of spatial locations is determined based on the distance between two asset events under unified spatial coordinates; when at least two of the candidate assets, event time, and spatial location meet the preset association conditions, the corresponding cross-modal candidate association is retained; when multiple cross-modal candidate associations compete for the same asset event, they are filtered sequentially according to the consistency of candidate assets, the proximity of spatial locations, and the proximity of event times to form a candidate event group.
[0009] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets integrating multimodal data described in this invention, the method for obtaining temporal consistency, spatial consistency, and causal consistency among asset events includes: reading the event time, source identifier, spatial location, and candidate assets of each multimodal asset event within a candidate event group; calling the corresponding historical time-series offset benchmark based on the source identifier, comparing the deviation between the time difference of asset events within the candidate event group and the historical time-series offset benchmark, and forming temporal consistency; comparing the spatial deviation between asset events within the candidate event group based on the spatial location distance under unified spatial coordinates, and forming spatial consistency; determining whether the asset events within the candidate event group conform to an interpretable event sequence relationship based on the historical movement records, spatial connectivity, and ledger change order of the candidate assets, and forming causal consistency; when temporal consistency, spatial consistency, and causal consistency all meet preset conditions, retaining the candidate event group for time-series offset estimation; when any consistency does not meet the preset conditions, recording the corresponding deviation source and passing the corresponding deviation source to the abnormal time-series offset judgment.
[0010] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets integrating multimodal data described in this invention, the method for estimating intermodal temporal offsets and judging abnormal temporal offsets includes: based on the consistency results output by the reliable spatiotemporal alignment process, statistically analyzing the event time differences in candidate event groups according to source identifiers and spatial regions; extracting the event time differences that occur within a sliding time window to form an estimated value of intermodal temporal offset; comparing the estimated value of intermodal temporal offsets with historical temporal offset benchmarks; determining that there is an abnormal temporal offset for the corresponding source identifier when the estimated value of intermodal temporal offsets continuously exceeds the historical allowable fluctuation range; determining that there is an abnormal temporal offset for the corresponding candidate event group when the estimated value of intermodal temporal offsets undergoes a sudden change in a short period of time; determining that there is an abnormal temporal offset for the corresponding candidate event group when the same candidate asset corresponds to multiple spatial locations that cannot be established simultaneously within the same time window or when the event sequence within the candidate event group is inconsistent with the historical movement record; and writing the abnormal temporal offset judgment result into the corresponding asset event for use in modal reliability scoring.
[0011] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets that integrates multimodal data as described in this invention, the calculation of the modal credibility score of each multimodal asset event includes: for each multimodal asset event, reading the corresponding confidence parameters, quality parameters, temporal consistency, spatial consistency, causal consistency, and abnormal temporal offset judgment results; when the quality parameters of an asset event meet preset conditions and form a stable cross-modal candidate association with asset events from different sources, increasing the corresponding modal credibility score; when an asset event deviates from the candidate event group in terms of time, space, or event sequence, decreasing the corresponding modal credibility score according to the degree of deviation; when an abnormal temporal offset is triggered by the corresponding source identifier, applying an abnormal penalty to the corresponding modal credibility score; when there are multiple mutually supporting asset events for the same candidate asset, selecting the asset event with higher cross-modal consistency to participate in the fusion; when there are conflicting asset events for the same candidate asset, sorting them in order of quality parameters, cross-modal consistency, abnormal temporal offset judgment results, and historical offset stability, and writing the sorting results into the credibility record of the corresponding asset event.
[0012] As a preferred embodiment of the intelligent identification and automatic inventory method for park assets that integrates multimodal data as described in this invention, the step of outputting asset inventory conclusions and generating supplementary collection or review tasks includes: integrating multimodal asset events belonging to the same candidate asset according to modal credibility scores; during asset confirmation, accumulating the modal credibility scores corresponding to the same candidate assets and determining the candidate assets that meet the confirmation conditions as inventory assets; during location confirmation, weighting and correcting the candidate spatial locations according to the modal credibility scores to obtain the integrated spatial locations; during the account-to-physical judgment, performing regional consistency judgment between the integrated spatial locations and the registered locations in the ledger, and forming an asset inventory conclusion based on the ledger status; when a candidate asset does not meet the confirmation conditions, the integrated spatial location is inconsistent with the registered location in the ledger, or the abnormal time series offset is not resolved, outputting a conclusion of suspected missing, suspected misalignment, suspected duplicate identification, or pending review; generating supplementary collection or review tasks based on the source of abnormal time series offsets, low-credibility asset events, and account-to-physical judgment results, and writing the supplementary collection or review results back to the corresponding asset events and historical time series offset benchmarks.
[0013] As a preferred embodiment of the intelligent identification and automatic inventory system for park assets integrating multimodal data described in this invention, the system includes: an event association module, a spatiotemporal verification module, and a reliable inventory module. The event association module collects image, RFID, positioning, point cloud, and ledger data of park assets to generate multimodal asset events. Based on candidate assets, event time, and spatial location of the multimodal asset events, it constructs cross-modal candidate association relationships between different modalities. The spatiotemporal verification module inputs the cross-modal candidate association relationships into a reliable spatiotemporal alignment process to obtain temporal consistency, spatial consistency, and causal consistency between asset events. Based on temporal consistency, spatial consistency, and causal consistency, it estimates inter-modal temporal offsets and identifies abnormal temporal offsets. The reliable inventory module combines abnormal temporal offsets, quality parameters, and cross-modal consistency to calculate the modal reliability score of each multimodal asset event. It then integrates multimodal asset events according to the modal reliability scores, outputs asset inventory conclusions, and generates supplementary collection or review tasks.
[0014] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a method for intelligent identification and automatic inventory of park assets that integrates multimodal data.
[0015] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a method for intelligent identification and automatic inventory of park assets that integrates multimodal data are implemented.
[0016] The beneficial effects of this invention are as follows: The intelligent identification and automatic inventory method for park assets, which integrates multimodal data, transforms image, RFID, positioning, point cloud, and ledger data into multimodal asset events with time, source, candidate assets, spatial location, confidence parameters, and quality parameters. This allows incomplete on-site evidence to be preserved and used in judgment with controlled weights, reducing missed or incorrect inventorying caused by missing candidate identities or spatial location supplementation. Through reliable spatiotemporal alignment and sliding window offset estimation, stable acquisition delays are distinguished from abnormal temporal offsets, allowing normal offsets such as camera caching, batch RFID uploads, and positioning refresh lags to be compensated, while combinations of abrupt changes, continuous boundary crossings, or physically unreachable events do not contaminate historical offset benchmarks. Through modal confidence scoring and supplementary acquisition verification and write-back, the original identification confidence, data quality, cross-modal consistency, source stability, and anomaly penalties are combined in the fusion output, making the inventory conclusion no longer dependent on fixed modal weights. It can output asset status with a confidence level under conditions of local modal anomalies, ledger lags, or identity conflicts, and continuously feeds back the verification results to subsequent inventory processes. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 The present invention provides an overall flowchart of a method for intelligent identification and automatic inventory of park assets that integrates multimodal data.
[0019] Figure 2 A schematic diagram of a computer device provided by the present invention. Detailed Implementation
[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0021] Reference Figure 1 As an embodiment of the present invention, a method for intelligent identification and automatic inventory of park assets integrating multimodal data is provided, comprising:
[0022] S1: Collect images, RFID, positioning, point cloud, and ledger data of assets in the park to generate multimodal asset events.
[0023] Furthermore, generating multimodal asset events includes unifying the collection time, binding source identifiers, and mapping spatial coordinates for different modal data; converting each piece of data eligible for inventory assessment into an asset event containing event time, source identifier, candidate asset, spatial location, confidence parameter, and quality parameter; candidate assets are formed from at least one of image recognition results, RFID results, or ledger records; spatial locations are formed from at least one of positioning results, point cloud mapping results, collection source deployment location, or ledger registration location; when a candidate asset is missing but a spatial location exists, the corresponding asset event is retained and added to the candidate association; when a spatial location is missing but a candidate asset exists, spatial mapping results are supplemented based on the source identifier or ledger registration location, and the supplemented results are marked as low-confidence spatial locations.
[0024] It should be noted that one specific approach to generating multimodal asset events involves considering that, in a park asset inventory scenario, various data sources typically have different data formats, upload frequencies, and spatial representations. For example, image data more easily reflects asset appearance and tag information, RFID data more easily reflects asset identity, location data more easily reflects real-time asset location, point cloud data more easily reflects spatial occupancy status, and ledger data more easily reflects registered identity and location. Directly fusing raw data from different modalities can easily lead to problems such as incomparable timestamps, inconsistent spatial coordinates, and inconsistent identity fields. Therefore, this step first converts data from different sources into a unified multimodal asset event.
[0025] In this step, the inputs are image data, RFID data, location data, point cloud data, and ledger data collected within the park. The data from different modalities are processed by standardizing the collection time, binding source identifiers, and mapping spatial coordinates, transforming data suitable for inventory assessment into multimodal asset events.
[0026] Any multimodal asset event is denoted as:
[0027]
[0028] in, Indicates the first A multimodal asset event. Indicates the time of the event. Indicates the source identifier. Indicates candidate assets. Indicates spatial location, Indicates the confidence parameter. Indicates quality parameters.
[0029] When the data collection time is unified, the local time carried by different data collection sources is converted to a unified time. If the data collection source provides a synchronized time, then the synchronized time is used as the event time. If the data source only provides local time, then the most recent clock correction record from the corresponding source is read, and the corrected time is used as the event time. If the data source lacks clock calibration records, the data reception time will be used as the event time. At the same time, the effective value of the corresponding time is reduced. This is because subsequent abnormal time series offset judgment does not require the acquisition end to be completely synchronized, but it is necessary for each asset event to be compared under at least the same time base.
[0030] When binding source identifiers, a source identifier is bound to each piece of original data. The source identifier at least distinguishes between the data modality and the acquisition source, enabling subsequent steps to call historical time-series offset benchmarks based on the source identifier, identify abnormal sources, and generate supplementary acquisition tasks for specific sources or modalities when needed.
[0031] During spatial coordinate mapping, the positions corresponding to different modes are transformed to a unified spatial coordinate system. Positioning data is then used to determine spatial location based on these coordinates. Point cloud data forms spatial locations based on the spatially occupied area or the target center. Image data forms spatial positions based on the calibration parameters of the acquisition source, the field of view of the acquisition source, and the relative position of the image recognition bounding box within the field of view. The ledger data is based on the registered area to form the registered spatial location. If the spatial location is a coordinate point, then... Recorded as coordinate vectors; if the spatial location is a region, then Record it as the center point of the region and associate it with the region boundary to enable spatial distance calculation and region consistency judgment.
[0032] Candidate Assets It is formed from at least one of image recognition results, RFID results, or ledger records. Image data is processed by target recognition or text recognition to obtain asset category, nameplate text, or asset number; RFID data is parsed by tag encoding to obtain tag number or asset number; ledger data is read from asset records to obtain asset number and registration location. If a candidate asset can correspond to a unique asset number, the candidate asset record is an identity candidate; if a candidate asset can only correspond to an asset category, the candidate asset record is a category candidate; if a candidate asset is missing but the spatial location exists, the corresponding asset event is not directly discarded, but is retained and entered into subsequent candidate associations.
[0033] Confidence parameters The value ranges from 0 to 1. The confidence parameter for image data is the confidence level of the target recognition or text recognition output; the confidence parameter for RFID data is determined based on whether the integrity of the tag encoding parsing and the reading strength are within the preset reading range; the confidence parameter for positioning data is determined based on whether the positioning error radius is within the tolerance of the corresponding area; the confidence parameter for point cloud data is determined based on whether the continuity and density of the point cloud targets reach the preset lower limit; the confidence parameter for ledger data is determined based on the interval between the ledger update time and the current inventory time, with a lower confidence parameter for longer intervals.
[0034] Since not all asset events have complete candidate assets and spatial locations, this step further introduces quality parameters. This is to prevent missing fields from directly causing asset events to be discarded, and also to prevent supplementary information from being equated with measured information in subsequent fusion. Quality parameters Calculated from temporal RMS, spatial RMS, and candidate RMS:
[0035]
[0036] in, Indicates the first Quality parameters for individual asset events; This represents the valid time value. It is set to 1 when the acquisition time is obtained after synchronization or correction, 0.5 when the received time is used instead, and 0 when the event time cannot be determined. The value represents the valid spatial location. It is 1 when the spatial location is directly formed by mapping from positioning, point cloud, or calibration image; 0.5 when it is supplemented by the deployment location of the data source or the location registered in the ledger; and 0 when the spatial location cannot be formed. This represents the valid candidate value. It is 1 when the candidate asset is a unique identity candidate, 0.5 when the candidate asset is a category candidate, and 0 when the candidate asset is missing. , and These represent the weights of the temporal effective value, spatial effective value, and candidate effective value, respectively, and their sum is 1. Preferably, , and Values are all taken in the range of 0.3 to 0.4, so that the event time, spatial location and candidate assets can all affect whether the asset event is suitable to participate in subsequent calculations.
[0037] When candidate assets are missing but spatial locations exist The value is set to 0, but the corresponding asset event is still retained because the corresponding event can be associated with other modal events through spatial location and event time. When the spatial location is missing but the candidate asset exists, the spatial mapping result is supplemented based on the source identifier or ledger registration location, and the result is made... A value of 0.5 is used to allow supplementary spatial locations to participate in spatial consistency assessment, but they will not participate with equal weight to spatial locations formed by positioning data or point cloud data. If neither the event time nor the source identifier can be determined, no multimodal asset event will be generated for the corresponding original data; if the event time can be generated substituted and the source identifier exists, a low-quality asset event will be generated and proceed to subsequent steps.
[0038] S2: Based on the candidate assets, event time, and spatial location of multimodal asset events, construct cross-modal candidate association relationships between different modalities.
[0039] Furthermore, constructing cross-modal candidate associations between different modalities includes: within a preset inventory time window, selecting multimodal asset events with different source identifiers for matching; during matching, first comparing the consistency between candidate assets, increasing the association priority when candidate assets correspond to the same asset record; when candidate assets cannot be directly confirmed, comparing the proximity between event times and the proximity between spatial locations; the proximity of event times is determined based on the difference between the time difference of two asset events and the historical offset benchmark between modalities; the proximity of spatial locations is determined based on the distance between two asset events under a unified spatial coordinate system; when at least two of the candidate assets, event time, and spatial location meet the preset association conditions, the corresponding cross-modal candidate association is retained; when multiple cross-modal candidate associations compete for the same asset event, they are filtered sequentially according to the consistency of candidate assets, the proximity of spatial locations, and the proximity of event times to form a candidate event group.
[0040] It should be noted that one specific approach to constructing cross-modal candidate associations between different modalities involves recognizing that in multimodal asset inventory, events of different modalities do not necessarily occur simultaneously, nor can they all provide complete asset identification. For example, RFID events may only provide tag identification with low spatial precision, image events may provide appearance category but not serial number, and location events may only provide coordinates. Matching based solely on a single field can easily miss valid associations; directly merging all events within a time window can easily lead to false associations. Therefore, this step constructs cross-modal candidate associations through three types of relationships: candidate assets, event time, and spatial location, ensuring that reliable spatiotemporal alignment is performed only between events that may describe the same asset.
[0041] In this step, the input is the set of multimodal asset events generated by S1, and the output is cross-modal candidate associations and candidate event groups. Cross-modal candidate associations indicate that asset events identified from different sources may describe the same candidate asset or the same on-site inventory object.
[0042] Within a preset inventory time window, multimodal asset events with different source identifiers are selected for matching. The preset inventory time window is preferably 30 seconds to 5 minutes; for scenarios with a high density of fixed collection sources, it is 30 seconds to 60 seconds; and for scenarios involving mobile inspection collection sources, it is 1 minute to 5 minutes. The above time window can cover asynchronous upload latency while avoiding the possibility of different asset events being mistakenly associated due to an excessively long window.
[0043] For any two asset events with different source identifiers and Calculate candidate association scores:
[0044]
[0045] in, Indicates asset events With asset events Candidate association scores between; Indicates the degree of consistency among candidate assets; Indicates the proximity of events in time; Indicates the degree of spatial proximity; , and These represent the weights of the candidate asset, event time, and spatial location in the candidate association, respectively, and their sum is 1. Preferably, in scenarios where the asset number or tag number is stable, Take a value of 0.5 to 0.7. and Take values ranging from 0.15 to 0.25; in scenarios with a large number of unlabeled assets, reduce... And improve This allows spatial location to have a greater impact on candidate associations.
[0046] Consistency of candidate assets Values are determined according to the following rules: when both candidate assets for two asset events are unique identity candidates and correspond to the same asset record. Set to 1; when one asset event is a unique identity candidate, another asset event is a category candidate, and the category of the unique identity candidate matches the category candidate in the ledger, Take 0.6; when both asset events are category candidates and belong to the same category, Take 0.4; when candidate assets are missing but there is no identity conflict, Take 0.2; when the candidate asset clearly corresponds to a different asset record, Set the value to 0. This hierarchical approach prioritizes unique identifiers in the association process while retaining the auxiliary roles of category and spatial identification for unlabeled assets.
[0047] proximity of events Determined based on the degree of deviation between the time difference of two asset events and the historical offset benchmark between modalities:
[0048]
[0049] in, Indicates asset events With asset events The degree of temporal proximity of the events between them; and These represent the event times of the two asset events respectively; and These represent the source identifiers of the two asset events; Source identifier Source Identifier Historical time-series offset reference; This indicates the coarse screening time tolerance used in the candidate association phase. This applies to the first run when there is no historical time-series offset reference. Take 0, Take 30 to 120 seconds; during subsequent runs, Updated via S4 writeback.
[0050] Spatial proximity Determined based on the distance between two asset events in a unified spatial coordinate system:
[0051]
[0052] in, Indicates asset events With asset events The degree of spatial proximity between them; and These represent the spatial locations of the two asset events, respectively. This represents the distance between two spatial locations in a unified spatial coordinate system. This indicates the regional spatial tolerance used in the candidate association phase; This indicates the spatial region to which two asset events belong or are most recently located. If two spatial locations are in the same region, their spatial proximity is high; if two spatial locations are in adjacent regions, their spatial proximity is calculated based on the distance between the center of the regions; if two spatial locations are in non-connected regions, then... Take 0.
[0053] When at least two of the following criteria are met: consistency of candidate assets, proximity of event time, and proximity of spatial location, or candidate correlation score, the candidate correlation score will be determined. Reaching the preset association threshold At that time, the corresponding cross-modal candidate association relationships are retained. A preset association threshold is set. The optimal threshold is between 0.6 and 0.75. A threshold below 0.6 is prone to introducing false associations, while a threshold above 0.75 is prone to missing asset events where candidate assets are missing but their spatial locations are reliable. Therefore, a threshold of 0.65 is preferred for general park inventory scenarios.
[0054] When multiple cross-modal candidate associations compete for the same asset event, they are selected sequentially based on the degree of consistency of candidate assets, the degree of spatial proximity, and the degree of temporal proximity of the events. Specifically, a candidate association graph is constructed with asset events as nodes and cross-modal candidate associations as edges; candidates with scores no lower than [a certain threshold] are retained. The edges are defined; sets of nodes that are interconnected and have at least two source identifiers are formed into candidate event groups; if the same asset event enters multiple candidate event groups simultaneously, the candidate event group with the highest total candidate association score is selected, and the other candidate event groups are marked as conflict candidates. Sets of nodes with only one source identifier do not enter the main process of trusted spatiotemporal alignment, but are treated as objects to be supplemented or reviewed.
[0055] S3: Input the cross-modal candidate association into the trusted spatiotemporal alignment process to obtain the temporal consistency, spatial consistency and causal consistency between asset events.
[0056] Furthermore, obtaining temporal consistency, spatial consistency, and causal consistency among asset events includes: reading the event time, source identifier, spatial location, and candidate assets of each multimodal asset event within the candidate event group; calling the corresponding historical time-series offset benchmark based on the source identifier, comparing the deviation between the time difference of asset events within the candidate event group and the historical time-series offset benchmark, and establishing temporal consistency; comparing the spatial deviation between asset events within the candidate event group based on the spatial distance under unified spatial coordinates, and establishing spatial consistency; determining whether the asset events within the candidate event group conform to an interpretable event sequence based on the historical movement records, spatial connectivity, and ledger change order of the candidate assets, and establishing causal consistency; when temporal consistency, spatial consistency, and causal consistency all meet preset conditions, the candidate event group is retained for time-series offset estimation; when any consistency does not meet the preset conditions, the corresponding deviation source is recorded and the corresponding deviation source is passed to the abnormal time-series offset judgment.
[0057] It should be noted that one specific approach to obtaining temporal, spatial, and causal consistency among asset events involves obtaining candidate event groups that can only indicate that multiple events are "potentially related," but cannot directly prove that these events can be considered valid inventory evidence for the same asset. This is because asset inventory in industrial parks frequently encounters issues such as RFID misreading, image occlusion, location drift, point cloud delays, or outdated ledgers. Directly merging candidate events based solely on high correlation scores can easily lead to high-confidence errors. Therefore, this step further performs credible spatiotemporal alignment of candidate event groups from three dimensions: time, space, and causality, ensuring that offset estimation and confidence scores are based on validated event relationships.
[0058] In this step, the input is a group of candidate events, and the output is the temporal consistency, spatial consistency, and causal consistency among asset events within the candidate event group. Temporal consistency and spatial consistency are used to quantify the degree of consistency between different asset events within the candidate event group in terms of time and space, while causal consistency is used to determine whether the sequential relationship of multiple asset events conforms to the logic of asset movement, spatial connectivity, and ledger changes.
[0059] For candidate event groups Any two asset events and Consistency in computation time:
[0060]
[0061] in, Indicates asset events With asset events Time consistency between them; and These represent the event times of the two asset events respectively; and These represent the source identifiers of the two asset events; Source identifier Source Identifier The historical time offset reference retained in the previous sliding time window; This indicates the time tolerance retained for the corresponding source identifier combination in the previous sliding time window. The initial value of the time tolerance is determined by the acquisition period and the network's allowable latency, preferably 2 to 3 times the larger value of the upload periods of the two corresponding sources, and not less than 5 seconds;
[0062] For candidate event groups Any two asset events and Computational space consistency:
[0063]
[0064] in, Indicates asset events With asset events Spatial consistency between them; This represents the uniform spatial coordinate distance between two spatial locations; Representing a spatial region Corresponding spatial tolerance; and These represent the spatial validity values formed by two asset events in S1. When the spatial location is directly formed by positioning, point cloud, or calibration image mapping, the spatial validity value is higher; when the spatial location is supplemented by the location registered in the ledger or the deployment location of the data source, the spatial validity value is lower. By multiplying by the spatial validity value, it is possible to avoid the supplemented spatial location and the measured spatial location being equally weighted in the spatial consistency judgment.
[0065] Spatial tolerance The inventory accuracy setting should be adjusted accordingly. For room-level inventory, a range of 1 to 3 meters is preferred; for shelf-level inventory, a range of 0.3 to 1 meter is preferred; and for floor-level inventory, a range of 5 to 10 meters is preferred. If two spaces are located on different floors and there is no accessible connection between the floors, the spatial consistency should be set to 0.
[0066] Causal consistency is used to handle event relationships that are difficult to explain solely by time and spatial distance. For example, two events may appear close in time and distance, but if the corresponding assets could not have moved from one area to another within that timeframe, or if the order of ledger changes clearly contradicts the order of on-site identification, then the candidate event group should not be directly considered a reliable fusion object. Therefore, for candidate event groups... Asset events within the system are sorted by event time to obtain adjacent event pairs; for each adjacent event pair, the shortest traversable path length between the two spatial locations is read. And read the maximum allowed movement speed corresponding to the candidate asset. .
[0067] If there is no passable path between two spatial locations, the causal consistency is classified as low-level; if there is a passable path, but the following conditions are met... If the causal consistency is low, then the causal consistency is determined to be low. This represents the time tolerance, preferably between 1 and 3 seconds. If a spatial movement relationship is established, but there is a single conflict between the order of changes in the ledger and the order of on-site identification, then the causal consistency is judged as medium level. If the spatial movement relationship, the time sequence of events, and the order of changes in the ledger can all be explained, then the causal consistency is judged as high level. Maximum permissible movement speed of fixed assets. The preferred speed is 0 to 0.5 m / s, the preferred speed for mobile devices is 1 to 3 m / s, and the speed for vehicle assets is calculated based on the park's speed limit.
[0068] Numericalize causal consistency as High-level events are assigned a value of 1, medium-level events are assigned 0.5, and low-level events are assigned 0. For candidate event groups... Group-level temporal consistency and group-level spatial consistency are formed by the average or minimum values of event pairs within the group, respectively. The minimum value is used for critical asset inventory to improve anomaly sensitivity; the average value is used for ordinary asset inventory to reduce the impact of individual noise events on overall judgment. To ensure a clear default implementation path, this embodiment uses the average value to form group-level consistency:
[0069]
[0070] in, Indicates candidate event group Group-level time consistency Indicates candidate event group Group-level space consistency; Indicates candidate event group The source identifies the collection of asset event pairs; Represents a set The number of event pairs in the data; and These represent asset events. With asset events Temporal and spatial consistency between them.
[0071] when Not lower than the time consistency threshold , Not lower than the spatial consistency threshold ,and Not lower than the causal consistency threshold At that time, retain candidate event groups Proceed to S4. Preferably, and Take 0.5, Set the value to 0.5. If any consistency threshold falls below the corresponding threshold, record the source of the deviation. The source of the deviation includes the source identifier, deviation type, and deviation magnitude. Deviation types include temporal deviation, spatial deviation, and causal deviation. The deviation magnitude is the difference between the corresponding consistency threshold and the actual consistency value. The deviation source is then entered into S4 for abnormal time series offset judgment.
[0072] S4: Estimate intermodal temporal offsets and identify anomalous temporal offsets based on temporal consistency, spatial consistency, and causal consistency.
[0073] Furthermore, estimating intermodal temporal offsets and identifying anomalous temporal offsets includes: based on the consistency results output by the trusted spatiotemporal alignment process, statistically analyzing the event time differences in candidate event groups according to source identifiers and spatial regions; extracting the event time differences occurring within a sliding time window to form an estimated intermodal temporal offset; comparing the estimated intermodal temporal offset with historical temporal offset benchmarks; determining that there is an anomalous temporal offset for the corresponding source identifier when the estimated intermodal temporal offset continuously exceeds the historical allowable fluctuation range; determining that there is an anomalous temporal offset for the corresponding candidate event group when the estimated intermodal temporal offset undergoes a sudden change within a short period of time; determining that there is an anomalous temporal offset for the corresponding candidate event group when multiple spatial locations corresponding to the same candidate asset cannot be simultaneously established within the same time window, or when the event order within the candidate event group is inconsistent with historical movement records; and writing the anomalous temporal offset judgment results into the corresponding asset event for use in modal credibility scoring.
[0074] It should be noted that a preferred approach to estimating intermodal temporal offsets and identifying anomalous offsets specifically includes the understanding that timing issues in park asset inventory do not always manifest as obvious errors. Camera video caching, RFID batch uploads, positioning refresh intervals, point cloud scanning cycles, and lag in ledger updates can all cause stable temporal offsets between different modal events. Simply requiring identical timestamps for different modal events would lead to the accidental deletion of numerous valid events; conversely, completely relaxing time constraints would result in the incorrect merging of asset states from different times. Therefore, this step estimates intermodal temporal offsets from trusted candidate event pairs using a sliding time window and determines whether the offsets are anomalous based on historical fluctuations.
[0075] In this step, the inputs are the consistency results, candidate event groups, and sources of deviation from the S3 output, and the outputs are the intermodal time series offset estimates, the abnormal time series offset judgment results, and the historical time series offset baseline update results.
[0076] To avoid confusion regarding the offset directions of different source identifier combinations, this embodiment pre-fixes the direction of the source identifier combinations. For any two source identifiers... and Determined according to preset source sorting Direction, and always using the source identifier as The asset event time minus the source identifier The time difference between asset events is calculated. The resulting intermodal time series offset has a fixed positive and negative direction, and subsequent historical benchmark updates and anomaly detection all follow the same direction.
[0077] In the sliding time window Inside, the source is identified as and Candidate event pairs are selected. The sliding time window is preferably 5 to 15 minutes; 5 minutes is used when there is a high density of fixed data sources, and 10 to 15 minutes is used when mobile inspection data sources are involved. Candidate event pairs participating in time series offset estimation must meet basic credibility conditions, including a candidate association score reaching a certain level. Spatial consistency is not lower than Causal consistency is no less than This avoids events with obvious spatial errors or causal contradictions from being included in the offset estimation.
[0078] Considering that the time difference of a single event pair may be affected by network jitter or a single misidentification, this step does not use the single time difference as the offset estimate. Instead, it uses the median of multiple reliable event pairs within the sliding window to form the offset estimate.
[0079]
[0080] in, Source identifier Source Identifier In the Estimated intermodal temporal offset within a sliding time window; This indicates taking the median of the time differences within the set; and These represent asset events. and asset events The time of the event; and These represent asset events. and asset events Source identifier; Indicates the first Within each sliding time window, the basic trust conditions are met and the source identifiers are respectively... and The candidate event pair set. Using the median can reduce the impact of occasional network latency, temporary occlusion, or single misreads on offset estimation.
[0081] when The number of candidate event pairs is less than the minimum sample size. At that time, the historical time series offset reference is not updated, and the source identifier is combined. Marked as insufficient sample size. Minimum sample size. A score of 5 is preferred; a score of 3 is acceptable for sparsely populated asset regions. Insufficient samples will not be directly identified as an abnormal time series shift, but the historical stability of the relevant source identifiers will be reduced in S5.
[0082] To distinguish between stable and abnormal offsets, this step further calculates the historical allowable fluctuation range. The historical allowable fluctuation range cannot rely solely on a fixed threshold because the upload cycles and network environments differ across data sources; nor can it depend entirely on historical fluctuations, as abnormal windows may amplify them. Therefore, this embodiment uses both historical offset fluctuations and the minimum tolerable offset as the judgment boundary:
[0083]
[0084] in, Source identifier Source Identifier In the The historical allowable fluctuation range of a sliding time window; This represents the fluctuation amplification factor, which is preferably set to 3; Indicates source identifier combination Offset fluctuation value within the most recent historical window; This represents the minimum tolerable offset, preferably between 2 and 5 seconds. Offset fluctuation value. It can be formed based on the median absolute deviation of the historical time series offset benchmark within the most recent normal window, making it less likely for individual abnormal windows to amplify the fluctuation range. The minimum tolerance offset can prevent normal jitter at the millisecond or low-second level from being misjudged as abnormal.
[0085] An abnormal timing offset is determined to exist when any of the following conditions are met: First, In continuous Established within a sliding time window, among which This indicates the historical time series offset reference retained in the previous sliding time window. Option 2 is preferred; option 1 is suitable for critical asset inventory scenarios; and option 3 is suitable for ordinary low-risk asset inventory scenarios. Second... This indicates a sudden change in the current offset estimate compared to the previous window; third, the same candidate asset may correspond to multiple spatial locations that cannot be simultaneously established within the same time window.
[0086] Multiple spatial locations that cannot be simultaneously determined are judged based on spatial connectivity and movement time. If there is no traversable path between two spatial locations, they are determined not to be simultaneously possible; if a traversable path exists, but the shortest traversable path length is limited... Greater than If both conditions cannot be met simultaneously, then the judgment is as follows: This indicates the maximum allowed movement speed for the candidate asset. This indicates the time tolerance. Although this judgment manifests as a spatial conflict, it is recorded as an external manifestation of a timing anomaly in the candidate event group in this step, because asynchronous uploads or incorrect timestamps can cause the same candidate asset to be associated with physically inaccessible locations within the same time window.
[0087] When no abnormal time series offset is triggered and the number of samples meets the requirements, update the historical time series offset baseline:
[0088]
[0089] in, Source identifier Source Identifier In the The historical time series offset baseline updated by a sliding time window; This indicates the historical time series offset reference retained in the previous sliding time window; This represents the estimated intermodal temporal offset for the current sliding time window; This represents the baseline update factor, preferably between 0.05 and 0.2. A smaller update factor results in a more stable historical baseline; a larger update factor allows the historical baseline to adapt more quickly to long-term equipment latency changes. When an abnormal timing offset is triggered, the offset estimate of the corresponding window is not used to update the historical timing offset baseline to avoid abnormal data contaminating the baseline.
[0090] The results of abnormal time series offset judgments are written to the corresponding asset event. The written information includes the source of the anomaly, the type of anomaly, the magnitude of the deviation, whether the sample size is insufficient, and whether it participated in the historical time series offset benchmark update. The results of abnormal time series offset judgments are then entered into S5 as the basis for anomaly penalties in the modal credibility scoring.
[0091] It should also be noted that estimating intermodal temporal offsets and identifying anomalous temporal offsets treats the time difference between different sources as a learnable and updatable state variable, rather than requiring complete synchronization at the acquisition end. In a park, camera caching, RFID batch uploading, positioning refresh cycles, and point cloud scanning cycles often create stable latency. Forcing alignment according to the original timestamps can misjudge asset events that should correspond as inconsistent; completely relaxing time conditions can lead to incorrect fusion of asset states at different times. This invention selects trusted event pairs that have passed spatial and causal verification within a sliding time window and uses the median to estimate intermodal temporal offsets, reducing the impact of occasional network jitter and single misreads on offset estimation. Furthermore, the historical allowable fluctuation range is jointly determined by historical offset fluctuations and the minimum tolerable offset, ensuring that normal jitter is not misjudged as anomalous, while continuous out-of-bounds movements or short-term abrupt changes can be identified. When no anomalies are triggered, the offset estimation is smoothly written back to the historical baseline; when anomalies are triggered, the corresponding window does not participate in the baseline update. This allows for the continued use of multimodal evidence under conditions of stable latency, while reducing the risk of abnormal timestamps or local delays contaminating subsequent inventory benchmarks.
[0092] S5: Calculate the modal credibility score for each multimodal asset event by combining anomalous time series offsets, quality parameters, and cross-modal consistency.
[0093] Furthermore, calculating the modal credibility score for each multimodal asset event includes: for each multimodal asset event, reading the corresponding confidence parameters, quality parameters, temporal consistency, spatial consistency, causal consistency, and abnormal temporal offset judgment results; when the quality parameters of an asset event meet preset conditions and form a stable cross-modal candidate association with asset events from different sources, increasing the corresponding modal credibility score; when an asset event deviates from the candidate event group in terms of time, space, or event sequence, decreasing the corresponding modal credibility score according to the degree of deviation; when an abnormal temporal offset is triggered by the corresponding source identifier, applying an abnormal penalty to the corresponding modal credibility score; when multiple mutually supporting asset events exist for the same candidate asset, selecting the asset event with higher cross-modal consistency to participate in the fusion; when conflicting asset events exist for the same candidate asset, sorting them in order of quality parameters, cross-modal consistency, abnormal temporal offset judgment results, and historical offset stability, and writing the sorting results into the credibility record of the corresponding asset event.
[0094] It should be noted that a preferred scheme for calculating the modal confidence score of each multimodal asset event specifically includes addressing the issue that, in the presence of temporal offsets, positioning errors, and local modal anomalies, using fixed weights to fuse data from different modalities can easily lead to a situation where a high confidence error in one modality dominates the inventory results. For example, RFID readings may be accurate but the location may be inaccurate, image recognition may have high confidence but the corresponding video frame may be lagging, and positioning data may be continuous but may exhibit drift. Therefore, this step does not directly use the original confidence score of a single modality, but instead incorporates the original confidence score, data quality, cross-modal consistency, historical stability of the source, and anomaly penalties into the modal confidence score, allowing subsequent fusion to be dynamically determined based on the current inventory scenario.
[0095] In this step, the inputs are the confidence parameters and quality parameters obtained in S1, the temporal consistency, spatial consistency and causal consistency obtained in S3, and the abnormal temporal offset judgment results obtained in S4. The output is the modal confidence score for each multimodal asset event.
[0096] For any asset event First, read the corresponding confidence parameters. Quality parameters Source identification Consistency results of participating candidate event groups and results of abnormal time series offset judgment. If an asset event... If multiple candidate event groups are involved, the candidate event group with the highest candidate association score will be used for scoring first; if the difference in candidate association scores between two candidate event groups is less than 0.1, the conflict mark will be retained, and identity or location conflict judgment will be triggered in S6.
[0097] Asset events The cross-modal consistency synthesis value is formed as follows:
[0098]
[0099] in, Indicates the first Cross-modal consistency composite value for individual asset events; Indicates the first The average time consistency between an asset event and other source-identified asset events within the same candidate event group; Indicates the first Average spatial consistency between an asset event and other source-identified asset events within the same candidate event group; Indicates the first Numericalized results of causal consistency of candidate event groups to which individual asset events belong; , and These represent the weights of temporal consistency, spatial consistency, and causal consistency in the overall cross-modal consistency value, respectively, and their sum is 1. In a typical campus inventory scenario, , and The values can be 0.3, 0.4, and 0.3 respectively; when the latency of the data source in the park fluctuates significantly, the value should be increased. When high spatial positioning accuracy is required, improve .
[0100] Source identification In the The historical stability of each sliding time window is determined as follows:
[0101]
[0102] in, Source identifier In the The historical stability corresponding to each sliding time window; Source identifier In recent The number of windows that form valid candidate associations within a sliding time window without triggering abnormal time offsets; Source identifier In recent The number of asset events within a sliding time window. Preferably, Take a range of 10 to 30. If... ,but A value of 0.5 indicates a neutral stability level when there are no historical records. Historical stability means that occasional anomalies will not completely negate the corresponding source, but long-term anomaly sources will be weighted less in the credibility score.
[0103] Abnormal penalty value Determined based on the source of the anomaly. If the anomaly's timing offset is clearly caused by an asset event. If triggered by the source identifier, then Set to 1; if the anomaly occurs during an asset event. If it belongs to the same candidate event group, but cannot be located to a single source identifier, then Take 0.5; if the abnormal time offset is not triggered simply due to insufficient samples, then... Set it to 0.2; if there are no abnormal time-series offsets and the sample size meets the requirements, then... Set it to 0. This setting can distinguish between explicit anomalies, indeterminate anomalies within a group, and insufficient samples, avoiding the direct equation of insufficient samples with anomalous attacks or equipment failures.
[0104] Modal credibility scores are calculated as follows:
[0105]
[0106] in, Indicates the first Modal credibility score for a multimodal asset event; This indicates that the calculation result within the parentheses is limited to the range of 0 to 1; Indicates the confidence parameter; Indicates quality parameters; This represents the cross-modal consistency composite value; Source identifier In the The historical stability of each sliding time window; Indicates the penalty value for an anomaly; , , , and These represent the weights of the corresponding items. In a typical park inventory scenario, , , and The abnormal penalty weights are set to 0.2, 0.2, 0.4, and 0.2 respectively. Set to 0.3; increase when the attack risk is high. When the quality of the data acquisition fluctuates significantly, improve... The aforementioned weights make cross-modal consistency the primary influencing factor, while preserving the roles of original identification confidence, data quality, and source stability.
[0107] An asset event is considered to have cross-modal support when its quality parameter reaches 0.5 or higher, it forms a candidate association with at least one asset event from a different source, and both its temporal and spatial consistency are not lower than the corresponding thresholds. Asset events with cross-modal support are considered to have higher quality parameters. Improve modal credibility scores. If asset events deviate in time, space, or sequence, consistency decreases, and this is addressed through [further measures]. Reduce modal confidence score. If the anomalous temporal offset is triggered by the corresponding source identifier, then... Impose abnormal penalties.
[0108] When multiple mutually supporting asset events exist for the same candidate asset, asset events with a modal credibility score of at least 0.5 are retained for S6 fusion. When conflicting asset events exist for the same candidate asset, they are fused according to their modal credibility scores. The data is sorted from highest to lowest score. If the difference between the highest and second-highest scores is not less than 0.15, the highest-scoring asset event is prioritized for fusion. If the difference is less than 0.15, a conflict marker is retained, and a review conclusion or supplementary data collection task is generated in S6. The sorting results are written to the credibility record of the corresponding asset event, serving as input for S6 fusion and write-back.
[0109] It should also be noted that the modal credibility score quantifies the credibility of each asset event in the current inventory scenario, rather than using fixed modal weights or single identification confidence levels. In actual inventory, a high initial confidence level for a particular modality does not necessarily indicate final credibility. For example, image recognition may have high confidence but the image may be lagging, RFID may show a genuine identity but a coarse location range, and location data may be continuous but drifting. This invention incorporates initial confidence parameters, quality parameters, cross-modal consistency, source historical stability, and anomaly penalties into the score. Cross-modal consistency integrates temporal, spatial, and causal verification results, giving higher scores to events that mutually support each other from other sources. Historical stability reflects the proportion of sources that normally participate in association within the most recent window, gradually reducing the weight of long-term anomalous sources. Anomaly penalties distinguish between explicit anomalies, uncertain anomalies within a group, and insufficient samples, avoiding equating insufficient samples directly with attacks or malfunctions. Through this scoring mechanism, the fusion stage is not dominated by locally high-confidence results of a single modality, but instead prioritizes asset events that are more consistent in terms of current time, space, and historical stability, reducing the erroneous inclusion of high-confidence mis-inventory and low-quality events in the fusion process.
[0110] S6: Integrate multimodal asset events according to modal credibility scores, output asset inventory conclusions and generate supplementary collection or review tasks.
[0111] Furthermore, the process of outputting asset inventory conclusions and generating supplementary collection or review tasks includes: fusing multimodal asset events belonging to the same candidate asset according to modal credibility scores; during asset confirmation, accumulating the modal credibility scores corresponding to the same candidate assets and identifying candidate assets that meet the confirmation conditions as inventory assets; during location confirmation, weighting and correcting the candidate spatial locations according to the modal credibility scores to obtain the fused spatial locations; during the account-to-physical verification, performing regional consistency judgment between the fused spatial locations and the ledger registration locations, and forming asset inventory conclusions based on the ledger status; when candidate assets do not meet the confirmation conditions, the fused spatial locations are inconsistent with the ledger registration locations, or abnormal time series offsets are not resolved, outputting conclusions of suspected missing, suspected misalignment, suspected duplicate identification, or pending review; generating supplementary collection or review tasks based on the source of abnormal time series offsets, low-credibility asset events, and account-to-physical verification results, and writing the supplementary collection or review results back to the corresponding asset events and historical time series offset benchmarks.
[0112] It should be noted that a preferred approach to outputting asset inventory conclusions and generating supplementary collection or review tasks specifically includes the following: After S5 scoring, different asset events have comparable modal credibility. However, the ultimate goal of asset inventory is not simply to select the highest-scoring event, but to confirm asset identity, correct asset location, determine consistency between accounts and physical assets, and form a closed-loop processing mechanism for uncertain conclusions. Especially when some modalities have temporal anomalies or identity conflicts, directly outputting definitive conclusions can easily lead to incorrect or missed inventory counts. Therefore, this step, through identity verification, location fusion, account-physical asset judgment, and task write-back, ensures that the inventory conclusions can utilize multimodal evidence while retaining subsequent processing paths for low-credibility or conflicting events.
[0113] In this step, the inputs are the modal credibility score and credibility record output by S5, and the outputs are the asset inventory conclusion, supplementary collection or review task, and write-back results.
[0114] First, group multimodal asset events that fall within the same conflict fusion scope into fusion groups. Fusion Group This includes asset events where candidate assets are identical or competing within the same inventory window. For candidate assets... Calculate the identity verification score:
[0115]
[0116] in, Indicates candidate assets Identity verification scoring; Indicates fusion group Candidate assets are Or it can be mapped to candidate assets A collection of asset events; Indicates the first Modal credibility score for individual asset events; Indicates candidate assets The sum of the modal credibility scores for the corresponding asset events; Indicates fusion group The sum of modal credibility scores for all asset events within the fusion group. If the denominator is 0, it indicates that there are no credible asset events within the fusion group, and the conclusion pending review is output directly.
[0117] When candidate assets Identity verification score Not lower than the identity verification threshold ,and The difference between the score of the second-highest candidate asset and the score of the identity verification is not less than the discrimination threshold. At that time, candidate assets Assets identified for inventory. Identity verification threshold. The preferred value is between 0.6 and 0.75, which serves as the distinction threshold. A threshold of 0.1 to 0.2 is preferred. Using a discrimination threshold can prevent premature conclusions about the identity of multiple candidate assets when their scores are close.
[0118] When confirming the location, select locations with calculable spatial locations and modal confidence scores not lower than the minimum participation threshold. Asset events form location fusion set Minimum participation threshold The preferred value is 0.3 to 0.5. If the spatial location is a regional identifier, the regional center point is used first in the location calculation, while the regional boundaries are retained for the accounting determination. The merged spatial location is obtained as follows:
[0119]
[0120] in, Indicates the spatial location of the merged entity; This represents the set of asset events in the fusion group that have a computable spatial location and whose modal credibility score reaches the minimum participation threshold; Indicates the first Modal credibility score for individual asset events; Indicates the first The spatial location of an asset event. If If the number of asset events is less than two, weighted correction will not be performed, the spatial location corresponding to the highest modal confidence score will be retained as the spatial location to be confirmed, and a supplementary sampling task will be generated.
[0121] When determining the balance sheet, spatial location will be taken into account. Location of ledger registration Perform a regional consistency assessment. If the merged spatial location falls within the boundary of the region corresponding to the registered location in the ledger, or meets the following conditions... If the physical inventory and the accounting records match, then the physical inventory and the accounting records are considered to be in the same position; otherwise, the physical inventory and the accounting records are considered to be incompatible. Indicates candidate assets The registration location in the ledger Using the spatial region in S3 The corresponding spatial tolerance avoids using different spatial judgment boundaries for different steps.
[0122] Abnormal time series offsets that were not resolved were judged based on the compensated time consistency and the anomaly penalty value. For combinations of source identifiers containing abnormal time series offsets, the inter-modal time series offset estimate obtained from S4 was used. After compensating for the corresponding event time, recalculate the time consistency; if the recalculated time consistency is still lower than expected... Or, abnormal penalty values for related asset events. If the value is not lower than 0.5, it is determined that the abnormal timing offset has not been resolved. If it has not been resolved, a deterministic normal conclusion is not output; instead, an inventory conclusion with a confidence level is output.
[0123] The asset inventory conclusion is output according to the following rules: When a candidate asset meets the identity confirmation conditions, its integrated spatial location is consistent with the location registered in the ledger, and there is no unresolved abnormal temporal offset, the conclusion that the asset exists and the ledger is consistent is output; when a candidate asset does not meet the identity confirmation conditions, but there is a corresponding asset record in the ledger, the conclusion that it is suspected to be missing is output; when a candidate asset meets the identity confirmation conditions, but its integrated spatial location is inconsistent with the location registered in the ledger, the conclusion that it is suspected to be misplaced is output; when the identity confirmation scores of different candidate assets are close and the distinction threshold is not met, the conclusion that it is suspected to be duplicated or pending review is output; when a candidate asset exists on-site but there is no corresponding asset record in the ledger, the conclusion that it is suspected to be a newly added asset is output.
[0124] Based on the source of the anomaly, low-reliability asset events, and the results of the inventory and physical inventory assessment, a supplementary acquisition or review task is generated. If the source of the anomaly is a time-series offset, the task records the source identifier, the target time window, and the modality type that needs to be re-acquired; if the source of the anomaly is a spatial conflict, the task records the candidate assets, the conflicting spatial location, and the spatially related modalities that need to be re-acquired; if the source of the anomaly is an identity conflict, the task records the candidate assets, the conflicting candidate objects, and the identity-related modalities that need to be re-acquired; if the source of the anomaly is a ledger inconsistency, the task records the candidate assets, the merged spatial location, and the ledger registration location, and generates a ledger review task. Each supplementary acquisition or review task must at least include the candidate assets to be processed, the target spatial location, the source of the anomaly, the modality type that needs to be re-acquired, and the asset event number that triggered the task.
[0125] The results of supplementary sampling or verification are written back to the corresponding asset event and historical time series offset benchmark. If the supplementary sampling results confirm that an asset event is misidentified, the confidence parameter and modal confidence score of the corresponding asset event are reduced to 0, and the historical stability of the corresponding source identifier in subsequent windows is reduced. If the supplementary sampling results confirm that a source identifier has a stable time delay, the time difference obtained from the supplementary sampling is used as a new time series offset estimate, and the historical time series offset benchmark of the corresponding source identifier combination is updated according to the historical time series offset benchmark update rules in S4. If the verification results confirm that the ledger registration position is lagging, the confidence parameter of the corresponding ledger asset event is reduced, and the on-site fusion spatial location is given priority in the judgment of the ledger and physical inventory in subsequent inventory counts. If the supplementary sampling results confirm that the original inventory conclusion is correct, the confidence record of the corresponding asset event is retained, and the historical stability of the corresponding source identifier is included in the normal window.
[0126] It should also be noted that by fusing multimodal asset events according to modal credibility scores and generating supplementary collection or review tasks, the inventory output is expanded from a single identification result to a process of identity verification, location correction, inventory and physical verification, and closed-loop write-back. Park inventory not only needs to determine "whether an asset was identified," but also whether the identification result is consistent with the ledger registration location, asset status, and abnormal temporal offsets. This invention first calculates the identity verification score of candidate assets within the same conflict fusion range, using confirmation thresholds and differentiation thresholds to prevent premature identity verification when multiple candidate asset scores are close; then, it selects the spatial event with the lowest credibility for weighted correction to form a fused spatial location, avoiding low-credibility locations from participating in location averaging; subsequently, it performs regional consistency judgment between the fused spatial location and the ledger registration location, and outputs conclusions such as existence, misalignment, missing, duplicate identification, pending review, or suspected addition based on unresolved temporal offsets. For identity conflicts, spatial conflicts, or temporal anomalies that cannot be confirmed, this invention generates a supplementary sampling or verification task that includes the source of the anomaly, the target location, and the resampling mode, and writes the supplementary sampling results back to the event credibility, historical time series offset benchmark, and source stability record, so that subsequent inventory will not repeatedly accept the confirmed error source or expired ledger information.
[0127] One embodiment of the present invention provides an intelligent identification and automatic inventory system for park assets that integrates multimodal data, including an event association module, a spatiotemporal verification module, and a trusted inventory module.
[0128] The event association module is used to collect image, RFID, positioning, point cloud, and ledger data of park assets to generate multimodal asset events. Based on the candidate assets, event time, and spatial location of the multimodal asset events, it constructs cross-modal candidate association relationships between different modalities. The spatiotemporal verification module is used to input the cross-modal candidate association relationships into a trusted spatiotemporal alignment process to obtain the temporal consistency, spatial consistency, and causal consistency between asset events. Based on the temporal consistency, spatial consistency, and causal consistency, it estimates the temporal offset between modalities and identifies abnormal temporal offsets. The trusted inventory module is used to calculate the modal credibility score of each multimodal asset event by combining abnormal temporal offsets, quality parameters, and cross-modal consistency. The multimodal asset events are fused according to the modal credibility scores to output the asset inventory conclusion and generate supplementary collection or review tasks.
[0129] Reference Figure 2 This embodiment also provides a computer device applicable to the intelligent identification and automatic inventory method for park assets that integrates multimodal data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent identification and automatic inventory method for park assets that integrates multimodal data as proposed in the above embodiment.
[0130] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0131] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the method for intelligent identification and automatic inventory of park assets that integrates multimodal data, as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
Claims
1. A method for intelligent identification and automatic inventory of park assets integrating multimodal data, characterized in that, include: Collect images, RFID, location, point cloud, and ledger data of park assets to generate multimodal asset events; Based on the candidate assets, event time, and spatial location of multimodal asset events, construct cross-modal candidate association relationships between different modalities; By inputting cross-modal candidate associations into a trusted spatiotemporal alignment process, temporal consistency, spatial consistency, and causal consistency among asset events are obtained. Based on temporal consistency, spatial consistency, and causal consistency, estimate intermodal temporal offsets and identify anomalous temporal offsets; By combining anomalous time series offsets, quality parameters, and cross-modal consistency, modal credibility scores are calculated for each multimodal asset event; By integrating multimodal asset events based on modal credibility scores, the system outputs asset inventory conclusions and generates supplementary collection or review tasks.
2. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 1, characterized in that: The events for generating multimodal assets include Different modal data are processed by unifying the acquisition time, binding the source identifier, and mapping the spatial coordinates. Each piece of data that can be used for inventory assessment is converted into an asset event that includes event time, source identifier, candidate asset, spatial location, confidence parameter, and quality parameter; Candidate assets are formed from at least one of image recognition results, radio frequency identification results, or ledger records; Spatial location is formed by at least one of the following: positioning results, point cloud mapping results, data collection source deployment location, or ledger registration location; When a candidate asset is missing but a spatial location exists, the corresponding asset event is retained and added to the candidate association; When a spatial location is missing but a candidate asset exists, the spatial mapping result is supplemented based on the source identifier or the location registered in the ledger, and the supplemented result is marked as a low-confidence spatial location.
3. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 2, characterized in that: The construction of cross-modal candidate association relationships between different modalities includes... Within the preset inventory time window, select multimodal asset events with different source identifiers for matching; During matching, the consistency between candidate assets is compared first, and the association priority is increased when candidate assets correspond to the same asset record. When candidate assets cannot be directly confirmed, compare the proximity of events in time and the proximity of events in location. The proximity of the events is determined based on the difference between the time difference of the two asset events and the historical offset benchmark between the modes; The proximity of two assets is determined by their distance in a unified spatial coordinate system. When at least two of the candidate assets, event time and spatial location meet the preset association conditions, the corresponding cross-modal candidate association relationship is retained. When multiple cross-modal candidate associations compete for the same asset event, they are sequentially selected based on the degree of consistency of candidate assets, the degree of proximity of spatial location, and the degree of proximity of event time to form a candidate event group.
4. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 3, characterized in that: The obtained temporal consistency, spatial consistency, and causal consistency among asset events include Read the event time, source identifier, spatial location, and candidate assets of each multimodal asset event within the candidate event group; Based on the source identifier, the corresponding historical time series offset benchmark is invoked, and the deviation between the time difference between asset events in the candidate event group and the historical time series offset benchmark is compared to form time consistency. Based on the spatial distance under a unified spatial coordinate system, the degree of spatial deviation between asset events within the candidate event group is compared to form spatial consistency. Based on the historical movement records, spatial connectivity, and ledger change order of candidate assets, determine whether the asset events within the candidate event group conform to an explainable sequence of events and form causal consistency. When temporal consistency, spatial consistency and causal consistency all meet the preset conditions, the candidate event group is retained for time series offset estimation. When any consistency condition is not met, the corresponding deviation source is recorded and the corresponding deviation source is passed to the abnormal timing offset judgment.
5. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 4, characterized in that: The estimation of inter-modal temporal offsets and the determination of anomalous temporal offsets include... Based on the consistency results output by the trusted spatiotemporal alignment process, the event time difference in the candidate event group is statistically analyzed according to the source identifier and spatial region. Extract the time difference of events occurring within the sliding time window to form an estimate of the intermodal temporal offset; Compare the estimated intermodal temporal offset with the historical temporal offset baseline; When the estimated intermodal temporal offset continuously exceeds the historical allowable fluctuation range, it is determined that there is an abnormal temporal offset in the corresponding source identifier; When the estimated intermodal temporal offset changes abruptly within a short period of time, it is determined that there is an abnormal temporal offset in the corresponding candidate event group; When multiple spatial locations corresponding to the same candidate asset cannot be established simultaneously within the same time window, or when the order of events within a candidate event group is inconsistent with the historical movement record, it is determined that there is an abnormal temporal offset in the candidate event group. The results of abnormal time-series offset judgments are written to the corresponding asset events for use in modal credibility scoring.
6. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 5, characterized in that: The calculation of the modal credibility score for each multimodal asset event includes, For each multimodal asset event, read the corresponding confidence parameters, quality parameters, temporal consistency, spatial consistency, causal consistency, and abnormal time series offset judgment results; When the quality parameters of an asset event meet the preset conditions and form a stable cross-modal candidate association with asset events from different sources, the corresponding modal credibility score is improved. When an asset event deviates from the candidate event group in terms of time, space, or event sequence, the corresponding modal credibility score is reduced according to the degree of deviation. When an abnormal time offset is triggered by the corresponding source identifier, an abnormal penalty is applied to the corresponding modality confidence score; When there are multiple mutually supporting asset events for the same candidate asset, the asset event with higher cross-modal consistency is selected to participate in the fusion. When there are conflicting asset events for the same candidate asset, they are sorted in order of quality parameters, cross-modal consistency, abnormal time series offset judgment results, and historical offset stability, and the sorting results are written into the credibility record of the corresponding asset event.
7. The method for intelligent identification and automatic inventory of park assets integrating multimodal data as described in claim 6, characterized in that: The output of asset inventory conclusions and the generation of supplementary collection or review tasks include: Multimodal asset events belonging to the same candidate asset are merged according to modal credibility scores; When confirming assets, accumulate the modal credibility scores corresponding to the same candidate assets, and determine the candidate assets that meet the confirmation conditions as inventory assets; During location confirmation, candidate spatial locations are weighted and corrected according to modal credibility scores to obtain fused spatial locations; When assessing the physical assets against the accounting records, the spatial location will be combined with the location registered in the ledger to determine regional consistency, and the asset inventory conclusion will be formed based on the ledger status. When a candidate asset fails to meet the confirmation criteria, the fusion spatial location is inconsistent with the ledger registration location, or the abnormal temporal offset is not resolved, the output will be a suspected missing, suspected misalignment, suspected duplicate identification, or pending review conclusion. Based on the source of abnormal time series offset, low-reliability asset events, and the results of the account-to-physical judgment, a supplementary collection or review task is generated, and the results of the supplementary collection or review are written back to the corresponding asset events and historical time series offset benchmarks.
8. A smart identification and automatic inventory system for park assets integrating multimodal data, employing the smart identification and automatic inventory method for park assets integrating multimodal data as described in any one of claims 1 to 7, characterized in that: Includes an event correlation module, a spatiotemporal verification module, and a trusted inventory module; The event association module is used to collect images, radio frequency identification, positioning, point cloud and ledger data of park assets to generate multimodal asset events; based on the candidate assets, event time and spatial location of the multimodal asset events, it constructs cross-modal candidate association relationships between different modalities; The spatiotemporal verification module is used to input cross-modal candidate associations into the trusted spatiotemporal alignment process to obtain temporal consistency, spatial consistency, and causal consistency between asset events; based on temporal consistency, spatial consistency, and causal consistency, it estimates the temporal offset between modes and identifies abnormal temporal offsets. The trusted inventory module is used to calculate the modal trust score of each multimodal asset event by combining abnormal time series offsets, quality parameters, and cross-modal consistency. By integrating multimodal asset events based on modal credibility scores, the system outputs asset inventory conclusions and generates supplementary collection or review tasks.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for intelligent identification and automatic inventory of park assets that integrates multimodal data as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for intelligent identification and automatic inventory of park assets that integrates multimodal data as described in any one of claims 1 to 7.