Electronic warehouse receipt cargo authenticity verification method based on multi-mode AI
By using multi-modal AI technology to fuse multi-source data and make adaptive decisions, the problems of low efficiency, insufficient accuracy and high fraud risk in traditional verification methods are solved, and efficient and accurate verification of goods quantity and identification of potential anomalies are achieved in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JUJUN TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies suffer from inefficiency, error-proneness, and insufficient stability and accuracy in verifying the quantity and authenticity of goods. Especially in complex warehousing environments, a single data source lacks cross-validation capabilities and cannot effectively address fraud risks.
By employing a multimodal AI approach, the simultaneous acquisition and fusion of multi-view visual data, depth data, weight data, and RFID data, combined with 3D reconstruction and instance segmentation techniques, and dynamically adjusting processing parameters and weight allocation strategies, cross-validation of multi-source data and adaptive decision-making are achieved.
It maintains a highly stable and accurate counting capability in complex warehousing environments, effectively resists the failure of a single data source or environmental interference, proactively identifies potential anomalies, improves the accuracy of verification and anti-counterfeiting capabilities, and supports efficient operation around the clock.
Smart Images

Figure CN121883045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of blockchain and information security technology, and in particular to a method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI. Background Technology
[0002] In sectors such as supply chain finance, commodity trading, and modern smart warehousing, electronic warehouse receipts have gradually replaced traditional paper certificates, becoming the core carrier for the digital representation, ownership confirmation, and transfer of goods assets. Essentially, they map physical inventory to reliable digital records, thereby supporting advanced business models such as movable asset financing, warehouse receipt pledging, and online transactions. Therefore, ensuring the accurate correspondence between the information contained in electronic warehouse receipts, especially the quantity and authenticity of goods, and physical inventory, is a fundamental guarantee for maintaining the credit system of the entire supply chain.
[0003] Currently, existing technologies have some problems in verifying the quantity and authenticity of goods. First, traditional manual inventory methods are inefficient, error-prone, and difficult to audit. Second, although visible light-based two-dimensional visual recognition technology has partially automated the process, it is significantly affected by changes in lighting, obstruction from stacked goods, and viewing angle limitations in complex warehousing environments, resulting in insufficient recognition stability and accuracy. Third, existing technologies still rely on a single data source. For example, using only RFID or relying solely on weight sensing solutions lacks cross-validation capabilities and cannot effectively address fraud risks such as label forgery, goods replacement, or empty boxes being filled with goods. It also has inherent defects in judging the consistency between physical goods and information.
[0004] Therefore, there is an urgent need for a multimodal AI-based method for verifying the authenticity of goods in electronic warehouse receipts to solve the above problems. Summary of the Invention
[0005] The purpose of this application is to provide a method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI, including the following steps: Multi-source sensing data of the target storage unit is collected synchronously based on a multi-modal sensing array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; Based on the multi-view visual data and depth data, three-dimensional information is obtained, and the three-dimensional information is used to perform three-dimensional reconstruction and instance segmentation of the cargo stack to obtain a point cloud-based cargo segmentation result and a first quantity estimate. Based on the multi-view visual data, depth data, weight data, and RFID data, the final verification quantity is obtained through dynamic weighted fusion. Obtain multimodal environment state information, and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information; The final verified quantity is compared with the quantity declared in the electronic warehouse receipt, and the cargo status is judged and the result is output based on the consistency between the multi-source sensing data.
[0006] Furthermore, the step of synchronously acquiring multi-source sensing data of the target storage unit based on a multi-modal sensing array includes: Deploy a multimodal sensor array that includes multi-view vision sensors, depth sensors, weight sensors and RFID read / write units, and perform spatial and temporal synchronous calibration; Based on the verification command, the multi-view vision sensor, depth sensor, weight sensor and RFID reading and writing unit are simultaneously triggered to collect the multi-view vision data, depth data, weight data and radio frequency identification data at the same time. The multi-view visual data, depth data, weight data, and RFID data are preprocessed.
[0007] Furthermore, the step of obtaining three-dimensional information based on the multi-view visual data and depth data, and performing three-dimensional reconstruction and instance segmentation of the cargo stack using the three-dimensional information to obtain a point cloud-based cargo segmentation result and a first quantity estimate includes: Based on the multi-view visual data and the depth data, a dense three-dimensional point cloud of the cargo pile is generated through stereo matching and coordinate transformation. The dense 3D point cloud is input into a pre-trained point cloud instance segmentation network, which assigns an instance identifier to each point in the point cloud. Based on the instance identifier, points belonging to the same independent physical cargo are aggregated into a single cargo point cloud cluster; The number of different cargo point cloud clusters is counted and used as the first quantity estimate. The spatial location information of each cargo point cloud cluster is then output.
[0008] Furthermore, the step of obtaining the final verification quantity based on the multi-view visual data, depth data, weight data, and RFID data through dynamic weighted fusion includes: Target detection is performed on the multi-view visual data, the number of detection boxes is counted, and after multi-view fusion and deduplication, the estimated number of visual boxes is obtained. Obtain stable weight readings from the weight data, and calculate the estimated weight quantity by dividing the stable weight readings by the pre-stored standard unit weight of the goods. The number of valid tags in the RFID data is counted to obtain an estimated number of RFID tags. The real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the RFID quantity estimate are evaluated respectively. Based on the real-time confidence level, dynamic fusion weights are assigned to the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the RFID quantity estimate. The quantity estimates are then weighted, summed, and rounded according to the dynamic fusion weights, and the output is the final verification quantity.
[0009] Furthermore, the step of acquiring multimodal environment state information and adaptively adjusting the processing parameters of instance segmentation and the weight allocation strategy of dynamic weighted fusion based on the multimodal environment state information includes: The multimodal environmental state information is acquired in real time. The state information includes at least ambient light intensity, point cloud clustering evaluation index reflecting the degree of cargo stacking and occlusion, and quality assessment index of each sensor data. Based on the point cloud clustering evaluation index, the operating mode of the point cloud instance segmentation network is dynamically selected. When the index shows complex stacking, the fine segmentation mode is enabled; otherwise, the fast segmentation mode is enabled. Based on the ambient light intensity and visual image quality evaluation indicators, adjust the weight allocation of the estimated visual quantity in the fusion process; Based on the point cloud clustering evaluation index and the depth data quality, adjust the weight allocation of the first quantity estimate in the fusion; Based on the stability index of the weight data and the reading success rate index of the RFID data, the fusion weights of the weight-derived quantity estimate and the RFID quantity estimate are adjusted respectively.
[0010] Furthermore, the step of comparing the final verified quantity with the quantity declared in the electronic warehouse receipt, and determining the cargo status and outputting the result based on the consistency between the multi-source sensing data, includes: Calculate the difference between the final verified quantity and the quantity declared in the electronic warehouse receipt; Analyze the consistency among the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate to identify whether there are contradictory evidence pairs that exceed a preset threshold; If the quantity difference is within the allowable range and all evidence pairs are consistent, the goods inspection is deemed to have passed. If the quantity difference exceeds the allowable range, it is determined to be a quantity abnormality alarm; If the difference in quantity is within the allowable range but there is a contradiction in the evidence, it is judged as a potential risk warning, and the risk type is identified according to the contradiction pattern. Generate a structured verification report that includes the number of verifications, the value of each piece of evidence, the status determination, and the risk type, and generate digital fingerprints of key process data for blockchain storage.
[0011] Furthermore, this invention also discloses an electronic warehouse receipt goods authenticity verification system based on multimodal AI, including: The acquisition module is used to synchronously acquire multi-source sensing data of the target storage unit based on a multi-modal sensor array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; The acquisition module is used to acquire three-dimensional information based on the multi-view visual data and depth data, and to perform three-dimensional reconstruction and instance segmentation of the cargo stack through the three-dimensional information to obtain cargo segmentation results and a first quantity estimate based on point cloud. The fusion module is used to obtain the final verification quantity based on the multi-view visual data, depth data, weight data and radio frequency identification data through dynamic weighted fusion. An adjustment module is used to acquire multimodal environment state information and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information. The output module is used to compare the final verified quantity with the quantity declared in the electronic warehouse receipt, and to determine the status of the goods and output the results based on the consistency between the multi-source sensing data.
[0012] Furthermore, the fusion module includes: The detection unit is used to perform target detection on the multi-view visual data, count the number of detection boxes, and obtain the visual quantity estimate after multi-view fusion and deduplication. The calculation unit is used to obtain the stable weight reading in the weight data and calculate the estimated weight by dividing the stable weight reading by the pre-stored standard unit weight of the goods. The statistics unit is used to count the number of valid tags in the RFID data and obtain an estimated value of the number of RFID tags. An evaluation unit is used to evaluate the real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate, respectively. The output unit is used to assign dynamic fusion weights to the first quantity estimate, visual quantity estimate, weight-derived quantity estimate and radio frequency identification quantity estimate based on the real-time confidence level, and to perform weighted summation and rounding on each quantity estimate according to the dynamic fusion weights, and output the final verification quantity.
[0013] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI.
[0014] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI.
[0015] The beneficial effects of this application are as follows: Firstly, this application introduces three-dimensional reconstruction and instance segmentation technology to directly analyze the stacking structure of goods from the physical space dimension, fundamentally overcoming the recognition bottleneck of traditional two-dimensional vision in occluded scenarios. By combining dynamic fusion and cross-validation of multimodal data such as vision, weight, and radio frequency, it can effectively resist the failure of a single data source or environmental interference, ensuring that it can still maintain a highly stable and accurate counting capability in complex real-world scenarios.
[0016] Secondly, this application not only verifies the quantity but also proactively identifies potential anomalies by analyzing the consistency logic between multi-source data. For example, when the RFID reading quantity and visual counting are significantly different but the weight data matches, the system can issue an early warning of the risk of goods replacement or tag anomalies, achieving in-depth prevention against complex fraud methods and providing a higher level of security for supply chain finance.
[0017] Third, this application dynamically adjusts core algorithm parameters and fusion strategies by real-time sensing of environmental conditions and data quality, thereby automatically adapting to different warehouse conditions and cargo types, reducing reliance on manual optimization. The overall solution achieves full-process automation from data collection and analysis to decision output, supporting efficient operation around the clock and improving warehouse management efficiency and operational automation levels. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a method flow proposed in an embodiment of this application.
[0019] Figure 2 This is a schematic diagram of the system structure proposed in one embodiment of this application.
[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0022] like Figure 1 As shown, this application provides a method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI, including the following steps: S1, synchronously collect multi-source sensing data of the target storage unit based on a multi-modal sensing array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; S2, obtain three-dimensional information based on the multi-view visual data and depth data, and perform three-dimensional reconstruction and instance segmentation of the cargo stack using the three-dimensional information to obtain cargo segmentation results and a first quantity estimate based on point cloud. S3, based on the multi-view visual data, depth data, weight data and radio frequency identification data, and obtain the final verification quantity through dynamic weighted fusion; S4, obtain multimodal environment state information, and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information; S5, compare the final verified quantity with the quantity declared in the electronic warehouse receipt, and determine the status of the goods and output the results based on the consistency between the multi-source sensing data.
[0023] As described in steps S1-S5 above, in supply chain finance and movable asset pledge supervision scenarios, electronic warehouse receipts, as digital certificates of ownership of goods, are directly related to financial risk control security due to their consistency with physical inventory. In actual warehousing environments, factors such as obstruction caused by stacked goods, random variations in lighting conditions, fraudulent activities like label forgery or goods substitution, and the limitations of a single data source lacking cross-validation make it difficult for traditional verification methods to balance accuracy and anti-counterfeiting capabilities. Therefore, it is necessary to address core issues such as verification bias caused by data monoliths and weak anti-counterfeiting measures due to the lack of effective cross-validation, ensuring the authenticity and reliability of verification results and meeting the needs of automated supervision.
[0024] Traditional verification methods suffer from low efficiency and susceptibility to human error and fraud. Two-dimensional visual recognition relies on planar information, which cannot overcome the limitations of stacked goods and is susceptible to occlusion. Changes in lighting directly cause significant fluctuations in recognition accuracy. Relying solely on RFID or weight data lacks cross-validation capabilities and is ill-suited to combat fraudulent activities such as empty-box repackaging and label forgery. This invention specifically employs multimodal data fusion and three-dimensional perception technology, combined with an adaptive adjustment strategy. It overcomes the limitations of a single data source through multi-source data cross-validation, utilizes three-dimensional spatial information to address occlusion issues, and dynamically adjusts the strategy to adapt to complex environments, comprehensively improving the accuracy, robustness, and anti-counterfeiting capabilities of verification.
[0025] This application introduces 3D reconstruction and instance segmentation technology to directly analyze the stacking structure of goods from the physical space dimension, fundamentally overcoming the recognition bottleneck of traditional 2D vision in occluded scenarios. By combining dynamic fusion and cross-validation of multimodal data such as vision, weight, and radio frequency, it can effectively resist the failure of a single data source or environmental interference, ensuring that it can still maintain a highly stable and accurate counting capability in complex real-world scenarios.
[0026] The core logic and principle of this method are as follows: Addressing the problems in traditional warehouse cargo verification, such as the susceptibility of single data sources to environmental interference, the inability of 2D vision to overcome stacking occlusion, and the difficulty of adapting fixed strategies to dynamic scenarios, this method focuses on multimodal data collaborative perception and adaptive intelligent decision-making. First, a uniformly calibrated multimodal sensor array is deployed to simultaneously collect multi-view visual data, depth data, weight data, and RFID data that comprehensively reflect the physical state of the cargo. This ensures the spatiotemporal consistency and comprehensiveness of the multi-source data, providing high-quality input for subsequent analysis. Then, 3D reconstruction is performed using the multi-view visual and depth data to generate a dense 3D point cloud that completely restores the spatial morphology of the cargo stack. A pre-trained point cloud instance segmentation network is used to accurately separate and count individual cargo, obtaining a first quantity estimate based on 3D spatial information. This fundamentally solves the quantity estimation bias problem caused by 2D visual occlusion. Subsequently, visual quantity estimates, weight-derived quantity estimates, and RFID quantity estimates are obtained based on the multi-source data. These estimates are then evaluated... The system dynamically assigns fusion weights to the real-time confidence levels of each estimate. Weighted summation and rounding yield the final verification quantity, leveraging the advantages of multiple sources. This avoids the limitations of single data and the rigidity of fixed weights. Simultaneously, it acquires multimodal environmental status information in real-time, including ambient light intensity, point cloud clustering evaluation indicators, and data quality assessment indicators from various sensors. The system dynamically adjusts the instance segmentation operation mode and processing parameters, as well as the weight allocation strategy for dynamic weighted fusion, enabling the system to adaptively match changes in the warehousing environment and the stacking status of goods, maintaining consistently high efficiency and accuracy. Finally, the final verification quantity is compared with the quantity declared in the electronic warehouse receipt. Consistency analysis between multi-source quantity estimates identifies contradictory evidence pairs, classifies goods verification as passed, issues quantity anomaly alarms, or trigger potential risk alarms, and generates a structured verification report. Key process data is used to generate digital fingerprints for blockchain storage, achieving comprehensive verification of the quantity and authenticity of goods while ensuring the traceability of the verification process and the credibility of the results. This provides reliable technical support for the supervision of movable property pledges in supply chain finance.
[0027] In one embodiment, the step of synchronously acquiring multi-source sensing data of the target storage unit based on a multimodal sensing array includes: S11 deploys a multimodal sensor array that includes multi-view vision sensors, depth sensors, weight sensors and RFID read / write units, and performs spatial and temporal synchronous calibration. S12, based on the verification command, the multi-view vision sensor, depth sensor, weight sensor and RFID reading and writing unit are simultaneously triggered to collect the multi-view vision data, depth data, weight data and radio frequency identification data at the same time. S13, preprocess the multi-view visual data, depth data, weight data and radio frequency identification data.
[0028] As described in steps S11-S13 above, through the deployment and calibration of the multimodal sensor array, synchronous acquisition and targeted data preprocessing, spatiotemporally consistent and reliable multi-source sensing data are obtained, providing high-quality input for subsequent 3D reconstruction, instance segmentation and dynamic fusion decision-making, and ensuring the accuracy and stability of the verification of the authenticity of goods in electronic warehouse receipts.
[0029] In warehouse cargo verification scenarios, the spatiotemporal consistency and data quality of multi-source sensing data are fundamental to subsequent analysis. Differences in the deployment locations of different sensors can lead to inconsistent spatial coordinates, and asynchronous data acquisition times can cause mismatches between the data and the actual state of the goods. Furthermore, issues in the raw data such as lighting interference, image distortion, depth holes, weight fluctuations, and label redundancy can directly amplify errors in subsequent processing, resulting in inaccurate quantity estimates and unreliable verification results. Therefore, it is necessary to address the problems of poor sensor coordination, spatiotemporal misalignment of data, and low quality of raw data to ensure the usability and relevance of multi-source data, providing reliable support for subsequent processes.
[0030] Traditional data acquisition methods have significant limitations. A single sensor cannot provide comprehensive sensing information, and asynchronous multi-sensor acquisition lacks unified calibration, leading to spatiotemporal data misalignment. Furthermore, the lack of targeted preprocessing after acquisition results in the direct use of noisy and distorted raw data for analysis, resulting in low accuracy and poor robustness in subsequent quantity estimation. By adopting a targeted approach that involves the unified deployment and calibration of multimodal sensor arrays, synchronous triggering of acquisition, and classification preprocessing, the spatiotemporal consistency problem of data is addressed at its source, systematically improving data quality and laying the foundation for subsequent high-precision verification.
[0031] A multimodal sensor array, comprising multi-view vision sensors, depth sensors, weight sensors, and RFID reader / writer units, is deployed and spatially and temporally synchronized for calibration. Based on the pallet size of the target storage unit, four global shutter cameras are arranged in a rectangular pattern on the gantry above the pallet to ensure comprehensive coverage of the stacked goods. The depth sensor is installed at the center of the vision sensor, maintaining a vertical distance of 50cm. The weight sensor is embedded in four support points at the bottom of the pallet. The antenna array of the RFID reader / writer unit is evenly deployed around the perimeter of the pallet area, covering the entire stacked goods area. Spatial synchronization calibration uses a checkerboard calibration board to acquire calibration images from different positions and angles. A perspective transformation algorithm is used to calculate the transformation matrix between each sensor and the world coordinate system, aligning the coordinate systems of all sensors to the world coordinate system. This ensures accurate correspondence of the coordinates of the same physical point in the visual and depth data. Temporal synchronization calibration uses pulse trigger signals with a set trigger frequency of 30Hz. All sensors are connected to the same synchronization controller, and upon receiving the same pulse signal, they simultaneously begin acquisition, eliminating data misalignment caused by time delay. It enables the coordinated operation of sensors. For example, after spatial calibration, the edge contour of the cargo captured by the vision sensor can be accurately matched with the corresponding depth value obtained by the depth sensor, avoiding spatial misalignment during 3D reconstruction. Time synchronization ensures that the collected data reflects the cargo status at the same time, providing a spatiotemporal basis for multi-source data fusion.
[0032] Based on verification commands, multiple sensors are synchronously triggered to collect multi-source sensing data at the same time. Verification commands are issued by the warehouse management system. When a pallet is transported into the preset verification area by a forklift and triggers the position sensor, the system immediately generates a verification command and sends it to the synchronization controller. The synchronization controller then sends trigger signals to all sensors. The vision sensor receives the signal and acquires an RGB image with a resolution of 1920×1080. The depth sensor synchronously acquires depth information, with a measurement range of 0.5-5m. The weight sensor records the load pressure data in real time at a sampling frequency of 10Hz. The RFID reader / writer unit initiates tag scanning, with a reading range covering the entire pallet area. Synchronous triggering ensures that multi-source data completely corresponds to the current state of the goods, avoiding data mismatches caused by goods movement or state changes due to time differences in acquisition. For example, it avoids the situation where the goods are in a certain position when visual data is acquired, but have slightly moved when weight data is acquired, leading to a loss of correlation and ensuring data relevance and timeliness.
[0033] Targeted preprocessing was performed on the multi-source sensing data. Visual data underwent illumination normalization using a gamma correction algorithm, adjusting image brightness to a uniform range of 80-120 cd / m². Distortion parameters were solved using the Zhang Zhengyou calibration method to correct radial and tangential distortion of the lens, improving the geometric accuracy of the image. Depth data was filtered using a 5×5 kernel with a standard deviation of 1.2 Gaussian filter to remove random noise. Neighborhood interpolation was used to fill data holes, calculating the mean based on the effective depth values within a 3×3 area around the hole to fill the hole area and ensure the integrity of the depth data. Weight data was filtered using a moving average filter with 10 sampling points to smooth out instantaneous fluctuations during transport and output a stable actual total weight reading. RFID tag data was processed by traversing the tag list, comparing unique tag identifiers, deleting duplicate identifiers, and retaining only unique valid tags. Preprocessing directly addresses the quality issues of the raw data. Illumination normalization ensures that visual data can clearly present the outline of goods in both strong and low light environments. Image distortion correction guarantees the accuracy of goods size measurement. Noise filtering and hole filling of depth data prevent missing parts of the goods surface during 3D reconstruction. Smoothing of weight data ensures the stability of weight derivation. Deduplication of RFID data avoids duplicate counting. These processes directly improve the accuracy of subsequent quantity estimation, provide reliable data support for dynamic fusion decision-making, and indirectly ensure the accuracy of the final verification results.
[0034] In one embodiment, the step of obtaining three-dimensional information based on the multi-view visual data and depth data, performing three-dimensional reconstruction and instance segmentation of the cargo stack using the three-dimensional information, and obtaining a point cloud-based cargo segmentation result and a first quantity estimate includes: S21, Based on the multi-view visual data and the depth data, a dense three-dimensional point cloud of the cargo pile is generated through stereo matching and coordinate transformation; S22, the dense 3D point cloud is input into a pre-trained point cloud instance segmentation network, wherein the segmentation network assigns an instance identifier to each point in the point cloud; S23, based on the instance identifier, aggregate points belonging to the same independent physical cargo into a single cargo point cloud cluster; S24, count the number of different cargo point cloud clusters, use them as the first quantity estimate, and output the spatial location information of each cargo point cloud cluster.
[0035] As described in steps S21-S24 above, the three-dimensional reconstruction and instance segmentation of the cargo stack are completed through the fusion processing of multi-view visual data and depth data. The spatial features of individual cargo are accurately extracted and the quantity is counted to obtain a reliable first quantity estimate, which provides a core three-dimensional perception basis for subsequent multimodal data fusion and breaks through the recognition limitations of two-dimensional vision in stacked occlusion scenarios.
[0036] In warehouse cargo verification scenarios, goods are often stored in stacks. The overlapping surfaces of adjacent goods make it impossible for two-dimensional vision to distinguish individual items, and it is difficult to accurately determine the actual quantity of goods based solely on planar images. Three-dimensional information can fully reflect the spatial shape and positional relationships of goods. 3D reconstruction can restore the three-dimensional structure of the goods stack, and instance segmentation can separate each individual item from the three-dimensional structure. Therefore, this step is necessary to upgrade two-dimensional perception to three-dimensional perception, solve the problem of quantity estimation bias caused by stacking occlusion, and provide a high-precision basic quantity reference for overall verification.
[0037] In traditional technologies, 2D visual recognition can only judge based on planar texture and contour. When faced with stacked and occluded goods, it is easy to misclassify multiple overlapping goods as a single item or miss occluded goods. Some 3D reconstruction methods generate sparse point clouds that cannot fully reflect the details of the goods. Simple clustering and segmentation algorithms rely only on spatial distance and tend to merge closely adjacent goods into one instance, resulting in large errors in quantity estimation. By using multi-view fusion to generate dense 3D point clouds and combining them with a pre-trained high-precision point cloud instance segmentation network, accurate separation of independent goods is achieved through feature extraction and instance label assignment. This completely solves the problems of occlusion and confusion of adjacent goods at the 3D spatial level, improving the accuracy of quantity estimation.
[0038] Based on preprocessed multi-view visual and depth data, a dense 3D point cloud of the cargo stack is generated through stereo matching and coordinate transformation. Stereo matching employs a semi-global matching algorithm, which calculates the similarity of corresponding pixels in images from different viewpoints and combines this with a global energy optimization strategy to accurately solve pixel-level disparity, ensuring the accuracy and continuity of disparity calculation. Coordinate transformation utilizes previously completed spatial synchronization calibration parameters to transform the 2D pixel coordinates and depth data from each viewpoint to a unified world coordinate system. Coordinate fusion then stitches and integrates the point cloud data from multiple viewpoints to generate a dense 3D point cloud covering all surfaces of the cargo stack. The point cloud density is set to no less than 5 points per cubic centimeter to ensure the complete representation of details such as the edges and recesses of the cargo. This process fills in blind spots in a single viewpoint by complementing multi-view data; for example, areas where the bottom of the cargo is obscured by pallets can be supplemented using depth data from side views, ensuring that the point cloud can completely reconstruct the 3D structure of the cargo stack and providing comprehensive spatial information support for subsequent segmentation.
[0039] A dense 3D point cloud is input into a pre-trained point cloud instance segmentation network, assigning an instance label to each point in the point cloud. The point cloud instance segmentation network adopts a PointNet++ variant architecture, which includes a sampling layer, a grouping layer, and a feature extraction layer. The sampling layer selects key sampling points in the point cloud using the farthest point sampling algorithm, reducing computation while preserving core spatial features. The grouping layer constructs a spherical neighborhood with a fixed radius of 0.1 meters centered on each sampling point, aggregating points within the neighborhood into a local feature set. The feature extraction layer extracts local and global features using three 1D convolutional kernels with kernel sizes of 1×16, 1×32, and 1×64, respectively. Activation functions are used to enhance feature representation capabilities, ultimately outputting the instance classification probability for each point. During network training, a labeled dataset containing different stacking methods and different cargo sizes is used. Each independent cargo is assigned a unique instance label in the labeled data. Training stops when the loss function value is below 0.01 to ensure the network has strong generalization ability. The assignment of instance identifiers is based on the spatial continuity, size characteristics and surface texture consistency of the goods. The preset minimum volume threshold for goods is 0.01 cubic meters. Point cloud clusters below this threshold are judged as noise and are not assigned instance identifiers to avoid interference factors such as dust and packaging debris from affecting the results.
[0040] Based on instance identifiers, points belonging to the same independent physical cargo are aggregated into a single cargo point cloud cluster. A connected component analysis algorithm is employed to traverse all points in the point cloud, grouping points with the same instance identifier and a spatial distance of less than 0.05 meters into the same cluster. Spatial distance constraints prevent the aggregation of points from different cargoes due to incorrect identifiers, while ensuring that discrete points of the same cargo can be completely aggregated. The aggregation process strictly adheres to the uniqueness of instance identifiers; each instance identifier corresponds to a unique cargo point cloud cluster. For example, two closely adjacent cube-shaped cargoes, even if their surfaces are almost touching, will still be separated into two independent point cloud clusters due to their different instance identifiers, completely resolving the problem of confusion between adjacent cargoes.
[0041] The system counts the number of different cargo point cloud clusters, using this as the first quantity estimate, and outputs the spatial location information of each cargo point cloud cluster. The spatial location information is obtained by calculating the coordinates of the center point of each point cloud cluster and the boundary of its circumscribed cuboid. The center point coordinates are the average of the three-dimensional coordinates of all points within the cluster, and the boundary of the circumscribed cuboid is determined by calculating the maximum and minimum values of the points within the cluster on the x, y, and z axes. When counting point cloud clusters, noisy clusters below a minimum volume threshold are automatically filtered out to ensure that the first quantity estimate only reflects the actual cargo quantity. For example, 10 stacked standard boxes, after segmentation and aggregation, will accurately result in 10 point cloud clusters, with a first quantity estimate of 10. The output spatial location information can provide a reference for subsequent multimodal data fusion. For instance, when there is a difference between the visual quantity estimate and the first quantity estimate, the spatial location information can be used to determine if there are visually missed occlusion areas, further improving the reliability of the overall verification.
[0042] In one embodiment, the step of obtaining the final verification count based on the multi-view visual data, depth data, weight data, and RFID data through dynamic weighted fusion includes: Target detection is performed on the multi-view visual data, the number of detection boxes is counted, and after multi-view fusion and deduplication, the estimated number of visual boxes is obtained. S31, obtain the stable weight reading from the weight data, and calculate the estimated weight by dividing the stable weight reading by the pre-stored standard unit weight of the goods. S32, count the number of valid tags in the RFID data to obtain an estimated number of RFID tags; S33, evaluate the real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate, respectively; S34. Based on the real-time confidence level, assign dynamic fusion weights to the first quantity estimate, visual quantity estimate, weight-derived quantity estimate, and radio frequency identification quantity estimate, and perform weighted summation and rounding on each quantity estimate according to the dynamic fusion weights, and output the final verification quantity.
[0043] As described in steps S31-S34 above, by estimating the quantity of multimodal sensing data separately, and dynamically allocating fusion weights based on the real-time confidence of each estimate, the final verified quantity is obtained through weighted summation and rounding. This fully leverages the complementary advantages of each data source, improves the reliability and robustness of cargo quantity estimation in complex warehousing environments, and provides core data support for cargo authenticity determination.
[0044] In electronic warehouse receipt cargo verification, quantity estimation from a single data source is easily affected by environmental and inherent limitations. Visual data is susceptible to interference from lighting and occlusion, weight data may deviate due to equipment fluctuations, and RFID data is easily affected by tag obstruction or malfunction. The reliability of each data source changes dynamically with the scenario. Relying solely on a single estimate or using fixed-weight fusion cannot adapt to changing scenarios, leading to significant quantity estimation deviations and affecting the accuracy of verification results. Therefore, it is necessary to integrate quantity estimation results from multiple data sources, dynamically adjusting weights based on the real-time performance of each data source to achieve complementary advantages and overcome the limitations of single-data estimation and the rigidity of fixed weights.
[0045] Traditional fusion methods have significant drawbacks. Fixed-weight fusion fails to consider the real-time quality differences among data sources. For example, in poor lighting conditions, the reliability of visual data decreases while the weights remain unchanged, leading to interference from low-quality data in the fusion results. Some methods directly average the estimates from multiple sources, failing to highlight the dominant role of high-quality data and reducing fusion accuracy. This solution employs a method combining source-specific quantity estimation, real-time confidence assessment, and dynamic weighted fusion. First, quantity estimates from each data source are obtained independently. Then, the reliability of each estimate is quantified, and weights are assigned according to reliability to achieve fusion. This ensures that high-quality data dominates the results, minimizes the impact of low-quality data, and improves the accuracy and adaptability of quantity estimation.
[0046] Object detection was performed on the preprocessed multi-view visual data. The number of detection boxes was counted, and the visual quantity estimate was obtained after multi-view fusion and deduplication. The YOLO model was used for object detection, consisting of 8 convolutional layers, 3 pooling layers, and 2 fully connected layers. The convolutional kernel sizes were 3×3 and 1×1, respectively. Max pooling was used in the pooling layers with a stride of 2. The fully connected layers output the coordinates of the detection boxes and their class confidence scores. A confidence threshold of 0.7 was set; detection boxes below this threshold were directly filtered out. Multi-view fusion and deduplication employed the Cross-Union Ratio (CUI) threshold method. The CUI of detection boxes under different views was calculated. When the CUI was greater than 0.5, the boxes were considered to be the same item, and only the detection box with the highest confidence score was retained. Finally, the number of remaining detection boxes was counted as the visual quantity estimate. For example, if 10, 10, 9, and 10 items are detected from four different viewpoints, and after deduplication, 10 detection frames are retained, the visual quantity estimate is 10. This process reduces missed detections and duplicate counts caused by occlusion from a single viewpoint through multi-view complementarity and deduplication, thereby improving the accuracy of visual estimation.
[0047] The method acquires stable weight readings from the weight data and calculates the estimated quantity by dividing these stable weight readings by the pre-stored standard unit weight of the goods. The stable weight readings are derived from pre-processed weight data, filtered using a moving average to eliminate instantaneous fluctuations and ensure numerical stability. The standard unit weight of the goods is the rated weight at the time of manufacture and is pre-stored in the system database, associated with each type of goods. During verification, the corresponding standard unit weight is retrieved based on the type of goods in the electronic warehouse receipt. For example, if the stable weight reading is 100kg and the standard unit weight is 10kg, the estimated quantity is 10. This method utilizes the physical relationship between weight and quantity, is unaffected by visual obstruction or lighting conditions, provides an objective reference for quantity estimation, and overcomes the limitations of visual and RFID data.
[0048] The number of valid RFID tags in the data is counted to obtain an estimated number of RFID tags. The number of valid tags is based on the pre-processed RFID data. In the pre-processing stage, redundant tags are deduplicated, retaining only unique tags. Tag validity is also verified through tag authentication to eliminate interference from counterfeit tags. The total number of valid tags is the estimated number of RFID tags. For example, if 10 valid tags remain after deduplication and verification, the estimated number of RFID tags is 10. Utilizing the uniqueness of RFID tags enables rapid counting of individual goods, providing accurate results when there are no obstructions and the tags are intact.
[0049] The real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the RFID quantity estimate are evaluated respectively. The confidence level of the first quantity estimate is calculated based on point cloud quality. A confidence level of 0.9 is set when the point cloud completeness is above 90%, and decreases by 0.1 for every 10% decrease in completeness. Point cloud completeness is determined by a combination of the percentage of valid points and the hole rate. The confidence level of the visual quantity estimate is positively correlated with image clarity. A confidence level of 0.8 is set when the image signal-to-noise ratio (SNR) is above 30 dB, and decreases by 0.1 for every 5 dB decrease in SNR. Image clarity is evaluated by calculating contrast and entropy values using the gray-level co-occurrence matrix. The confidence level of the weight-derived quantity estimate is negatively correlated with weight data stability. A confidence level of 0.8 is set when the weight fluctuation coefficient is below 5%, and decreases by 0.1 for every 5% increase in fluctuation coefficient. The weight fluctuation coefficient is the percentage difference between the values before and after the moving average filtering. The confidence level of the RFID quantity estimate is positively correlated with the tag reading success rate. A confidence level of 0.7 is set when the reading success rate is above 95%, and decreases by 0.1 for every 5% decrease in success rate. The reading success rate is the ratio of the number of valid tags to the number of tags declared in the electronic warehouse receipt. Confidence assessment provides an objective basis for weight allocation by quantifying the real-time performance of each data source.
[0050] Dynamic fusion weights are assigned to the four quantity estimates based on real-time confidence levels. The weighted summation and rounding of each estimate are then performed to output the final verified quantity. Weight allocation uses normalization; the weight of each estimate is equal to its confidence level divided by the sum of the four confidence levels, ensuring the total weights equal to 1. The weighted summation is then rounded to the nearest integer to obtain the final verified quantity. For example, if the first quantity estimate is 10 with a confidence level of 0.9, the visual quantity estimate is 10 with a confidence level of 0.8, the weighted derivation quantity estimate is 10 with a confidence level of 0.8, and the RFID quantity estimate is 10 with a confidence level of 0.7, the calculated weights are 0.28, 0.25, 0.25, and 0.22 respectively. The weighted summation result is 10, and the final verified quantity is 10. If poor lighting causes the visual confidence score to drop to 0.3, while other confidence scores remain unchanged, the weights are adjusted to 0.33, 0.11, 0.29, and 0.27. In this case, the primary quantity and weighted data dominate the fusion result, avoiding interference from low-quality visual data and ensuring the accuracy of the final verified quantity. This dynamic fusion strategy can adapt to changes in the quality of various data sources in real time, fully leveraging the dominant role of high-quality data and the supplementary role of low-quality data, thereby improving the robustness and accuracy of quantity estimation.
[0051] In one embodiment, the step of acquiring multimodal environment state information and adaptively adjusting the processing parameters of instance segmentation and the weight allocation strategy of dynamic weighted fusion based on the multimodal environment state information includes: S41, acquire the multimodal environment status information in real time. The status information includes at least ambient light intensity, point cloud clustering evaluation index reflecting the degree of cargo stacking and occlusion, and quality assessment index of each sensor data. S42, Based on the point cloud clustering evaluation index, dynamically select the operating mode of the point cloud instance segmentation network. When the index shows complex stacking, enable the fine segmentation mode; otherwise, enable the fast segmentation mode. S43, adjust the weight allocation of the estimated visual quantity in the fusion according to the ambient light intensity and visual image quality evaluation index; S44, adjust the weight allocation of the first quantity estimate in the fusion according to the point cloud clustering evaluation index and the depth data quality; S45, based on the stability index of the weight data and the reading success rate index of the RFID data, adjust the fusion weight of the weight-derived quantity estimate and the RFID quantity estimate respectively.
[0052] As described in steps S41-S45 above, by collecting multimodal environment status information in real time, dynamically adjusting the operation mode and processing parameters of point cloud instance segmentation, as well as the dynamic fusion weight of multimodal quantity estimates, the system can adaptively match changes in the warehousing environment and the stacking status of goods, ensuring that it can maintain high segmentation accuracy and reliable fusion effect in different scenarios, and further improving the overall robustness and adaptability of electronic warehouse receipt goods authenticity verification.
[0053] Warehouse environments exhibit significant dynamic changes. Ambient light intensity fluctuates with time, weather, and the status of warehouse lighting equipment. The stacking method and degree of occlusion of goods vary depending on storage needs, and the data quality from various sensors also changes dynamically. If the processing parameters for point cloud instance segmentation and the multimodal fusion weights remain fixed, they will be unable to adapt to these changes: a fixed, fast segmentation mode will lead to insufficient segmentation accuracy when stacking is complex; a fixed high visual weight will introduce low-quality data interference in poor lighting conditions, ultimately causing quantity estimation errors and unreliable verification results. Therefore, it is necessary to perceive the environment and data status in real time, adjust the system processing strategy accordingly, and solve the problem that fixed parameters and weights cannot adapt to dynamic changes in the scenario, ensuring the system continues to operate stably in complex and ever-changing warehouse environments.
[0054] Traditional point cloud instance segmentation uses a single operating mode and fixed parameters, processing data according to the same standard regardless of the complexity of the stacked goods. This leads to insufficient segmentation accuracy when the stack is complex, and wasted computing power when the stack is simple. Furthermore, multimodal fusion weights are often preset fixed values, failing to consider the impact of environmental changes on the quality of data sources. Low-quality data still participates in fusion with fixed weights, reducing the reliability of the fusion results. By adopting a solution that incorporates multimodal environment state awareness, dynamic switching of segmentation modes, and adaptive adjustment of fusion weights, the system first captures the scene and data state in real time, and then adjusts processing parameters and weight allocation based on the state quantification results. This ensures that the system's processing strategy accurately matches the real-time scene, fundamentally solving the rigidity problem of fixed strategies.
[0055] Real-time acquisition of multimodal environmental status information includes at least ambient light intensity, point cloud clustering evaluation metrics reflecting cargo stacking and occlusion levels, and quality assessment metrics for each sensor data. Ambient light intensity is collected in real-time by a light intensity sensing module attached to a multi-view vision sensor, with the sampling frequency consistent with the visual data acquisition frequency to ensure time synchronization between light and visual data. The point cloud clustering evaluation metrics are obtained by calculating the cluster dispersion of the dense 3D point cloud of cargo stacking. The dispersion is calculated by dividing the average distance between points within a cluster and the cluster center by the average distance between clusters; a higher value indicates denser cargo stacking and more severe occlusion. Among the quality assessment metrics for each sensor data, the visual image quality assessment metric is the image signal-to-noise ratio (SNR), obtained by calculating the ratio of image signal strength to noise strength. The RFID data quality assessment metric is the read success rate, calculated by the ratio of the number of valid tags to the number of electronic warehouse receipt declaration tags. The depth data quality assessment metric is the point cloud integrity rate, calculated by the ratio of the number of valid points to the theoretical number of complete points. The weight data quality assessment metric is the stability index, i.e., the weight fluctuation coefficient, calculated by the percentage difference between the weight data before and after filtering. The real-time acquisition of this status information provides a comprehensive and accurate quantitative basis for subsequent strategy adjustments, ensuring that the adjustments are clearly targeted.
[0056] Secondly, based on the point cloud clustering evaluation index, the operating mode of the point cloud instance segmentation network is dynamically selected. When the index indicates complex stacking, the fine segmentation mode is activated; otherwise, the fast segmentation mode is activated. The judgment threshold of the point cloud clustering evaluation index is calibrated to 0.7 through offline experiments. When the clustering dispersion is greater than 0.7, it is judged as complex stacking, and the fine segmentation mode is activated, adjusting the number of iterations of the point cloud instance segmentation network to 50 and the clustering threshold to 0.3. Increasing the number of iterations improves the sufficiency of feature extraction, while lowering the clustering threshold enhances the ability to distinguish closely adjacent items. When the clustering dispersion is less than or equal to 0.7, it is judged as simple stacking, and the fast segmentation mode is activated: the number of network iterations is adjusted to 20 and the clustering threshold is adjusted to 0.5. This reduces computing power consumption and improves processing efficiency while ensuring that the segmentation accuracy meets the requirements. For example, when goods are tightly stacked, the cluster dispersion is 0.8. Enabling the fine segmentation mode can accurately separate adjacent tightly stacked goods and avoid merging and counting. When goods are loosely stacked, the cluster dispersion is 0.5. Enabling the fast segmentation mode can complete the segmentation process within 1 second, balancing efficiency and accuracy.
[0057] Based on ambient light intensity and visual image quality assessment metrics, the weighting of visual quantity estimates in the fusion process is adjusted. The suitable range for ambient light intensity is set at 500-1000 lux, and the suitable threshold for visual image signal-to-noise ratio (SNR) is set at 30 dB. When the ambient light intensity is within the suitable range and the image SNR is higher than 30 dB, the visual data quality is reliable, and the fusion weight of the visual quantity estimates is adjusted to 0.3-0.4. When the ambient light intensity is lower than 200 lux or higher than 1500 lux, and the image SNR is lower than 20 dB, the visual data quality deteriorates, and the fusion weight is adjusted to 0.1-0.2. For example, on a sunny midday, the light intensity in a warehouse is 800 lux, and the image SNR is 35 dB; the visual weight is adjusted to 0.35 to fully utilize the advantages of visual data. At night, a warehouse lighting malfunction results in an light intensity of 150 lux and an image SNR of 18 dB; the visual weight is adjusted to 0.15 to reduce the interference of low-quality visual data on the fusion results.
[0058] Based on point cloud clustering evaluation metrics and depth data quality, the weight allocation of the first quantity estimate in the fusion process is adjusted. Depth data quality is quantified by point cloud completeness rate, with a suitable threshold set at 90%. When the point cloud clustering evaluation metrics show complex stacking and a point cloud completeness rate higher than 90%, the reliability of the 3D segmentation result is high, and the fusion weight of the first quantity estimate is adjusted to 0.4-0.5. When the point cloud clustering evaluation metrics show simple stacking but a point cloud completeness rate lower than 70%, the 3D segmentation result is affected by data quality, and the fusion weight is adjusted to 0.2-0.3. For example, if the goods stacking is complex and the point cloud completeness rate is 95%, the weight of the first quantity estimate is adjusted to 0.45, making it the dominant basis for the fusion result. If the goods stacking is simple but the point cloud completeness rate is 65%, the weight is adjusted to 0.25 to reduce the impact of low-quality 3D data.
[0059] Finally, based on the stability index of weight data and the read success rate index of RFID data, the fusion weight of the weight-derived quantity estimate and the RFID quantity estimate are adjusted respectively. The weight data stability index uses the weight fluctuation coefficient as the quantification standard, with a suitable threshold set at 5%. When the weight fluctuation coefficient is below 5%, the weight data is reliable, and the fusion weight of the weight-derived quantity estimate is adjusted to 0.15-0.2; when the weight fluctuation coefficient is above 10%, the weight data exhibits significant fluctuations, and the weight is adjusted to 0.05-0.1. The suitable threshold for the RFID data read success rate is set at 95%. When the read success rate is above 95%, the RFID data is reliable, and the fusion weight of the RFID quantity estimate is adjusted to 0.05-0.1; when the read success rate is below 80%, the RFID data has many omissions or faults, and the weight is adjusted to 0.02-0.05. For example, if the weight fluctuation coefficient is 3%, the weight of the estimated quantity derived from the weight is adjusted to 0.18. If the success rate of RFID data reading is 98%, the weight of the estimated quantity of RFID is adjusted to 0.08. This allows high-quality weight and RFID data to effectively supplement the deficiencies of 3D and visual data. If the weight fluctuation coefficient is 12%, the weight is adjusted to 0.08. If the success rate of RFID quantity reading is 75%, the weight is adjusted to 0.03. This avoids the interference of fluctuating weight data and low success rate RFID data on the fusion results.
[0060] In one embodiment, the step of comparing the final verified quantity with the quantity declared in the electronic warehouse receipt, and determining the cargo status and outputting the result based on the consistency between the multi-source sensing data, includes: S51, calculate the difference between the final verified quantity and the quantity declared in the electronic warehouse receipt; S52, analyze the consistency between the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate and the radio frequency identification quantity estimate, and identify whether there are contradictory evidence pairs that exceed a preset threshold; If the quantity difference is within the allowable range and all evidence pairs are consistent, the goods inspection is deemed to have passed. If the quantity difference exceeds the allowable range, it is determined to be a quantity abnormality alarm; If the difference in quantity is within the allowable range but there is a contradiction in the evidence, it is judged as a potential risk warning, and the risk type is identified according to the contradiction pattern. S53 generates a structured verification report containing the number of verifications, the value of each piece of evidence, the status determination, and the risk type, and generates digital fingerprints of key process data for blockchain storage.
[0061] As described in steps S51-S53 above, by accurately comparing the final verified quantity with the declared quantity of the electronic warehouse receipt, and combining the consistency analysis of multi-source quantity estimates, a multi-dimensional determination of the cargo status is achieved, a structured verification report is generated, and key data is stored on the blockchain to ensure the accuracy of the verification results, the comprehensiveness of risk identification, and the traceability of the process, providing the final decision output and credible evidence support for the verification of the authenticity of goods in electronic warehouse receipts.
[0062] In supply chain finance and movable asset pledge supervision scenarios, the consistency between electronic warehouse receipts and physical goods is not only reflected in quantity matching, but also implicitly in the authenticity and legality of the goods' status. Quantity comparison alone cannot identify fraudulent activities such as counterfeit labels or empty containers despite consistent quantities, and traditional verification results lack tamper-proof evidence, easily leading to disputes. Furthermore, inconsistencies between multi-source data often foreshadow potential risks; ignoring these inconsistencies and relying solely on quantity for judgment can result in missed risks, impacting financial risk control security. Therefore, it is necessary to determine the status of goods through a two-dimensional approach of quantity comparison and consistency analysis, combined with an evidence preservation mechanism, to address the limitations and insufficient credibility of single-quantity judgments, comprehensively ensuring the completeness and authority of verification.
[0063] Traditional verification methods have significant drawbacks. They focus solely on comparing the final quantity with warehouse receipts, neglecting consistency information across multiple data sources. This leads to the inability to identify potential fraud and risks. Verification results are often recorded in simple text, lacking structured presentation, which hinders subsequent traceability and analysis. Furthermore, critical process data is not immutably preserved, leaving no credible evidence in case of disputes. This step specifically employs a two-dimensional approach: quantity comparison and consistency analysis. It utilizes tiered status output, structured report generation, and blockchain-based notarization to ensure both quantitative matching and the identification of potential risks from data discrepancies. The notarization mechanism guarantees the credibility of the results, comprehensively enhancing the practicality and authority of the verification process.
[0064] The system calculates the difference between the final verified quantity and the quantity declared in the electronic warehouse receipt. The final verified quantity is derived from the rounded result of dynamic weighted fusion and is the core quantity output after multi-source data collaborative optimization. The quantity declared in the electronic warehouse receipt is the quantity of goods explicitly recorded in the electronic warehouse receipt, retrieved by the system from the electronic warehouse receipt database of the warehouse management platform. During retrieval, a unique match is made using the warehouse receipt number to ensure data accuracy. The difference is calculated using the absolute difference method, meaning the difference equals the absolute value of the final verified quantity and the quantity declared in the electronic warehouse receipt. For example, if the final verified quantity is 10 and the electronic warehouse receipt quantity is 10, the difference is 0; if the final verified quantity is 9 and the electronic warehouse receipt quantity is 10, the difference is 1. This directly quantifies the difference between the physical goods and the quantity recorded in the warehouse receipt, providing a basic quantitative indicator for subsequent status determination.
[0065] The consistency among the initial quantity estimate, visual quantity estimate, weight-derived quantity estimate, and RFID quantity estimate is analyzed to identify any contradictory evidence pairs exceeding a preset threshold. Consistency analysis is performed by comparing each quantity estimate pairwise. The preset contradiction threshold is 1, meaning that any pair of quantity estimates with an absolute difference greater than 1 is considered a contradictory evidence pair. For example, if the initial quantity estimate is 10, the visual quantity estimate is 10, the weight-derived quantity estimate is 10, and the RFID quantity estimate is 8, the difference between the RFID quantity estimate and the other three estimates is 2, exceeding the preset threshold and forming three contradictory evidence pairs. If the estimates are 10, 10, 9, and 10 respectively, and the difference between any pairwise values does not exceed 1, then there are no contradictory evidence pairs. Cross-validation of multi-source data uncovers potential problems that cannot be detected by a single quantity comparison. For example, if the RFID quantity is significantly lower than other estimates, it may indicate missing or counterfeited tags, providing a basis for risk identification.
[0066] The status of goods is determined based on the quantity difference and the presence of contradictory evidence pairs. The allowable range for the quantity difference is ±1, calibrated through offline experiments and industry standards. If the quantity difference is within the allowable range and all evidence pairs are consistent, it indicates that the quantity of the physical goods matches the warehouse receipt record, multi-source data corroborates each other, there is no potential risk, and the goods verification is deemed successful. If the quantity difference exceeds the allowable range, it indicates a substantial difference between the quantity of the physical goods and the warehouse receipt record, possibly indicating a shortage or false reporting of goods, and is thus identified as a quantity anomaly alarm. If the quantity difference is within the allowable range but contradictory evidence pairs exist, it indicates that the quantity appears to match, but there are conflicts in the multi-source data, potentially indicating potential fraud such as label forgery or empty container filling, and is thus identified as a potential risk alarm, with the risk type identified based on the contradiction pattern. For example, if the RFID quantity estimate is significantly lower than other estimates, it is identified as a risk of incomplete RFID tag coverage or forgery; if the weight-derived quantity estimate is significantly lower than the visual and three-dimensional estimates, it is identified as a risk of empty container filling. This tiered judgment mechanism clearly defines the quantity matching situation and accurately identifies potential risks, avoiding missed risk assessments due to a single judgment.
[0067] A structured verification report is generated, including the number of verifications, the values of each piece of evidence, the status determination, and the risk type. Key process data is then used to generate digital fingerprints for blockchain storage. The structured verification report uses a fixed data format, clearly listing the final verification quantity, the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, the RFID quantity estimate, the quantity difference, the number of contradictory evidence pairs, the status determination result, and the risk type (marked as none if there is no risk). This ensures the report information is clear, standardized, and easy to review and analyze later. Key process data includes multi-view image fingerprints, point cloud summaries, stable weight values, a list of valid RFID tags, dynamic fusion weights, and status determination results. Digital fingerprint generation uses the SHA256 algorithm to hash the key process data, obtaining a unique digital fingerprint of fixed length, which is then uploaded to the consortium blockchain network for storage. Blockchain storage leverages its decentralized and tamper-proof characteristics to ensure the authenticity and integrity of key data. In case of disputes, the credibility of the verification process can be verified by comparing the digital fingerprint stored on the blockchain with the original data, providing solid evidentiary support for supply chain finance risk control.
[0068] like Figure 2 As shown, this application also discloses an electronic warehouse receipt goods authenticity verification system based on multimodal AI, including: The acquisition module 1 is used to synchronously acquire multi-source sensing data of the target storage unit based on a multi-modal sensing array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; The acquisition module 2 is used to acquire three-dimensional information based on the multi-view visual data and depth data, and to perform three-dimensional reconstruction and instance segmentation of the cargo stack through the three-dimensional information to obtain cargo segmentation results and a first quantity estimate based on point cloud. Fusion module 3 is used to obtain the final verification quantity based on the multi-view visual data, depth data, weight data and radio frequency identification data through dynamic weighted fusion; Adjustment module 4 is used to acquire multimodal environment state information and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information. Output module 5 is used to compare the final verified quantity with the quantity declared in the electronic warehouse receipt, and to determine the status of the goods and output the results based on the consistency between the multi-source sensing data.
[0069] In one embodiment, the fusion module includes: The detection unit is used to perform target detection on the multi-view visual data, count the number of detection boxes, and obtain the visual quantity estimate after multi-view fusion and deduplication. The calculation unit is used to obtain the stable weight reading in the weight data and calculate the estimated weight by dividing the stable weight reading by the pre-stored standard unit weight of the goods. The statistics unit is used to count the number of valid tags in the RFID data and obtain an estimated value of the number of RFID tags. An evaluation unit is used to evaluate the real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate, respectively. The output unit is used to assign dynamic fusion weights to the first quantity estimate, visual quantity estimate, weight-derived quantity estimate and radio frequency identification quantity estimate based on the real-time confidence level, and to perform weighted summation and rounding on each quantity estimate according to the dynamic fusion weights, and output the final verification quantity.
[0070] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI.
[0071] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI.
[0072] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0074] The above description is merely a preferred embodiment of this application and does not limit the scope of this application. Any equivalent results or equivalent process transformations made based on the content of this application specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.
Claims
1. A method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI, characterized in that, Includes the following steps: Multi-source sensing data of the target storage unit is collected synchronously based on a multi-modal sensing array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; Based on the multi-view visual data and depth data, three-dimensional information is obtained, and the three-dimensional information is used to perform three-dimensional reconstruction and instance segmentation of the cargo stack to obtain a point cloud-based cargo segmentation result and a first quantity estimate. Based on the multi-view visual data, depth data, weight data, and RFID data, the final verification quantity is obtained through dynamic weighted fusion. Obtain multimodal environment state information, and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information; The final verified quantity is compared with the quantity declared in the electronic warehouse receipt, and the cargo status is judged and the result is output based on the consistency between the multi-source sensing data.
2. The method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI according to claim 1, characterized in that, The steps for synchronously acquiring multi-source sensing data of the target storage unit based on a multimodal sensor array include: Deploy a multimodal sensor array that includes multi-view vision sensors, depth sensors, weight sensors and RFID read / write units, and perform spatial and temporal synchronous calibration; Based on the verification command, the multi-view vision sensor, depth sensor, weight sensor and RFID reading and writing unit are simultaneously triggered to collect the multi-view vision data, depth data, weight data and radio frequency identification data at the same time. The multi-view visual data, depth data, weight data, and RFID data are preprocessed.
3. The method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI according to claim 1, characterized in that, The steps of obtaining three-dimensional information based on the multi-view visual data and depth data, performing three-dimensional reconstruction and instance segmentation of the cargo stack using the three-dimensional information, and obtaining a point cloud-based cargo segmentation result and a first quantity estimate include: Based on the multi-view visual data and the depth data, a dense three-dimensional point cloud of the cargo pile is generated through stereo matching and coordinate transformation. The dense 3D point cloud is input into a pre-trained point cloud instance segmentation network, which assigns an instance identifier to each point in the point cloud. Based on the instance identifier, points belonging to the same independent physical cargo are aggregated into a single cargo point cloud cluster; The number of different cargo point cloud clusters is counted and used as the first quantity estimate. The spatial location information of each cargo point cloud cluster is then output.
4. The method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI according to claim 1, characterized in that, The step of obtaining the final verification quantity based on the multi-view visual data, depth data, weight data, and RFID data through dynamic weighted fusion includes: Target detection is performed on the multi-view visual data, the number of detection boxes is counted, and after multi-view fusion and deduplication, the estimated number of visual boxes is obtained. Obtain stable weight readings from the weight data, and calculate the estimated weight quantity by dividing the stable weight readings by the pre-stored standard unit weight of the goods. The number of valid tags in the RFID data is counted to obtain an estimated number of RFID tags. The real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the RFID quantity estimate are evaluated respectively. Based on the real-time confidence level, dynamic fusion weights are assigned to the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the RFID quantity estimate. The quantity estimates are then weighted, summed, and rounded according to the dynamic fusion weights, and the output is the final verification quantity.
5. The method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI according to claim 4, characterized in that, The steps of acquiring multimodal environment state information and adaptively adjusting the processing parameters of instance segmentation and the weight allocation strategy of dynamic weighted fusion based on the multimodal environment state information include: The multimodal environmental state information is acquired in real time. The state information includes at least ambient light intensity, point cloud clustering evaluation index reflecting the degree of cargo stacking and occlusion, and quality assessment index of each sensor data. Based on the point cloud clustering evaluation index, the operating mode of the point cloud instance segmentation network is dynamically selected. When the index shows complex stacking, the fine segmentation mode is enabled; otherwise, the fast segmentation mode is enabled. Based on the ambient light intensity and visual image quality evaluation indicators, adjust the weight allocation of the estimated visual quantity in the fusion process; Based on the point cloud clustering evaluation index and the depth data quality, adjust the weight allocation of the first quantity estimate in the fusion; Based on the stability index of the weight data and the reading success rate index of the RFID data, the fusion weights of the weight-derived quantity estimate and the RFID quantity estimate are adjusted respectively.
6. The method for verifying the authenticity of goods in electronic warehouse receipts based on multimodal AI according to claim 5, characterized in that, The steps of comparing the final verified quantity with the quantity declared in the electronic warehouse receipt, and determining the cargo status and outputting the results based on the consistency between the multi-source sensing data, include: Calculate the difference between the final verified quantity and the quantity declared in the electronic warehouse receipt; Analyze the consistency among the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate to identify whether there are contradictory evidence pairs that exceed a preset threshold; If the quantity difference is within the allowable range and all evidence pairs are consistent, the goods inspection is deemed to have passed. If the quantity difference exceeds the allowable range, it is determined to be a quantity abnormality alarm; If the difference in quantity is within the allowable range but there is a contradiction in the evidence, it is judged as a potential risk warning, and the risk type is identified according to the contradiction pattern. Generate a structured verification report that includes the number of verifications, the value of each piece of evidence, the status determination, and the risk type, and generate digital fingerprints of key process data for blockchain storage.
7. A multimodal AI-based electronic warehouse receipt goods authenticity verification system, characterized in that, include: The acquisition module is used to synchronously acquire multi-source sensing data of the target storage unit based on a multi-modal sensor array, wherein the multi-source sensing data includes multi-view visual data, depth data, weight data and radio frequency identification data; The acquisition module is used to acquire three-dimensional information based on the multi-view visual data and depth data, and to perform three-dimensional reconstruction and instance segmentation of the cargo stack through the three-dimensional information to obtain cargo segmentation results and a first quantity estimate based on point cloud. The fusion module is used to obtain the final verification quantity based on the multi-view visual data, depth data, weight data and radio frequency identification data through dynamic weighted fusion. An adjustment module is used to acquire multimodal environment state information and adaptively adjust the processing parameters of the instance segmentation and the weight allocation strategy of the dynamic weighted fusion based on the multimodal environment state information. The output module is used to compare the final verified quantity with the quantity declared in the electronic warehouse receipt, and to determine the status of the goods and output the results based on the consistency between the multi-source sensing data.
8. The electronic warehouse receipt goods authenticity verification system based on multimodal AI according to claim 7, characterized in that, The fusion module includes: The detection unit is used to perform target detection on the multi-view visual data, count the number of detection boxes, and obtain the visual quantity estimate after multi-view fusion and deduplication. The calculation unit is used to obtain the stable weight reading in the weight data and calculate the estimated weight by dividing the stable weight reading by the pre-stored standard unit weight of the goods. The statistics unit is used to count the number of valid tags in the RFID data and obtain an estimated value of the number of RFID tags. An evaluation unit is used to evaluate the real-time confidence levels of the first quantity estimate, the visual quantity estimate, the weight-derived quantity estimate, and the radio frequency identification quantity estimate, respectively. The output unit is used to assign dynamic fusion weights to the first quantity estimate, visual quantity estimate, weight-derived quantity estimate and radio frequency identification quantity estimate based on the real-time confidence level, and to perform weighted summation and rounding on each quantity estimate according to the dynamic fusion weights, and output the final verification quantity.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.