Container number tallying identification method
Through multi-sensor fusion and deep learning models, combined with dynamic Bayesian network and blockchain evidence storage technology, the problems of low accuracy and high operating costs of container number identification in complex environments are solved, and efficient and reliable container number identification and verification are achieved.
Patent Information
- Application Number
- CN202510922133.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing container number identification technology has low recognition accuracy in complex environments, making it difficult to meet the efficient operation needs during peak hours of ports, and lacks a dynamic verification mechanism, resulting in an increase in operating costs and time costs.
Multi-sensor fusion of distributed stereoscopic vision arrays, millimeter-wave radars and lidars is adopted, and combined with reinforcement learning and deep learning models, multi-modal data collaborative acquisition, heterogeneous data fusion, character area intelligent positioning and adaptive segmentation are carried out, and multi-dimensional verification is carried out in combination with dynamic Bayesian networks. Finally, the accuracy and traceability of the identification results are ensured through blockchain evidence storage technology.
It realizes high-precision container number identification in complex environments, has environmental adaptive sensing capabilities, multi-level verification decision-making, ensures the accuracy and reliability of identification results, meets the high-speed passing needs during peak hours of ports, and reduces operating costs.
Smart Images

Figure CN120411989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of container intelligent tallying, and specifically to a method for tallying and identifying container numbers. Background Art
[0002] In the container operation scenarios of modern ports and logistics centers, the accurate identification of container numbers is the core link of the tallying work, which is directly related to the efficiency, safety, and traceability of cargo transportation. The traditional manual identification method relies on tally clerks to check each container one by one. During the peak operation period, not only is the work intensity extremely high, but it is also extremely easy to have omissions due to fatigue. In actual operation, manual identification often has problems such as incorrect character recording and number confusion, resulting in the disorder of the cargo loading and unloading sequence. Seriously, it may even lead to chaos in transportation scheduling, causing the wrong flow of goods, bringing great uncertainty and risks to logistics transportation.
[0003] Although the existing automatic identification technologies have improved the efficiency to a certain extent, there are still many limitations. For the identification method based on single vision, in actual port operations, the identification effect is often affected by environmental factors. For example, at noon in summer, the strong sunlight directly shines on the surface of the container, and the resulting reflection overexposes the character area, making it difficult for the camera to capture clear images; while at night or in a dimly lit warehouse, the images taken in low-light environments have a lot of noise and the character features are blurred. These situations greatly reduce the identification accuracy. In actual applications, after many logistics parks adopt single-camera identification systems, identification errors frequently occur due to lighting problems, resulting in incorrect shipment of goods, seriously affecting customer satisfaction and corporate reputation.
[0004] Although the identification scheme of multi-sensor fusion introduces depth information to assist in identification, in actual applications, the immaturity of the data fusion algorithm makes it difficult to effectively integrate the data collected by each sensor. For example, in an automated terminal, even if the equipment is equipped with a vision camera and a lidar at the same time, due to the low data fusion efficiency and insufficient feature extraction, it still takes a long time to process a single container number and cannot meet the requirements of efficient operation during the peak period of the port. In addition, the existing technology lacks a dynamic verification mechanism. Facing abnormal situations such as common stains covering the surface of the container, characters being worn due to long-term use, and partial occlusion caused by stacking, it is difficult to accurately identify, and often requires manual secondary verification, greatly increasing the operating cost and time cost. Therefore, we provide a method for tallying and identifying container numbers to solve the above technical problems. Summary of the Invention
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0006] A method for tallying and identifying container numbers includes the following steps:
[0007] S1. System Initialization and Parameter Configuration: Deploy a distributed stereo vision array, millimeter-wave radar, and lidar and perform initial calibration, configure reinforcement learning parameters, pre-train an improved YOLOv7-tiny object detection model, a Transformer-based cGAN character segmentation model, and a hybrid recognition model integrated with a lightweight CNN-Transformer-capsule network, construct a character structure knowledge graph and initialize the parameters of the dynamic Bayesian network.
[0008] Among them, the initial calibration includes: performing internal and external parameter calibration on the distributed stereo vision array through the Zhang's calibration method, with the internal parameter calibration accuracy reaching 0.1 pixel level and the external parameter calibration error controlled within ±0.5°. The pre-training model adopts a multi-modal pre-training strategy based on contrastive learning, and enhances the generalization ability of the model by constructing image-point cloud-text contrastive learning tasks. Among them, the temperature parameter of the contrastive learning is set to 0.07, and the positive-negative sample ratio is 1:3.
[0009] S2. Multi-modal Data Cooperative Acquisition: Through a distributed stereo vision array with a self-calibration mechanism, combined with the multi-sensor fusion positioning of millimeter-wave radar and lidar, synchronously collect RGB images, TOF depth images, millimeter-wave point cloud data, and lidar point cloud data on multiple sides of the container; use a dynamic threshold triggering and adaptive exposure control strategy based on a deep Q network to obtain multi-modal data containing the container number area.
[0010] Among them, the dynamic threshold triggering strategy uses a double-layer Q network architecture to achieve multi-objective optimization, including a target network and a real network. The parameters of the target network are updated every 100 steps, and the real network uses an Adam optimizer with a learning rate of 0.001; the adaptive exposure control strategy dynamically adjusts the exposure time by calculating the image entropy value, and the exposure time range is 1ms - 100ms, and the adjustment step size is 0.1ms.
[0011] S3. Heterogeneous Data Fusion and Preprocessing: Adopt a cross-modal feature fusion network based on the Transformer architecture, combined with a dynamic routing mechanism, to perform feature-level fusion of multi-source data; use a noise reduction model based on a generative adversarial network and a geometric distortion correction algorithm based on deep learning to preprocess the fused data.
[0012] Among them, the cross-modal feature fusion network contains a multi-head self-attention mechanism with 8 heads. The dynamic routing mechanism adaptively allocates fusion weights by calculating the quantum entanglement similarity of different modal features, and the similarity calculation uses the quantum state inner product formula; the noise reduction model consists of a generator and a discriminator. The generator adopts a U-Net structure, including 5 downsampling blocks and 5 upsampling blocks, and the discriminator adopts a PatchGAN structure. The loss function for adversarial training is LSGAN.
[0013] S4. Intelligent Character Region Localization: Based on the improved YOLOv7-tiny object detection model, integrating spatio-temporal attention mechanism, character structure knowledge graph, and semantic association analysis based on graph neural network, detect the character regions in the preprocessed data, and output the localization results including bounding box coordinates, confidence, character category probability, and semantic association score;
[0014] Among them, the improved YOLOv7-tiny model adopts a collaborative optimization architecture of progressive knowledge distillation and dynamic network pruning. The temperature parameter of knowledge distillation is set to 4, the pruning ratio is 30%, and the number of model parameters after pruning is reduced by 40%. The spatio-temporal attention mechanism captures the changing features of character regions at different time steps through a spatio-temporal feature pyramid structure. The feature pyramid contains 3 scales, corresponding to 1 / 4, 1 / 8, and 1 / 16 resolutions of the original image respectively;
[0015] S5. Adaptive Character Segmentation and Normalization: Use the cGAN character segmentation algorithm based on Transformer, combined with morphological post-processing of topological structure analysis; through the adaptive normalization network based on the differentiable deformation module, achieve precise segmentation and normalization of character images;
[0016] Among them, the differentiable deformation module adopts a hybrid deformation model combining bilinear interpolation and thin plate spline transformation; the morphological post-processing of topological structure analysis includes calculating topological invariants such as the Euler number and the number of holes in the character connected region;
[0017] S6. Multi-Model Collaborative Character Recognition: Adopt a hybrid model integrating lightweight CNN, Transformer, and capsule network, combined with bidirectional knowledge distillation technology and dynamic weight fusion strategy, perform character classification, output the character probability distribution, and combine them in spatial order to form a preliminary recognition result;
[0018] Among them, bidirectional knowledge distillation adopts an adversarial knowledge transfer mechanism. The teacher model guides the student model to learn through soft labels, and the student model optimizes the feature expression of the teacher model through adversarial training feedback. The discriminator of adversarial training adopts a multi-layer perceptron, including 2 hidden layers, with 256 neurons in each layer; the dynamic weight fusion strategy adopts a reinforcement learning optimization framework, and dynamically adjusts the model weights with the recognition accuracy and computational efficiency as the optimization goals. The reward function of reinforcement learning is recognition accuracy × 0.8 + computational speed × 0.2;
[0019] S7. Multi-Dimensional Intelligent Verification and Decision-Making: Construct a verification and decision-making model based on a dynamic Bayesian network, combined with format compliance verification, check code correctness verification, spatio-temporal consistency verification, multi-modal data consistency verification, and trend prediction verification based on historical data, to evaluate the credibility of the container number recognition result; when an abnormal verification result occurs in a certain item, start a targeted verification enhancement strategy;
[0020] Among them, the dynamic Bayesian network adopts an inference mechanism that fuses the evidence theory and Bayesian optimization. The basic probability assignment function of the evidence theory is obtained through historical data statistics, and the acquisition function of Bayesian optimization is the expected improvement (EI) function; the targeted verification enhancement strategy adopts a meta-verification framework, which automatically selects the optimal verification enhancement scheme according to the verification anomaly type. The meta-verification framework includes 5 verification enhancement strategies, and the strategy selection probability is optimized by a genetic algorithm;
[0021] S8, Result Output and Blockchain Archiving: Output the verified container number and related information to the tallying system. At the same time, through the blockchain archiving technology based on zero-knowledge proof, combined with the data processing method of homomorphic encryption, tamper-proof archiving is carried out;
[0022] Among them, the zero-knowledge proof adopts the zk-STARKs technology, and the generated proof length is 1024 bits, and the verification time is less than 100 ms; the homomorphic encryption adopts the BGV homomorphic encryption scheme, which supports addition and multiplication operations on ciphertexts, and the encryption parameter is set as the modulus q = 2 40 , and the plaintext space is Z_2 10 ; The blockchain archiving is implemented by sidechain technology to handle high-frequency transactions.
[0023] Furthermore, in S1, the character structure knowledge graph predefines constraint conditions such as the character spacing being 1.2 - 1.5 times the character height and the aspect ratio being 0.8 - 1.2, etc., to filter out candidate boxes that do not conform to the structural rules; the nodes of the knowledge graph include character types, character sizes, character spacings, etc., and the edges include spatial relationships, semantic relationships, etc. The construction of the knowledge graph uses the Neo4j graph database, and the attributes of the nodes and edges are determined by combining expert annotation and machine learning.
[0024] Furthermore, in S2, the self-calibration mechanism, through the checkerboard calibration board and IMU deployed in the transportation channel, performs spatio-temporal calibration on the stereo vision array, millimeter-wave radar, and lidar in real time based on the Kalman filter algorithm. The state transition matrix of the Kalman filter is determined by the motion model of the IMU, and the observation matrix is determined by the corner detection results of the checkerboard; the distributed stereo vision array adopts a non-uniform redundancy design, and dynamically adjusts the sensor layout density according to the recognition requirements of the key areas of the container. The sensor density in the key area is 2 times that in the non-key area.
[0025] Furthermore, in S3, the geometric distortion correction algorithm adopts a differentiable mapping model based on diffeomorphism, and optimizes the mapping parameters by minimizing the image reprojection error. The threshold of the reprojection error is set to 1 pixel; both the generator and discriminator of the noise reduction model adopt the LeakyReLU activation function with a slope of 0.2, and the momentum parameter of the batch normalization layer is set to 0.9.
[0026] Furthermore, in S4, the semantic association analysis of the graph neural network uses a graph convolutional network (GCN), which includes 2 graph convolutional layers, the output dimension of each layer is 128, and the activation function is ReLU; the dynamic update of the character structure knowledge graph is achieved through incremental learning. After each recognition result update, the confidence scores of the knowledge graph are recalculated, and the nodes and edges with confidence scores lower than 0.5 are deleted.
[0027] Furthermore, in S5, the Transformer-based cGAN character segmentation algorithm introduces an adaptive receptive field module in the generator. The size of the receptive field is dynamically adjusted according to the different morphological features of the characters, and the adjustment range is from 3×3 to 11×11; the adaptive normalization network uses an attention-guided spatial transformation network. The attention weight map is generated by the Softmax function, and the parameters of the spatial transformation network are learned through backpropagation.
[0028] Furthermore, in S6, the lightweight CNN of the hybrid model uses the MobileNetV3 architecture, the Transformer uses 6 layers of encoders, the number of capsules in the capsule network is 16, and the capsule dimension is 32; the reinforcement learning optimization framework of the dynamic weight fusion strategy uses the PPO algorithm. Both the policy network and the value network use multi-layer perceptrons, the size of the hidden layer is 256, and the training batch size is 64.
[0029] Furthermore, in S7, the format compliance verification is implemented through a probabilistic finite state automaton. The state transition probability of the automaton is obtained by statistical analysis of historical data, and the threshold of state transition is set to 0.8; the spatio-temporal consistency verification is constructed using a spatio-temporal graph convolutional network. The graph convolutional network includes 3 spatio-temporal convolutional layers, the time step is set to 5, and the spatial neighborhood size is 3×3.
[0030] Furthermore, in S8, the side chain of blockchain evidence storage uses the Practical Byzantine Fault Tolerance (PBFT) consensus mechanism, the number of consensus nodes is 5, and the tolerance number of Byzantine nodes is 1; the evidence storage data is organized in a Merkle tree structure. The depth of the Merkle tree is 10, the leaf nodes store data hash values, and the parent nodes store combinations of child node hash values.
[0031] Furthermore, in S9, the construction of the dynamic training dataset adopts an active learning strategy. The most valuable error cases are selected through uncertainty sampling. The uncertainty measure uses the prediction entropy value, and the entropy threshold is set to 0.9; the parameter optimization of the reward function of reinforcement learning uses the Trust Region Policy Optimization (TRPO) algorithm. The size of the trust region is set to 0.01, and the optimization step size is set to 0.5.
[0032] A system for a container number tally recognition method, including a hardware layer and a software layer; the hardware layer includes a multi-modal perception unit, an edge computing unit, a trigger and control unit, and an auxiliary unit; the software layer includes a multi-modal data processing module, an intelligent recognition module, an intelligent verification and decision-making module, an interface and evidence storage module, and a system management module. Among them, the edge computing unit adopts a heterogeneous computing architecture, integrating GPU, FPGA, NPU, and a dedicated AI acceleration chip, and realizes the intelligent allocation of computing tasks through a task scheduling optimization engine. The task scheduling optimization engine adopts a genetic algorithm, with a population size of 50 and an iteration number of 100.
[0033] A computer-readable storage medium stores a computer program. The program includes a system initialization and parameter configuration module, a multi-modal data collaborative acquisition control module, a heterogeneous data fusion preprocessing module, a character area intelligent positioning module, an adaptive character segmentation and normalization module, a multi-model collaborative character recognition module, a multi-dimensional intelligent verification and decision-making module, and a blockchain evidence storage interface module; when the computer program is executed by a processor, it is used for the steps of the container number tally recognition method. Among them, the multi-model collaborative character recognition module adopts an uncertainty quantification method for model integration, estimates the uncertainty of the recognition result through the Monte Carlo Dropout technique, sets the probability of Dropout to 0.5, and the number of Monte Carlo samplings is 10.
[0034] The present invention provides a container number tally recognition method, having the following beneficial effects:
[0035] 1. Full-chain high-precision recognition: A three-level cascaded recognition architecture is adopted to achieve precise processing. The improved YOLOv7-tiny model combines knowledge distillation and spatio-temporal attention to effectively solve the problem of missed detection of small character targets during high-speed movement; the Transformer-cGAN segmentation model realizes the complete segmentation of rusty and sticky characters through dynamic receptive fields and topological analysis; the lightweight integrated model fuses local texture, global semantics, and spatial hierarchical features, combines a dynamic weight strategy, and realizes the precise distinction of easily confused characters, constructing a full-process high-precision processing system from object detection to semantic recognition.
[0036] 2. Environment adaptive perception: The multi-modal data acquisition module realizes environmental robustness through the cooperation of hardware and algorithms. The distributed vision array and IMU are jointly calibrated to quickly complete spatio-temporal calibration during device vibration to ensure the precise registration of multi-source data; the dynamic threshold triggering strategy combines a double-layer Q network to adaptively adjust sensor parameters and maintain stable data acquisition quality in a wide illuminance range and extreme weather. Together with GAN noise reduction and distortion correction technologies, an environmental adaptation closed-loop from data acquisition to preprocessing is formed.
[0037] 3. Multi-level verification decision-making. A three-level verification architecture constructs a data security protection system. The basic verification layer filters invalid data through format compliance verification and spatio-temporal trajectory analysis; the meta-verification framework of the enhanced verification layer integrates a multi-strategy dynamic response mechanism to trigger quick review for low-confidence results; the blockchain evidence storage layer uses zero-knowledge proof and homomorphic encryption technologies to achieve immutable evidence storage of verified data, comprehensively ensuring the accuracy and traceability of tally data from format verification to exception handling and then to data evidence storage.
[0038] 4. Edge-cloud collaboration architecture. The hardware layer adopts heterogeneous computing design to achieve efficient processing. Cross-modal feature fusion and deep learning inference are completed on edge devices, and the millisecond-level processing speed meets the high-speed passing requirements of hundreds of containers per hour in the port; the edge computing architecture reduces network dependence and maintains stable operation in complex communication environments, building an intelligent tally infrastructure with both real-time performance and reliability, providing technical support for the automated operation of smart ports.
[0039] The organic combination of the above technical solutions enables the present invention to achieve a significant improvement in recognition performance in complex scenarios such as high speed, strong light, and rust, constructs a fully automated process from data acquisition to decision-making and evidence storage, and provides a practical solution for the intelligent upgrade of port tally operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the overall system flow chart of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Embodiment: As Figure 1 shown, a container number tally recognition technology based on multi-modal fusion and intelligent decision-making is disclosed. In the container number tally recognition system, the design of each model and algorithm closely follows the data processing and decision-making process, and each link from multi-modal data acquisition, feature extraction, character recognition to result verification is carefully designed and optimized to achieve the recognition goal of high precision and high robustness. The design processes of each key model and algorithm will be elaborated in detail below, and specific examples will be given in combination with actual application scenarios.
[0043] 1. Multi-modal data acquisition:
[0044] 1.1. Self-calibration mechanism algorithm:
[0045] (1) Hardware Deployment: Install a checkerboard calibration board at a fixed position in the container transportation channel. At the same time, integrate an IMU (Inertial Measurement Unit) into the sensor device. The checkerboard calibration board serves as the reference for visual calibration, and the IMU is used to monitor the motion state of the device in real time. Specifically as follows:
[0046] At key positions of the gantry and sensor brackets in the container transportation channel, evenly install 10 Xsens MTi-G-700 inertial measurement units (IMUs). At the same time, install 12 Basler acA4112-8gm industrial cameras in a matrix distribution on both sides and the top of the channel, equipped with Schneider 12mm industrial lenses, with the distortion rate controlled within <0.1% to capture high-quality images, thus forming a distributed stereo vision array that can obtain container parameters and images from multiple angles. Among them, the distributed stereo vision array performs internal and external parameter calibration through the Zhang's calibration method. The internal parameter calibration accuracy reaches 0.1 pixel level, and the external parameter calibration error is controlled within ±0.5°, and the translation error is ±0.8 cm.
[0047] The Distributed Stereo Vision Array is a stereo vision system based on the distributed deployment of multiple cameras (or visual sensors). By collaborating on the perception data of multiple independent vision nodes, it realizes functions such as high-precision reconstruction of three-dimensional scenes, target positioning, and tracking. And the distributed stereo vision array adopts a non-uniform redundancy design. In the key area where the container number is located, 4 industrial cameras are deployed per square meter, and 2 are deployed per square meter in non-critical areas. This layout has been tested by simulating the actual port scene, which can reduce the hardware cost while ensuring the recognition accuracy. In the self-calibration mechanism, a checkerboard calibration board is set every 5 meters in the transportation channel. The IMU selects a high-precision inertial measurement unit (such as Xsens MTi-30), and based on the Kalman filter algorithm, it performs spatio-temporal calibration on the stereo vision array, millimeter-wave radar, and lidar in real time. The state transition matrix is constructed through the acceleration and angular velocity data collected by the IMU, and the observation matrix is constructed through the detection results of the checkerboard corner points to achieve precise calibration between sensors and ensure the spatial consistency of multi-modal data.
[0048] The IMU is installed at the center of the gantry beam, the four corners of the sensor bracket and other positions vulnerable to vibration to ensure comprehensive perception of the device motion. The industrial cameras are arranged at intervals of 2 meters on both sides of the channel, and one is installed every 3 meters on the top to form an all-round visual coverage.
[0049] (2) Data Acquisition: During the operation of the sensor, the IMU collects data such as acceleration and angular velocity at a frequency of 100 Hz, and the visual sensor periodically captures checkerboard images.
[0050] (3) State Prediction: Based on the data collected by the IMU, use the state transition equation of the extended Kalman filter algorithm Predict the state of the sensor at the current moment.
[0051] Formula analysis: This equation is used to predict the state of the sensor at a certain moment , where is the state transition matrix that describes the motion law of the sensor. For example, when the vehicle-mounted sensor moves with the vehicle, it can be constructed according to the vehicle kinematic model and is used to convert the state at the previous moment to the current moment; is the control input matrix. When there are manual or system control motion instructions for the sensor, represents the influence of the control input on the state; is the process noise, which is used to characterize the interference factors that cannot be accurately modeled in actual motion, such as the bumps and vibrations during vehicle driving. It is usually assumed to follow a Gaussian distribution , is the process noise covariance matrix, where it is a control quantity set manually or by the system and is used to describe the active control or influencing factors of the external world on the motion state of the sensor.
[0052] For example, when a container transport vehicle passes over a speed bump, the IMU collects the sudden change data of acceleration as , combined with the and preset according to the vehicle motion characteristics, and uses this equation to predict the state of the sensor after vibration, providing a reference for subsequent calibration.
[0053] (4) Observation update: By detecting the corner points of the checkerboard image and combining with Zhang's calibration method to calculate the external parameters of the sensor, an observation equation is constructed, and the predicted state is fused with the observation result to update the spatio-temporal parameters of the sensor and achieve self-calibration.
[0054] Formula analysis: This equation is used to establish the relationship between the predicted state of the sensor and the actual observation data . is the observation matrix, which maps the state variables to the observation space. For example, when performing visual calibration through the checkerboard image, it can convert the spatial attitude parameters of the sensor into the coordinates of the checkerboard corner points in the image according to the principle of Zhang's calibration method; is the observation noise, which reflects the errors existing in actual observations, such as the measurement deviations caused by the resolution limit of the camera image, light interference, etc. It is usually also assumed to follow a Gaussian distribution , is the observation noise covariance matrix.
[0055] For example, after a vision sensor captures an image of a checkerboard, the detected corner coordinates are used as , and through the known and the predicted state the observation value is calculated and compared with the actual to update the sensor parameters.
[0056] Example: Taking the container transportation channel in a port as an example, when a transport vehicle passes over a speed bump, the strong vibration generated causes the vision sensor installed on the gantry to shift. At this time, the IMU collects the acceleration of the device in the X-axis direction reaching and the change in the angular velocity in the Y-axis is at a frequency of 100 Hz in real time. Based on this data, the extended Kalman filter algorithm predicts the offset angle and displacement of the sensor in space. At the same time, the vision sensor captures an image of a checkerboard, and by detecting the corners and combining with Zhang's calibration method, the actual external parameter change is calculated. The system fuses the predicted state with the actual observation result, adjusts the rotation and translation parameters of the sensor, and completes self-calibration within 2 seconds to ensure the accuracy of the subsequent captured image data.
[0057] The Zhang Zhengyou calibration method is used to calibrate the internal parameters of the devices in the distributed stereo vision array. 20 sets of calibration images (resolution 1920×1080) are collected, and the internal parameter matrix is calculated. Among them, the focal length parameter reaches an accuracy of 0.1 pixel level, are the focal lengths of the camera in the x and y axis directions (unit: pixel) respectively, reflecting the scaling ability of the lens to the image, and the positioning error of the principal point coordinates is less than 0.5 pixel.
[0058] 1.2. Timestamp synchronization mechanism:
[0059] Time calibration of multi-modal perception devices through the timestamp synchronization mechanism is the core link to ensure the spatio-temporal consistency of multi-source data. In specific implementation, a combination of hardware timestamps and software algorithms is used: Each device embeds a high-precision timestamp at the moment of data acquisition, and global clock synchronization is achieved through PTP or GPS time service. The edge computing unit monitors the time deviation of each device in real time, and dynamically adjusts the time offset based on the Kalman filter algorithm to control the time error within ±10 ms.
[0060] The time calibration process includes: ① Aligning the clock reference when the device is initialized; ② Periodically sending synchronization signals during operation (e.g., once per second); ③ Carrying timestamps and marking the acquisition order during data acquisition. Through this mechanism, clock drift between multiple devices can be eliminated, the spatio-temporal misalignment of multi-modal data caused by nanosecond-level time differences can be avoided, and an accurate time correspondence relationship can be ensured when fusing RGB images, point cloud data, etc., providing a reliable time reference for subsequent character positioning and multi-model collaborative recognition, and improving the accuracy of data fusion and the system robustness in complex scenarios.
[0061] 1.3. Dynamic threshold triggering strategy algorithm:
[0062] The dynamic threshold triggering strategy adopts a double-layer Q-network architecture. The target network copies parameters from the real network every 100 steps. The real network uses the Adam optimizer with a learning rate set to 0.001. The optimization objective function is set to "maximize the detection accuracy × 0.7 + minimize the false detection rate × 0.3". This weight allocation is determined through multiple simulation experiments and can effectively achieve multi-objective optimization. The adaptive exposure control strategy dynamically adjusts the exposure time according to the image entropy value. When the image entropy value is lower than the threshold of 80, the exposure time is increased in steps of 0.1 ms; when it is higher than the threshold of 120, the exposure time is reduced to ensure that the images collected under different lighting conditions are all clear and usable. Specifically as follows:
[0063] (1) State definition: Determine the state space, including environmental characteristics in 8 dimensions such as the current frame image entropy value (reflecting image clarity, range 0 - 8), depth map variance (measuring the stability of depth data, range 0 - 1000), and millimeter-wave radar point cloud density (the number of point clouds per unit volume, range 0 - 50 points / m³).
[0064] The trigger and control unit uses 3 Advantech UNO-3083 intelligent controllers, which are respectively deployed at the channel monitoring room, the device power supply box, and the edge computing node. The multi-modal perception unit includes 16 Basler industrial cameras, which are symmetrically installed in two rows on both sides of the channel; 8 Intel RealSense D455 TOF depth sensors are spaced between the cameras; 4 Continental ARS408-21 millimeter-wave radars are installed above the channel entrances and exits; 2 Velodyne VLP-16 lidars are arranged at a high position in the middle of the channel. The NDT algorithm is used to perform point cloud registration on the millimeter-wave radar and the lidar, and the registration error is less than 5 cm. Specifically as follows:
[0065] Synchronously collect 100 groups of calibration data of the Continental ARS408-21 millimeter-wave radar and the Velodyne VLS-128 lidar, and perform point cloud registration based on the NDT (Normal Distributions Transform) algorithm.
[0066] Configure the registration parameters with a grid resolution of 0.1 m and 50 iterations. Finally, achieve high-precision registration with a root mean square error < 5 cm. The time synchronization module uses the IEEE 1588 Precision Clock Protocol to achieve multi-sensor nanosecond-level synchronization through a hardware timestamp counter, and the synchronization accuracy is controlled within ±10 μs.
[0067] Industrial cameras are installed at a height of 5 meters above the ground on both sides of the channel, with two rows spaced 1.5 meters apart vertically; the TOF depth sensors are installed staggered with the cameras to ensure data complementarity; the millimeter-wave radar is installed horizontally above the entrance and exit to cover the vehicle passage area; the lidar is installed on a column 8 meters high in the middle of the channel to obtain a wider scanning range.
[0068] (2) Network architecture: Adopt a double-layer deep Q-network (DQN) architecture, including a target network and a real network. The parameters of the target network are copied and updated from the real network every 100 steps. The real network is trained using the Adam optimizer, and the learning rate is set to 0.001.
[0069] (3) Reward function design:
[0070] Design the reward function , comprehensively consider the recognition accuracy rate, acquisition efficiency, and false trigger rate, and guide the algorithm to learn the optimal trigger strategy. In the formula, is the recognition accuracy rate, that is, the ratio of the number of correctly recognized container numbers to the total number of recognized numbers, which is used to measure the correctness of the system's recognition of container numbers and reflects the core functional performance of the system; is the acquisition efficiency, which can be defined as the number of effectively acquired container numbers per unit time or the average time to complete a complete acquisition-recognition process; is the false trigger rate, that is, the proportion of the system's false trigger of data acquisition or generation of incorrect recognition results; the coefficients 0.6, 0.3, and -0.1 are the weight coefficients of the recognition accuracy rate, acquisition efficiency, and false trigger rate respectively, which are set according to the system requirements and actual application scenarios, and reflect the relative importance of each index in the system optimization goal. In the container number tally recognition system, since recognition accuracy is the primary goal, a higher weight is given to the recognition accuracy rate; the acquisition efficiency is the second; and false triggers will have a negative impact, so a negative weight is given to the false trigger rate.
[0071] In the dynamic threshold triggering strategy, the double DQN updates the network parameters according to the reward value R obtained after each action execution (such as adjusting the acquisition interval, exposure time, etc.), and adjusts the triggering strategy. For example, if the current strategy results in a low recognition accuracy, the reward value will be correspondingly reduced, and the algorithm will adjust the strategy through learning to improve the accuracy, thereby guiding the system to automatically optimize the data acquisition timing and parameter settings in a complex environment, and improving the acquisition efficiency and reducing the false triggering rate while ensuring the recognition accuracy.
[0072] (4) Training process: By interacting with the environment, continuously collect data, calculate the reward value, update the network parameters, and gradually optimize the triggering strategy.
[0073] Implementation example: Taking the early peak period of an automated terminal as an example, the traffic flow of container trucks in the channel reaches twice that of normal days, the vehicle speeds are generally fast and the vehicle spacing is small. Three Advantech intelligent controllers interact in real time through gigabit network ports to jointly monitor the environmental status. The entropy value of the current frame image is calculated to be 2.2 (low clarity), the variance of the depth map reaches 700 (large fluctuation), and the point cloud density of the millimeter-wave radar is 28 points / m3. Based on the dynamic threshold triggering strategy of the deep Q network, the controller quickly shortens the acquisition interval of 16 industrial cameras from 200 ms to 100 ms through 8 GPIO interfaces, increases the exposure time of 8 Intel RealSense D455 from 6 ms to 15 ms, and at the same time controls the data output frequency of 4 millimeter-wave radars and 2 lidars to be increased to 30 Hz through the gigabit network port. This series of adjustments are completed within 300 milliseconds, enabling the system to effectively reduce the acquisition of invalid data in the high-speed moving container scenario, significantly improve the processing efficiency, and ensure the smooth operation during the peak period.
[0074] 2. Related models and algorithms for data fusion and preprocessing:
[0075] 2.1 Cross-modal feature fusion network:
[0076] The cross-modal feature fusion network is based on the Transformer architecture and includes a multi-head self-attention mechanism with 8 heads. The dynamic routing mechanism adaptively allocates fusion weights by calculating the quantum entanglement similarity of different modal features (using the quantum state inner product formula). Among them, the quantum state represents a high-dimensional mapping based on the feature vector. This method has been verified by theory and experimental tests and can effectively capture the internal correlations of different modal data.
[0077] The edge computing unit adopts a heterogeneous architecture, including 2 NVIDIA RTX 3060 Embedded GPUs for parallel processing of images and data; 3 Xilinx Zynq UltraScale+ MPSoC FPGAs responsible for real-time data processing and control logic; 4 Cambricon MLU270 NPUs dedicated to deep learning inference; 2 Huawei Ascend 310 dedicated AI acceleration chips for processing specific algorithm tasks. Among them, 2 NVIDIA GPUs and 4 Cambricon NPUs are installed in the high-performance computing module of the edge computing node and are connected to the main board through PCIe 4.0 interfaces; 3 Xilinx FPGAs are distributed at the front end of data acquisition to achieve real-time preprocessing of data; 2 Huawei Ascend 310 chips are integrated on the algorithm acceleration board to work in coordination with other devices. Specifically as follows:
[0078] (1)Feature extraction layer: The ResNet-18 backbone network is used to extract features from RGB images, TOF depth images, millimeter-wave point cloud data, and lidar point cloud data respectively, and convert the original data into feature vectors.
[0079] (2)Multi-head self-attention mechanism: A multi-head self-attention mechanism with 8 heads is adopted, and the dimension of each head is 64. By calculating the attention weights between feature vectors of different modalities, the network can simultaneously focus on multiple feature subspaces and capture long-range dependencies between data.
[0080] (3)Dynamic routing mechanism: Based on the calculation of quantum entanglement similarity, the quantum state inner product formula is adopted:
[0081] Calculate the similarity of features of different modalities, adaptively allocate fusion weights, and achieve feature-level fusion.
[0082] Formula analysis: In quantum mechanics, a quantum state can be represented by a vector, and the inner product is used to measure the similarity between two quantum states. This formula calculates the square of the modulus of the inner product of two quantum states and to obtain the similarity measure between them. In the cross-modal feature fusion network, the features of different modalities are analogized to quantum states, and this formula is used to calculate the quantum entanglement similarity between features to adaptively allocate fusion weights;
[0083] and : respectively represent the "quantum state" vectors formed after encoding two different modality data (such as RGB image feature vectors and lidar point cloud feature vectors), and these vectors contain the feature information of the corresponding modality data;
[0084] : represents the quantum state and The inner product, whose calculation method is to sum the products of corresponding elements (in the vector representation), and its result reflects the similarity degree and directional relationship of two eigenvectors in the feature space;
[0085] : The similarity value between the finally obtained two-modal features, whose value range is between 0 and 1. The closer the value is to 1, the more similar the features of the two modalities are, and a higher weight should be given during the feature fusion process. When a certain modal feature (such as point cloud) is more important for target recognition in the current scene (such as a character occlusion scene), its similarity increases and the corresponding weight automatically increases; conversely, the weight of the secondary modality (such as a low-resolution image) automatically decreases to avoid interference of invalid features on the fusion result.
[0086] (4) Feature enhancement layer: Further enhance the fused features through convolution operations and output the final fused features.
[0087] Implementation example: In a container yard, the system collects container data with a severely reflective surface and partially blocked by goods. Two NVIDIA GPUs first use the ResNet-18 backbone network to extract the eigenvector features of the RGB images collected by 16 Basler industrial cameras and the point cloud data obtained by 2 Velodyne VLP-16 lidars respectively. Subsequently, on 4 Cambrian NPUs, 8 Transformer architectures (hidden dimension 512) calculate the attention weights of different modal features in parallel. Through analysis, it is found that in the character localization task, the importance weight of the lidar point cloud feature for target recognition reaches 0.75; then, 2 Huawei Ascend 310s execute the dynamic routing mechanism, further adjust the weights according to the quantum entanglement similarity formula, increase the lidar point cloud feature weight to 0.85, and correspondingly reduce the RGB image feature weight. Finally, 3 Xilinx FPGAs complete the convolution operation to enhance the features and output the fused features within 500 milliseconds. This feature accurately represents the position and partially visible texture information of the container number characters, providing high-quality data support for subsequent character recognition and analysis.
[0088] [[ID=1_{4}]]2.2. GAN-based noise reduction model:
[0089] The noise reduction model uses a U-Net structure generator and a PatchGAN structure discriminator. The adversarial training uses the LSGAN loss function. The activation functions of the generator and the discriminator are LeakyReLU (slope 0.2), and the momentum parameter of the batch normalization layer is set to 0.9. These parameter configurations have shown good denoising effects in a large number of data denoising experiments. The geometric distortion correction algorithm is based on a diffeomorphic differentiable mapping model. By minimizing the image reprojection error (threshold set to 1 pixel), the mapping parameters are optimized to correct the image distortion, ensuring the accuracy of data preprocessing, removing the noise in multi-modal data, improving the data quality, and providing a reliable input for subsequent processing. Specifically as follows:
[0090] (1) Generator design: The U-Net structure is adopted, which includes 5 downsampling blocks and 5 upsampling blocks. The downsampling blocks extract image features through convolution and pooling operations, and the upsampling blocks restore image details through transposed convolution and skip connections, learning the mapping relationship between the noisy image and the clean image.
[0091] The training of the generator and the discriminator is based on 2 NVIDIA RTX 3060 Embedded GPUs. The generator with the U-Net (5 downsampling + 5 upsampling blocks) structure and the discriminator with the PatchGAN structure are used. During the data processing, the batch size is set to 32 and the learning rate is set to 2e-4.
[0092] Deployment method: 2 NVIDIA GPUs are installed in the deep learning training module of the edge computing node and are connected to the storage device through a high-speed data bus to ensure fast data transmission.
[0093] (2) Discriminator design: The PatchGAN structure is adopted. The input image is divided into multiple local patches, and it is judged whether each patch comes from a real clean image or an image generated by the generator, thereby guiding the generator to generate more realistic images.
[0094] (3) Loss function design: The LSGAN (Least Squares Generative Adversarial Network) loss function is adopted, including the generator loss and the discriminator loss. Through adversarial training, the generator can effectively remove the data noise.
[0095] (4) Training process: The generator and the discriminator are alternately trained, and the network parameters are continuously adjusted until the generator can generate high-quality denoised images.
[0096] (5) Diffeomorphic differentiable mapping model: Let the input image be , and the output corrected image be . The pixel coordinate mapping relationship is established through the differentiable mapping function : , where is the distorted image coordinate. is the corrected coordinate, are learnable mapping parameters (such as translation, rotation, scaling, non - linear distortion coefficients).
[0097] (6) Reprojection error optimization: For each pixel in the image , calculate its corrected coordinate , and reproject it back to the original image coordinate system to obtain . The reprojection error is: ; Minimize the reprojection error through the backpropagation algorithm, and update the mapping parameter to make the error lower than the 1 - pixel threshold: ; When correcting the visual image, synchronize the device attitude information (such as pitch angle, yaw angle) in the reference IMU data, embed the motion parameters measured by the inertial measurement into the mapping model to achieve real - time compensation for dynamic distortion. For example, when the IMU detects that the device vibration causes the camera to tilt by 5°, the algorithm automatically adjusts the rotation component in the mapping parameter to offset the perspective distortion caused by the tilt.
[0098] (7) Initialize the distortion model. Obtain the camera internal parameters (focal length, principal point coordinates) and external parameters (rotation matrix, translation vector) through the Zhang's calibration method, and establish an initial distortion model (such as the Brown - Conrady model).
[0099] Construct a differentiable mapping network. Adopt the convolutional neural network (CNN) architecture (such as the encoder - decoder structure of U - Net), input the distorted image, and output the pixel - level offset<00's calibration method, and establish an initial distortion model (such as the Brown - Conrady model).
[0099] Construct a differentiable mapping network. Adopt the convolutional neural network (CNN) architecture (such as the encoder - decoder structure of U - Net), input the distorted image, and output the pixel - level offset , that is: , where is the learnable mapping network, are the network parameters.
[0100] End - to - end training: Use a paired distorted - corrected image dataset (such as synthetic distorted images + real - scene calibration images), and train with the reprojection error as the loss function to ensure that the offset output by the network can accurately cancel the distortion.
[0101] Real - time correction inference: For the distorted images collected in real - time, predict the coordinate offset of each pixel through the trained network to generate the corrected image, and control the processing delay within 100 ms (to adapt to the real - time operation requirements of the port).
[0102] Implementation example: During the operation period in rainy weather, the RGB images of containers collected by 16 Basler industrial cameras were covered with a large amount of noise due to dense raindrop adhesion and strong light reflection, resulting in a serious decline in image quality and extremely low clarity, which brought great difficulties to subsequent recognition. After these noisy images were input into the denoising model, 2 GPUs worked in parallel. The U-Net generator gradually extracted the noise features and potentially useful information in the images through downsampling blocks, and the upsampling blocks combined with skip connections to gradually restore the image details; the PatchGAN discriminator divided the image into 16×16 patches for true / false discrimination and continuously fed back information to optimize the generator. After 50 epochs of training, the generator effectively removed the raindrop noise, the image entropy value increased significantly from 3.0 to 5.2, the character edges became clear and sharp, and the overall image quality was greatly improved, providing high-quality data for subsequent character recognition and improving the recognition accuracy.
[0103] 3. Character recognition related models:
[0104] 3.1 Improved YOLOv7-tiny object detection model:
[0105] The model is deployed on 4 Huawei Ascend 310 dedicated AI acceleration chips of the edge computing node, adopting a parallel computing architecture. The input size is set to 640×640. After 30% pruning, the number of parameters is reduced from 6.3M to 3.8M, the model volume is greatly reduced, and the inference speed is significantly improved; the 4 Huawei Ascend 310 chips are integrated on the dedicated acceleration board of the edge computing node and are connected to the data processing module through high-speed interfaces to achieve fast data transmission and efficient computing.
[0106] Basic model selection: Using YOLOv7-tiny as the basic model, this model structure is lightweight and suitable for fast inference on edge computing devices.
[0107] Progressive knowledge distillation: Using a pre-trained large object detection model (such as YOLOv7) as the teacher model, setting the temperature parameter to 4, and transferring the soft label knowledge of the teacher model to the student model (improved YOLOv7-tiny) to guide the student model to learn richer feature expressions.
[0108] Dynamic network pruning: According to the requirements of computing resources and detection accuracy, the Taylor expansion formula is used to evaluate the importance of each convolutional kernel, and the model is pruned by 30% to remove redundant connections and parameters, reducing the number of model parameters by 40% and improving the detection speed while ensuring the detection accuracy.
[0109] Introduction of spatio-temporal attention mechanism: On the basis of the Feature Pyramid Network (FPN), a spatio-temporal feature fusion module is added. Through the spatio-temporal feature pyramid structure (including 3 scales, corresponding to the 1 / 4, 1 / 8, and 1 / 16 resolutions of the original image respectively), combined with the spatio-temporal cross-attention module, the changing features of the container number character area at different time steps are captured.
[0110] Analysis: Introduce the analysis of the time dimension in the object detection process to enhance the feature expression ability for fast-moving and small-size character targets, and solve the problem of high miss detection rate of dynamic targets by traditional single-frame detection models.
[0111] Spatio-temporal feature pyramid structure: On the basis of the Feature Pyramid Network (FPN), the feature fusion in the time dimension is added to construct multi-scale and multi-temporal feature expressions;
[0112] Spatio-temporal cross-attention mechanism: By calculating the attention weights of spatial positions and time steps, the key feature areas of dynamic targets are highlighted.
[0113] 3.1.1. The spatio-temporal feature pyramid structure is as follows:
[0114] Multi-scale spatial features, following the FPN architecture of YOLOv7-tiny, output 3 scales of spatial feature maps (resolutions are 1 / 4, 1 / 8, 1 / 16 of the original image respectively), corresponding to detecting large, medium, and small-size character targets.
[0115] Time dimension expansion. For the feature maps of each spatial scale, introduce the historical features of continuous V frames (take V = 5) to form a spatio-temporal feature cube (spatial dimension H×W, time dimension T); perform convolution operations on the spatio-temporal feature cube through 3D convolution (kernel size 3×3×3, time step 1) to extract cross-frame motion features (such as the displacement trajectory of characters in continuous frames).
[0116] 3.1.2. The spatio-temporal cross-attention module is as follows:
[0117] Attention calculation logic: (1) Spatial attention: Within a single-frame feature map, generate a spatial attention weight map through convolution to highlight the areas where characters may exist (such as the fixed positions of container numbers); (2) Time attention: At the same spatial position in continuous frames, calculate the feature similarity to generate time attention weights (such as when a character appears at a certain position in 3 consecutive frames, the weight increases).
[0118] Cross-fusion: Multiply the spatial and time attention weights to obtain the spatio-temporal joint attention weight. The formula is: Among them, represents element-wise multiplication, and the final weight is used to adjust the fusion intensity of spatio-temporal features. Denote the spatial attention weight matrix, with dimension H×W, representing the importance of each spatial position (i, j) in the image for object detection (the larger the value, the more important). Denote the temporal attention weight matrix, with dimension T×H×W, representing the stability of the features at the same spatial position (i, j) in consecutive T frames in the temporal dimension (the larger the value, the more stable).
[0119] Combined with LSTM: At the top layer of the spatio-temporal feature pyramid (small-size feature maps), connect a 2-layer LSTM network (hidden layer dimension 256) to perform sequential modeling on the temporal features and capture the motion patterns of characters over a long time span (such as the continuous displacement when a truck passes at high speed).
[0120] 3.1.3. Dynamic Object Trajectory Tracking:
[0121] Through spatio-temporal feature fusion, the module can track the motion trajectory of characters in 5 consecutive frames of images, calculate their displacement vectors in the images (with an accuracy of ±2 pixels), and combine the motion data of the IMU (such as vehicle acceleration) to achieve precise positioning of high-speed moving targets.
[0122] Enhancement of small target features. For small character targets with a pixel area <50px², traditional single-frame detection is prone to missed detection due to insufficient features. The spatio-temporal feature fusion module accumulates weak features from multiple frames (such as blurred pixel points at the same position in 2 consecutive frames) to form detectable strong response regions in the spatio-temporal feature cube, improving the detection accuracy of small targets.
[0123] Improvement of anti-interference ability. Filter out temporary noises (such as single-frame pseudo-features caused by sudden changes in light) through the temporal attention mechanism, and only retain the features that stably exist in multiple consecutive frames, effectively suppressing false detections in dynamic scenarios. For example, when a truck passes through a strong light area, the module can exclude the interference caused by single-frame overexposure through feature comparison of 3 consecutive frames and maintain the stability of the detection results.
[0124] 3.1.4. Implementation Process and Parameter Configuration:
[0125] Input data: 5 consecutive frames of images (resolution 640×640), extract spatial features through the backbone network (CSPDarknet); IMU data (acceleration, angular velocity) are used as prior temporal motion information and input into the spatio-temporal cross-attention module.
[0126] Feature processing: The spatio-temporal feature pyramid generates spatio-temporal feature maps at 3 scales (sizes are 160×160×5, 80×80×5, 40×40×5); the spatio-temporal cross-attention module outputs the weighted feature maps, which are connected to the detection head (classification layer + regression layer) of YOLOv7-tiny.
[0127] Semantic Association Analysis of Graph Neural Network: The character region detection results are constructed into a graph structure, where the nodes represent character candidate boxes and the edges represent the semantic relationships between characters. A graph convolutional network (GCN) is adopted, which contains 2 graph convolutional layers, the output dimension of each layer is 128, the activation function is ReLU, and the node features are updated through graph convolutional operations to mine the semantic associations between characters, filter out the candidate boxes that do not conform to the rules, and output accurate character region localization results.
[0128] Implementation Example: A container truck traveling at a speed of 60 kilometers per hour enters the recognition area. A computing cluster composed of 4 Huawei Ascend 310 chips receives 5 consecutive frames of image data. The improved YOLOv7-tiny model utilizes the rich features learned by progressive knowledge distillation and quickly identifies the character candidate area within 80 ms. The spatio-temporal attention mechanism associates through 3-level feature pyramids (1 / 4, 1 / 8, 1 / 16 scales) and 5-frame historical data to capture the rapid displacement of characters in the image and accurately track the character movement trajectory. The semantic association analysis module of the graph neural network constructs the candidate boxes into a graph structure, analyzes them through 2-layer graph convolutional network (output dimension 128), combines with the pre-constructed semantic rules of container number characters, filters out the mis-detection areas that do not conform to the rules, and finally accurately outputs the character region bounding box. In the high-speed movement scenario, it effectively improves the detection accuracy, successfully identifies the container number, and ensures the accurate recording of logistics information.
[0129] In the pre-training stage of the improved YOLOv7-tiny object detection model, based on the COCO dataset, 5000 additional container number images with different illuminations, angles, and sharpness are added to expand the training set. A multi-modal pre-training strategy based on contrastive learning is adopted to construct an image-point cloud-text contrastive learning task. The contrastive loss is calculated through the InfoNCE loss function, and the temperature parameter is set to 0.07. This value has been verified through multiple groups of experiments and can achieve the optimal effect in balancing feature discrimination and model stability. The positive-negative sample ratio is set to 1:3, which can effectively improve the model's ability to extract container number features and generalization performance.
[0130] Among them, the multi-modal pre-training strategy refers to, in the initial stage of model training, through fusing multiple modal data (such as images, point clouds, texts) for joint learning, enabling the model to capture the associated features between different modalities, thereby enhancing the generalization ability for complex scenarios. This strategy is mainly applied to the pre-training stage of the improved YOLOv7-tiny object detection model. The specific implementation method is as follows:
[0131] Data composition: 5,000 new RGB images of container numbers with different lighting (strong light / weak light), angles (tilted / straight ahead), and clarity (clear / blurred) were added; point cloud data from millimeter-wave radar and lidar were simultaneously collected to characterize the three-dimensional geometric features of the container; and the character semantic information of the container number (such as the ISO 6346 standard format rules) was extracted.
[0132] Through cross-modal comparative learning, the association between image pixels, point cloud spatial distribution and text semantics is established, so that the model can learn to understand the characteristics of container numbers from multiple dimensions.
[0133] Contrastive learning task design: forces the model to learn similar features of the same target in different modalities, while widening the feature distance between different targets. The details are as follows:
[0134] Positive sample pair: image-point cloud-text data combination of the same container number;
[0135] Negative sample pairs: any combination of modal data with different container numbers (positive and negative sample ratio is 1:3)
[0136] Loss function: InfoNCE loss function (noise contrast estimation loss) is used, and the calculation formula is:
[0137] ;
[0138] Where N is the number of query samples; K is the number of negative samples in the query samples, is the feature vector of the i-th query sample, is the feature vector of the i-th positive sample (point cloud / text features of the same sample), is the feature vector of the i-th negative sample (arbitrary modal features of different samples), =0.07 is the temperature parameter, which controls the concentration of characteristic distribution.
[0139] During the knowledge distillation stage, the teacher model is a complete YOLOv7 model with a temperature parameter set to 4. This temperature value is most effective in balancing knowledge transfer efficiency and model convergence speed, and soft labels are used to guide student model learning. Dynamic network pruning removes unimportant connections at a ratio of 30%, reducing the number of model parameters by 40%, significantly improving detection speed while maintaining high detection accuracy.
[0140] 3.2 Transformer-based cGAN character segmentation model:
[0141] The model runs on 2 Cambrian MLU270 NPUs (with a computing power of 128 TOPS). The Transformer part adopts a 4-head multi-head self-attention mechanism, and the adaptive receptive field can be dynamically adjusted within the range of 3×3 - 11×11 to adapt to different character shapes. Among them, 2 Cambrian MLU270 NPUs are installed in the AI inference module of the edge computing node and are closely connected to the data storage and transmission module to ensure efficient data processing and fast transmission. Specifically as follows:
[0142] (1) Generator design: An adaptive receptive field module is introduced in the generator to dynamically adjust the receptive field size according to the stroke complexity and morphological features of the characters (range from 3×3 to 11×11). At the same time, the multi-head self-attention mechanism is used to capture the long-range dependencies of the characters and improve the segmentation accuracy.
[0143] (2) Discriminator design: A PatchGAN structure similar to the GAN-based denoising model is adopted to judge the authenticity of the segmentation result.
[0144] (3) Post-processing of topological structure analysis: Calculate topological invariants such as the Euler number and the number of holes in the connected region of the characters (Euler number threshold -1 to 1, number of holes threshold 0 to 2), judge the integrity of the characters, repair the broken characters, and remove the noise points.
[0145] (4) Training process: Through adversarial training, continuously optimize the parameters of the generator and the discriminator to enable the generator to accurately segment the characters.
[0146] Implementation example: When a container with a long service life, rusty surface and severely adhered characters enters the recognition range, 2 Cambrian MLU270 NPUs work together. The generator of the cGAN model first detects the complex stroke structure of the adhered characters through the adaptive receptive field module, automatically adjusts the receptive field to 9×9 to better capture the character details; the multi-head self-attention mechanism fully plays its role, accurately identifies the character boundaries, and initially segments the character regions. However, some characters are broken. The topological structure analysis module calculates that the Euler number of the connected region of the characters is -1 and the number of holes is 1, and determines them as broken characters. Morphological operations are used to repair them and remove the surrounding noise points. After a series of processes, a clear and complete character segmentation image is finally output within 300 milliseconds, providing high-quality input for subsequent character recognition. Even for such severely damaged characters, it can be effectively segmented, laying a foundation for accurate recognition.
[0147] 3.3, Lightweight CNN-Transformer-Capsule Network Integrated Model:
[0148] The lightweight CNN (MobileNetV3), Transformer, and capsule network are deployed in a collaborative processing architecture of 2 Cambrian MLU270 NPUs and 2 NVIDIA RTX3060 Embedded GPUs. The PPO algorithm is used for dynamic weight fusion, and the batch size is set to 64. Among them, the 2 Cambrian MLU270 NPUs are responsible for the computing tasks of the Transformer and capsule network, and are installed in the AI acceleration module of the edge computing node; the 2 NVIDIA RTX3060 Embedded GPUs process the tasks related to MobileNetV3 and are deployed in the image processing module. The two achieve data sharing and collaborative computing through a high-speed data bus. Specifically as follows:
[0149] (1)Lightweight CNN design: The MobileNetV3 architecture is adopted, and its inverted residual structure and SE attention mechanism are used to efficiently extract local character features and reduce the model's computational complexity.
[0150] (2)Transformer design: A 6-layer encoder structure is adopted, and the multi-head self-attention mechanism is used to capture the global semantic information of characters and understand the context relationship between characters.
[0151] (3)Capsule network design: 16 capsules are set, and the dimension of each capsule is 32, which is used to extract the spatial hierarchical structure of characters and capture the relative position and pose information between character components.
[0152] (4)Bidirectional knowledge distillation: The adversarial knowledge transfer mechanism is adopted. ResNet50 is used as the teacher model to transfer knowledge to the integrated model. The student model optimizes the feature expression of the teacher model through adversarial training feedback. At the same time, a contrastive learning mechanism is introduced to enhance the model's ability to distinguish character features. The adversarial training process is as follows:
[0153] Feature extraction: The teacher model extracts features from the input image (such as the output of the pooling layer of ResNet50); the student model extracts features from the same image (the fused features of the integrated model).
[0154] Discriminator training: Input the mixed features , and the discriminator learns to distinguish the source. The loss function is cross-entropy: , where is the output probability of the discriminator (the closer to 1, the more likely the feature comes from the teacher model), and E is the mathematical symbol for "taking the average according to the distribution of samples" (mathematical expectation, representing the average operation on random variables (sample distributions)). Generalize the loss of a single sample to the statistical law of the entire dataset, so that the loss function can reflect the performance of the model on all data.
[0155] Student model training forces the student features to approximate the teacher features through adversarial loss, and the loss function is as follows: That is, the student model tries to deceive the discriminator so that is close to 1.
[0156] Indirect optimization of the teacher model. During the adversarial process, the feature distribution of the teacher model will be adjusted by the feedback of the discriminator, improving the ability to distinguish features of difficult samples (such as blurred characters).
[0157] In two-way knowledge distillation, the specific process of knowledge transfer is as follows:
[0158] Feature alignment loss combines the traditional knowledge distillation loss (soft label loss ) and the adversarial loss to form a comprehensive loss: , where (weights that can be determined through experiments) balances semantic knowledge transfer and adversarial feature alignment.
[0159] Discriminator decision boundary. Ideally, the discriminator cannot distinguish between teacher and student features, that is , indicating that the student model has fully learned the feature distribution of the teacher.
[0160] (5) Dynamic weight fusion strategy: Based on the reinforcement learning optimization framework, trained using the PPO algorithm, the state space includes 5 dimensions such as the prediction confidence and calculation time consumption of each sub-model, and the action space is the weight allocation of each model (0 - 1). Design the reward function to dynamically adjust the output weights of each model with the recognition accuracy and calculation efficiency as the goals, achieving accurate classification and recognition of characters. In the formula, represents the recognition accuracy of the model for container number characters, with a value range between 0 and 1, used to measure the correctness of the model's recognition result; represents the time taken for the model to process one recognition task, and the unit can be time units such as seconds; then converts the time consumption into an index in the same direction as the accuracy, that is, the shorter the time consumption, the larger the value of this part.
[0161] Through this reward function, during the training process of the reinforcement learning algorithm, according to the recognition accuracy and calculation time consumption obtained after each decision (weight allocation of each sub-model), the corresponding reward value will be calculated, guiding the model to learn the optimal weight allocation strategy for each sub-model, so that in different input scenarios, the model can ensure a high recognition accuracy while minimizing the calculation time as much as possible, achieving fast and accurate recognition of container number characters.
[0162] Implementation example: For the recognition of a container number, MobileNetV3 on 2 NVIDIA GPUs quickly extracts local texture features of characters, such as the subtle wear marks and unique printing textures at the edges of characters; the Transformer on 2 Cambrian NPUs captures global semantic information through 6 layers of encoders to understand the sequential relationship and overall structure among characters; the capsule network deeply analyzes the spatial hierarchy of character components to determine the inclination angle and relative position of characters. During the two-way knowledge distillation process, the pre-trained ResNet50 is used as the teacher model to guide the integrated model to learn. During the recognition process, when encountering easily confused characters such as '9' and '6', the dynamic weight fusion strategy adjusts the weight of the Transformer from 0.3 to 0.5 in real time according to the prediction confidence and calculation time consumption of each sub-model, enabling it to play a greater role in comprehensive judgment. Finally, the integrated model accurately recognizes the container number within 120 ms, effectively avoiding recognition errors caused by character confusion and ensuring the accurate entry of container information.
[0163] When constructing the character structure knowledge graph, the Neo4j graph database is used to manually label 2000 groups of container number character samples, clarify node attributes (including character type, size, spacing, etc.) and edge attributes (covering spatial relationships, semantic relationships, etc.), and predefined constraints such as the character spacing being 1.2 - 1.5 times the character height and the aspect ratio being 0.8 - 1.2. These parameters are obtained based on industry standards for container number characters and statistics from a large number of actual samples, providing a precise screening basis for subsequent character region localization.
[0164] Summary: In the hybrid model, the lightweight CNN adopts the MobileNetV3 architecture, the Transformer has 6 layers of encoders, the capsule network contains 16 capsules with a dimension of 32, the two-way knowledge distillation adopts an adversarial knowledge transfer mechanism, the teacher model guides the student model to learn through soft labels, the student model optimizes the teacher model through adversarial training feedback, the discriminator of the adversarial training uses a multi-layer perceptron (2 hidden layers, 256 neurons in each layer), the dynamic weight fusion strategy is optimized based on a reinforcement learning framework and uses the PPO algorithm, and the reward function is 'Recognition accuracy × 0.8 + Computational speed × 0.2'. The weight setting of this reward function has been verified through multiple groups of comparison experiments and can effectively balance recognition accuracy and computational efficiency. The policy network and value network use multi-layer perceptrons (hidden layer size 256, training batch size 64) to dynamically adjust the model weights and improve the overall recognition performance.
[0165] 4. Intelligent verification and decision-making related models:
[0166] 4.1 Dynamic Bayesian network verification and decision-making model:
[0167] The dynamic Bayesian network adopts an inference mechanism that integrates evidence theory and Bayesian optimization. The basic probability assignment function of evidence theory is determined by statistical analysis of historical data. The Bayesian optimization acquisition function is the expected improvement (EI) function. Format compliance verification is achieved through a probabilistic finite state automaton. The state transition probability is statistically calculated based on historical data, and the threshold is set to 0.8. This threshold can effectively distinguish between compliant and non-compliant container number formats. Spatiotemporal consistency verification uses a spatiotemporal graph convolutional network (with 3 spatiotemporal convolutional layers, a time step of 5, and a spatial neighborhood size of 3×3). The targeted verification enhancement strategy adopts a meta-verification framework, which includes 5 verification enhancement strategies (such as increasing the sampling times, switching the recognition model, etc.). The genetic algorithm is used to optimize the strategy selection probability to ensure the accuracy and reliability of the verification decision. The recognition results are multi-dimensionally verified through probabilistic inference to evaluate the credibility of the results and make reasonable decisions for abnormal situations. Specifically as follows:
[0168] The model runs on a distributed computing cluster composed of 3 Advantech UNO-3083 intelligent controllers (Intel Core i7-8550U processors, 16GB DDR4 memory). The dynamic Bayesian network contains 15 nodes (5 verification dimensions + 10 state nodes), and data sharing and collaborative computing are achieved through network communication. The 3 Advantech intelligent controllers are respectively deployed in the channel monitoring center, the equipment management room, and the edge computing station, and are connected through a gigabit local area network to form a distributed computing architecture to ensure efficient data transmission and collaborative processing.
[0169] Network structure construction: According to the verification requirements of container number recognition, a dynamic Bayesian network containing multiple nodes and edges is constructed. The nodes represent different verification dimensions (such as format compliance, spatiotemporal consistency, etc.) and recognition results, and the edges represent the probabilistic dependence relationships between the nodes.
[0170] (1) Parameter initialization: By analyzing historical container number recognition data, the prior probability distribution of each node is statistically calculated to initialize the network parameters. For format compliance verification, a probabilistic finite state automaton containing 12 states is constructed, and the state transition probability matrix is obtained through training with 100,000 historical data, and the threshold is set to 0.8.
[0171] (2) Evidence theory fusion: Combine evidence theory with the Bayesian network, and obtain the basic probability assignment function of evidence theory through statistical analysis of historical data to fuse evidence from different sources and improve the accuracy of inference.
[0172] (3) Bayesian optimization: Adopt the expected improvement (EI) function to dynamically update the network structure and parameters according to real-time verification data, and optimize the inference ability of the network.
[0173] (4) Decision mechanism design: When the recognition result passes the verification, confirmation information is output; when an anomaly occurs, a targeted enhancement strategy of the meta-verification framework is triggered.
[0174] Implementation example: The system identifies the container number "ABC1234X", and a cluster composed of 3 Advantech intelligent controllers works collaboratively to verify this recognition result. In the format compliance verification, the probability finite state automaton judges that the probability of the number conforming to the ISO6346 standard format is 0.9 according to the state transition probability matrix trained from historical data; the spatio-temporal consistency verification module analyzes its motion trajectory in the past 5 frames, and the verification probability reaches 0.95; however, in the multi-modal data consistency verification, when calculating the Wasserstein distance of the feature vectors of the RGB image, depth image, and point cloud data, a large difference is found, and the verification probability is only 0.4. The dynamic Bayesian network synthesizes the information and probability relationship of each node for reasoning, and finally judges that there is an anomaly in this recognition result, triggering the multi-perspective supplementary acquisition strategy of the meta-verification framework, and taking measures to further confirm the accuracy of the recognition result.
[0175] 4.2 Implementation of the meta-verification framework:
[0176] The meta-verification framework is started, and the strategy library contains 5 strategies: OCR post-processing (Python script calls Tesseract-OCR), multi-perspective data supplementary acquisition (controls the DJI Matrice300RTK drone), historical data comparison (queries the MySQL database), sensor parameter recalibration, and manual intervention prompt.
[0177] The genetic algorithm (population of 50 individuals, 100 generations of iteration) optimizes the strategy selection. The fitness of each strategy is calculated according to historical data, and finally it is determined to preferentially adopt the multi-perspective data supplementary acquisition strategy. The system sends instructions to the DJI drone to acquire images from 3 angles such as the side and top view.
[0178] After the supplementary acquired data is processed by the edge computing device, it is re-recognized and verified, and it is confirmed that the original result is incorrect, and the correct number is "ABC8765B".
[0179] The verified container number and related information are output to the tallying system. At the same time, through the blockchain deposit evidence technology based on zero-knowledge proof, combined with the data processing method of homomorphic encryption, non-tamperable deposit evidence is carried out; among them, zero-knowledge proof adopts the zk-STARKs technology to generate a 1024-bit proof, and the verification time is less than 100 ms. Homomorphic encryption adopts the BGV homomorphic encryption scheme, and the modulus q = 2 40 and the plaintext space is Z_2 10, supporting ciphertext operations to ensure data privacy. The blockchain evidence storage adopts the side-chain technology. The side-chain uses the Practical Byzantine Fault Tolerance (PBFT) consensus mechanism (the number of consensus nodes is 5, and the tolerance number of Byzantine nodes is 1). The evidence storage data is organized in a Merkle tree structure (depth 10). The side-chain block size is 4MB, and the confirmation time is less than 2 seconds. The verified container numbers and related information are output to the tallying system and stored as evidence to ensure the security and traceability of the data.
[0180] 5. Implementation of system self-learning related algorithms:
[0181] The construction of the dynamic training dataset adopts the active learning strategy, measures uncertainty through the prediction entropy value (threshold 0.9), and selects the most valuable error cases. The online learning strategy is based on the Model-Agnostic Meta-Learning (MAML) framework. The number of tasks in the meta-training stage is 1000, the meta-learning rate is 0.01, the number of gradient updates in the meta-testing stage is 2 times. The parameter update of the dynamic Bayesian network adopts the variational inference method, constructs the variational lower bound by minimizing the KL divergence, and the parameter optimization of the reinforcement learning reward function adopts the Trust Region Policy Optimization (TRPO) algorithm, with the trust region size of 0.01 and the optimization step size of 0.5. Continuously optimize the system performance and improve the recognition accuracy.
[0182] The implementation examples of the above models and algorithms are closely combined with the actual port operation scenarios, and the full-process operations from data collection to result verification are detailedly demonstrated to ensure the integrity and operability of the overall technical solution, providing a practical technical reference for the container number tallying recognition in the field of intelligent logistics.
[0183] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.
Claims
1. A method for tallying and identifying container numbers, characterized in that, The recognition method includes the following steps: S1. Deploy multi-modal perception devices and perform initial calibration, and configure reinforcement learning parameters; Pre-train an improved YOLOv7-tiny object detection model, a Transformer-based cGAN character segmentation model, and a hybrid recognition model integrated with a lightweight CNN-Transformer-capsule network, construct a character structure knowledge graph, and initialize the parameters of the dynamic Bayesian network; S2. Through the multi-modal perception devices with a self-calibration mechanism, combined with multi-sensor fusion positioning, synchronously collect multi-modal data on multiple sides of the container; S3. Based on the improved YOLOv7-tiny object detection model, fuse the spatio-temporal attention mechanism, the character structure knowledge graph, and the semantic association analysis based on the graph neural network to detect the character regions in the multi-modal data; S4. Use the Transformer-based cGAN character segmentation model, combined with morphological post-processing of topological structure analysis; through the adaptive normalization network, perform character image segmentation and normalization; S5. Adopt a hybrid recognition model integrated with a lightweight CNN-Transformer-capsule network, combined with the bidirectional knowledge distillation technology and the dynamic weight fusion strategy, perform character classification, and output the recognition result; S6. Construct a verification decision model based on the dynamic Bayesian network, combined with a multi-dimensional verification strategy, to evaluate the credibility of the recognition result; when an abnormal verification result occurs in a certain item, start a targeted verification enhancement strategy.
2. The method for container number tallying and identification according to claim 1, wherein: In S1, the multi-modal perception devices include a distributed stereo vision array, a millimeter-wave radar, and a lidar, and a timestamp synchronization mechanism is used to calibrate the time of the multi-modal perception devices; Among them, the distributed stereo vision array is calibrated for internal and external parameters by the Zhang's calibration method; The millimeter-wave radar and the lidar use the NDT algorithm for point cloud registration.
3. The method for tallying and identifying container numbers according to claim 2, characterized in that: In S2, the acquisition of multi-modal data is through the distributed stereo vision array with a self-calibration mechanism, combined with the multi-sensor fusion positioning of the millimeter-wave radar and the lidar, synchronously collect the RGB images, TOF depth images, millimeter-wave point cloud data, and lidar point cloud data on multiple sides of the container, and use the dynamic threshold triggering and adaptive exposure control strategy based on the deep Q network to obtain multi-modal data containing the container number region; Among them, the dynamic threshold triggering strategy uses a double-layer Q network architecture to achieve multi-objective optimization, including a target network and a real network. The parameters of the target network are updated every L steps, and the real network uses the Adam optimizer, where L > 100; The adaptive exposure control strategy dynamically adjusts the exposure time by calculating the image entropy value.
4. A method for tallying and identifying container numbers according to claim 1, characterized in that: In S2, the self-calibration mechanism uses a checkerboard calibration board and an IMU deployed in the transportation channel to perform spatio-temporal calibration on the distributed stereo vision array, millimeter-wave radar, and lidar in real time based on the Kalman filter algorithm; among them, the state transition matrix of the Kalman filter is determined by the motion model of the IMU, and the observation matrix is determined by the corner detection results of the checkerboard; The distributed stereo vision array adopts a non-uniform redundancy design, dynamically adjusts the sensor layout density according to the recognition requirements of the key areas of the container, and the sensor density in the key areas is 2 to 3 times that of the non-key areas.
5. A method for tallying and identifying container numbers according to claim 1, characterized in that: Before character area detection, a cross-modal feature fusion network based on the Transformer architecture is adopted. Combining with a dynamic routing mechanism, multi-source data is fused at the feature level. A noise reduction model based on the generative adversarial network and a geometric distortion correction algorithm based on deep learning are used to preprocess the fused data. Among them, the cross-modal feature fusion network contains a multi-head self-attention mechanism. The dynamic routing mechanism adaptively allocates fusion weights by calculating the quantum entanglement similarity of different modal features, and the similarity calculation uses the quantum state inner product formula. The noise reduction model consists of a generator and a discriminator. The generator adopts a U-Net structure, which contains 5 downsampling blocks and 5 upsampling blocks. The discriminator adopts a PatchGAN structure, and the loss function for adversarial training is LSGAN. The geometric distortion correction algorithm adopts a differentiable mapping model based on diffeomorphism, and optimizes the mapping parameters by minimizing the image reprojection error. The threshold of the reprojection error is set to 1 pixel. Both the generator and the discriminator of the noise reduction model adopt the LeakyReLU activation function.
6. The method for tallying and identifying container numbers according to claim 1, wherein: In S3, the improved YOLOv7-tiny model adopts a collaborative optimization architecture of progressive knowledge distillation and dynamic network pruning. The spatio-temporal attention mechanism captures the change features of the character area at different time steps through the spatio-temporal feature pyramid structure. The feature pyramid contains 3 scales, corresponding to 1 / 4, 1 / 8, and 1 / 16 resolutions of the original image respectively. The semantic association analysis of the graph neural network adopts a graph convolutional network, which contains 2 graph convolutional layers, the output dimension of each layer is 128, and the activation function is ReLU. The dynamic update of the character structure knowledge graph is realized through incremental learning. After each recognition result is updated, the confidence scores of the knowledge graph are recalculated, and the nodes and edges with confidence scores lower than 0.5 are deleted.
7. A method for tallying and identifying container numbers according to claim 1, characterized in that: In S4, through an adaptive normalization network based on a differentiable deformation module, the segmentation and normalization of character images are realized. The differentiable deformation module adopts a hybrid deformation model combining bilinear interpolation and thin plate spline transformation. The adaptive normalization network adopts an attention-guided spatial transformation network. The attention weight map is generated through the Softmax function, and the parameters of the spatial transformation network are learned through backpropagation.
8. A method for tallying and identifying container numbers according to claim 1, characterized in that: In S5, bidirectional knowledge distillation adopts an adversarial knowledge transfer mechanism. The teacher model guides the student model to learn through soft labels, and the student model optimizes the feature expression of the teacher model through adversarial training feedback. The discriminator for adversarial training adopts a multi-layer perceptron, which contains 2 hidden layers, with 256 neurons in each layer. The dynamic weight fusion strategy adopts a reinforcement learning optimization framework, dynamically adjusts the model weights with the recognition accuracy and computational efficiency as the optimization objectives, and the reward function of reinforcement learning is recognition accuracy × 0.8 + computational speed × 0.
2.
9. The method for tallying and identifying container numbers according to claim 1, wherein: In S6, the dynamic Bayesian network adopts an inference mechanism that combines the evidence theory and Bayesian optimization. The basic probability assignment function of the evidence theory is obtained through historical data statistics, and the acquisition function of Bayesian optimization is the expected improvement function; The targeted verification enhancement strategy adopts a meta-verification framework, which automatically selects the optimal verification enhancement scheme according to the verification anomaly type. The meta-verification framework includes 5 verification enhancement strategies, and the strategy selection probability is optimized by a genetic algorithm; The format compliance verification is implemented through a probabilistic finite state automaton, and the state transition probability of the automaton is obtained through historical data statistics; The spatio-temporal consistency verification is constructed using a spatio-temporal graph convolutional network. The graph convolutional network includes 3 spatio-temporal convolutional layers, the time step is set to 5, and the spatial neighborhood size is 3×3.
10. The method for container number tallying and identification according to claim 1, wherein This method further includes the following steps: S7. Output the verified container number and related information to the tallying system. At the same time, through the blockchain deposit evidence technology based on zero-knowledge proof, combined with the data processing method of homomorphic encryption, perform non-tamperable deposit evidence; Among them, the zero-knowledge proof adopts the zk-STARKs technology; The homomorphic encryption adopts the BGV homomorphic encryption scheme, which supports addition and multiplication operations on ciphertexts; The blockchain deposit evidence uses the side chain technology for high-frequency transaction processing; The side chain of the blockchain deposit evidence adopts the practical Byzantine fault tolerance consensus mechanism.
Citation Information
Patent Citations
End-to-end container number detection and identification method based on instance segmentation
CN117854076A
Anti-noise welding seam feature recognition method and system based on lightweight neural network
CN118674985A
Continual text recognition using prompt-guided knowledge distillation
US12033408B1
Character recognition method and apparatus, device, medium, and product
WO2023078070A1
Cited By
Baking temperature adjusting method, system and equipment for coating production line and medium
CN120714874A
A method, system, device, and medium for adjusting bake temperature of a painting line
CN120714874B
Data identification method and system
CN120726361A
Information processing method for port container number tallying
CN120782401A
An information processing method for port container sorting
CN120782401B