Container automatic identification and positioning method and system based on end-to-end multi-task learning algorithm
By using end-to-end multi-task learning algorithms and multimodal data fusion, the synchronous output of the 6-DOF pose of the container and the 3D coordinates of the keyhole was achieved. This solved the problems of visual blind spots and robustness in automatic container identification and positioning, improved the identification accuracy and robustness, reduced the system complexity, and ensured real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for automatic identification and positioning of containers suffer from problems such as visual blind spots, difficulty in aligning fasteners, insufficient robustness, high computational complexity, and poor real-time performance, making it difficult to achieve high-precision automatic identification and positioning of multi-layer containers.
Employing an end-to-end multi-task learning algorithm, this method achieves synchronous output of the 6-DOF pose of a container and the 3D coordinates of a keyhole through multimodal data input and a dual-branch decoder. A multimodal adaptive fusion mechanism and a hierarchical fusion strategy are designed, and cross-modal fusion is performed at each layer of the feature pyramid using a self-attention mechanism. Finally, lightweight model transfer is achieved through a knowledge distillation transfer method.
It improves the accuracy and robustness of automatic container identification and positioning, reduces system complexity, ensures real-time performance and autonomous learning capabilities, supports keyhole opening and closing status recognition, and solves the shortcomings of multi-module algorithm error accumulation and low data transmission efficiency.
Smart Images

Figure CN121811380A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of logistics automation and computer vision technology, and in particular relates to an automatic container identification and positioning method and system based on an end-to-end multi-task learning algorithm. Background Technology
[0002] As a core pillar industry supporting global logistics and trade, the port container crane industry has experienced rapid development in recent years, driven by technological innovation, market demand, and policy. Container stackers, as core equipment in port logistics, combine handling, stacking, and loading / unloading functions. With their strong adaptability, they cover diverse scenarios such as global trade and cross-border supply chains, demonstrating significant market potential. However, traditional stacker operation suffers from the following technical pain points in empty container handling operations within port logistics scenarios: reliance on manual adjustment of the spreader's twistlocks to align with the container's lifting holes by the operator, resulting in blind spots, difficulty in aligning the latches, and the need for repeated calibration, all of which demand high operator skills. Therefore, the development of a multi-layer container automatic identification and positioning system has significant industry demand and practical implications.
[0003] Accurate automatic identification and positioning of multi-layer containers is the foundation for realizing the intelligentization of port lifting equipment such as stackers. It is necessary to simultaneously meet the requirements of high-precision identification, millimeter-level positioning, multi-layer stacking analysis, and robustness in harsh environments.
[0004] Chinese invention patent CN118083809A proposes a method, system, terminal, and storage medium for controlling container front-end lifting. This method combines image recognition technology, distance sensors to collect specific distance data, and positioning and navigation technology to achieve full-process control of container front-end lifting, and can also realize 3D display, improving visualization effects. However, this method requires complex collaborative work between its modules, multi-stage processing introduces error accumulation and delays, reliance on rule-based modules leads to insufficient robustness, and complex system integration increases engineering costs.
[0005] Chinese invention patent CN111899301A discloses a deep learning-based 6D workpiece pose estimation method. The core shortcomings of the image-point cloud fusion module in this method lie in data alignment accuracy, sensor noise robustness, computational complexity, and engineering adaptability. In practical applications, it is necessary to improve module performance by designing a multimodal fusion network architecture and introducing adaptive environmental compensation mechanisms (such as illumination normalization and point cloud denoising). Furthermore, the computational complexity of dense fusion networks conflicts with real-time performance, making it difficult to guarantee algorithm computational efficiency.
[0006] Chinese invention patent CN118429421A discloses a dual-modal fusion 6D pose estimation method based on Masked Point-Transformer. This method makes up for the shortcomings of a single sensor by dual-modal fusion. However, reverse analysis shows that single vision is limited by texture, lighting and depth ambiguity, and single laser point cloud is limited by sparse data and semantic missing. Both are insufficient in robustness, generalization ability and information integrity in complex environments.
[0007] The paper "Research on Container Pose Measurement Based on Binocular Vision" uses a binocular camera for container pose recognition, but is limited by the limitations of depth measurement, texture dependence, computational complexity, and especially the motion blur problem in moving scenes. Summary of the Invention
[0008] The purpose of this invention is to provide a method and system for automatic identification and positioning of containers based on an end-to-end multi-task learning algorithm, which can output the 6-DOF pose of the container and the 3D coordinates of the lock hole in a single inference.
[0009] To achieve the above objectives, the present invention employs the following technical solution:
[0010] This invention provides an automatic container identification and positioning method based on an end-to-end multi-task learning algorithm, comprising:
[0011] Acquire multimodal data of containers and perform spatiotemporal synchronization;
[0012] The multimodal data is input into the multimodal feature pyramid backbone network to extract feature data of different modalities, and hierarchical cross-modal feature fusion is performed to obtain hierarchical fused features;
[0013] A dual-branch model is constructed, and the 6-DOF pose of the container and the 3D coordinates of the keyhole are predicted based on the hierarchical fusion features.
[0014] Preferably, the acquisition of container multimodal data and spatiotemporal synchronization includes:
[0015] Multimodal data of the container were acquired using an RGBD camera, LiDAR, IMU, and infrared thermal imaging sensor, respectively.
[0016] Hardware triggering ensures consistent sampling times for multimodal data, achieving time synchronization, and linear interpolation algorithms are used to align asynchronous data.
[0017] Spatial synchronization is achieved by establishing a mapping relationship between the local coordinate system and the global coordinate system through spatial coordinate transformation.
[0018] Preferably, the multimodal feature pyramid backbone network includes:
[0019] The first branch is used to process RGB images and depth maps acquired by the RGBD camera and extract a set of multi-scale feature maps.
[0020] The second branch is used to process the lidar projection map and extract feature maps.
[0021] The third branch is used to process infrared thermal imaging data to obtain feature maps.
[0022] Preferably, the first branch uses a CNN backbone network ConvNeXt as a feature extractor, the ConvNeXt network comprising There are modules with different downsampling rates, and each module includes depthwise separable convolution, layer normalization, and GELU activation function;
[0023] The first branch ultimately yields a multi-scale feature map set { }
[0024] Preferably, the point cloud in the 3D space collected by the lidar is projected onto a plane aligned with the RGB image to form a projection image with multiple feature channels.
[0025] The second branch uses sparse 3D convolution to extract features from the projection map.
[0026] Preferably, the third branch fuses infrared thermal imaging data with feature maps from the first branch that have similar semantic levels, including:
[0027] Infrared thermal imaging data were adjusted to match the desired values using bilinear interpolation. Same size, For the first branch Each feature map;
[0028] The adjusted infrared thermal imaging data was then added element by element. Integration.
[0029] Preferably, the hierarchical cross-modal feature fusion includes:
[0030] Flatten the feature maps extracted from the three branches into vectors respectively. Through linear projection Map to the same dimension, add position encoding, and input to the Transformer encoder;
[0031] By adopting a hierarchical fusion strategy, the Transformer encoder achieves cross-modal feature fusion at each layer of the feature pyramid through a self-attention mechanism.
[0032] Preferably, the dual-branch model includes a pose detection head and a keyhole detection head, both of which take the layered fusion features as input;
[0033] The pose detection head is used to predict the 6-DOF parameters of the container, including the three-dimensional translation vector. With rotation matrix ;
[0034] The keyhole detection head is used to predict the absolute three-dimensional coordinates of container keyholes. and open / closed state .
[0035] Preferably, the pose detection head maps the high-level semantic features output from the feature pyramid using a multilayer perceptron, and outputs... ,in, To apply the Rodriguez transformation, the rotation matrix is... The resulting rotation vector;
[0036] The keyhole detection head processes the high-level semantic features output from the feature pyramid using a multilayer perceptron, and the output is represented as follows: ,in The calculation is as follows:
[0037] ,
[0038] in, The predicted three-dimensional coordinates of the four corner points of the container. Corner point The three-dimensional coordinates From the corner point The estimated rotation matrix, This represents the offset vector of the keyhole relative to a corner point in a predefined standard template. This represents the adjustment offset of the model prediction relative to the standard template. This indicates the adjusted keyhole offset.
[0039] Preferably, the following loss function is set during the prediction process of the dual-branch model:
[0040] ,
[0041] ,
[0042] ,
[0043] in, For loss function, Loss for pose prediction To predict loss for keyholes, For geometric constraint loss, This is a dangling constraint term. and For hyperparameters, To determine the container pose based on the predicted position and and keyhole coordinates The coordinates of the keyhole in the world coordinate system are obtained based on the principle of rigid body transformation. This refers to the coordinates of the same keyhole in the world coordinate system predicted by the keyhole detection head. To smooth out the L1 loss, The number of container pairs detected. and The first The lock holes of the lower and upper layers of the container are at the corners. coordinate, This is the safe distance threshold.
[0044] Preferably, the method further includes:
[0045] The knowledge distillation transfer method is used to transfer the teacher model to a lightweight student model, where the teacher model refers to the two-branch model. During the transfer process,
[0046] A learnable cross-modal feature projection module is constructed to achieve dimensionality matching of the feature maps of the teacher model and the student model; the projection is represented as:
[0047] ,
[0048] in, This represents the multimodal fusion feature map of the teacher model. for Projected feature map Indicates upsampling, For learnable parameters, the following loss function is used during the learning process. :
[0049] ,
[0050] in, To select the distillation layer, It is the Frobenius norm. Feature map of student model;
[0051] Construct the task relationship matrix between the pose detection head and the keyhole detection head of the teacher model. and :
[0052] ,
[0053] ,
[0054] in Temperature coefficient;
[0055] Determine the dependencies between distillation tasks:
[0056] ;
[0057] in, For divergence;
[0058] The transfer teacher model must satisfy the following constraints:
[0059] ,
[0060] ,
[0061] in, and Depend on Analysis shows that, For the local coordinates of the keyhole, It is the keyhole index. Number of keyholes;
[0062] The overall training objective of the student model is to integrate task supervision and distillation loss.
[0063] ,
[0064] in, The total training loss for the student model. For the loss of original pose and coordinates, , , , These are adjustable weighting coefficients.
[0065] This invention also provides an automatic container identification and positioning system based on an end-to-end multi-task learning algorithm, used to implement the above-mentioned automatic container identification and positioning method based on an end-to-end multi-task learning algorithm, the system comprising:
[0066] The data processing module is used to acquire container multimodal data and perform spatiotemporal synchronization;
[0067] The multimodal feature extraction and fusion module is used to input the multimodal data into the multimodal feature pyramid backbone network, extract feature data of different modalities, and perform hierarchical cross-modal feature fusion to obtain hierarchical fused features;
[0068] The branch prediction module is used to construct a dual-branch model and predict the 6-DOF pose of the container and the 3D coordinates of the keyhole based on the hierarchical fusion features.
[0069] Preferably, the system further includes:
[0070] A transfer model is used to transfer a teacher model to a lightweight student model using a knowledge distillation transfer method. The teacher model refers to the two-branch model. During the transfer process...
[0071] A learnable cross-modal feature projection module is constructed to achieve dimensionality matching of the feature maps of the teacher model and the student model; the projection is represented as:
[0072] ,
[0073] in, This represents the multimodal fusion feature map of the teacher model. for Projected feature map Indicates upsampling, For learnable parameters, the following loss function is used during the learning process. :
[0074] ,
[0075] in, To select the distillation layer, It is the Frobenius norm. Feature map of student model;
[0076] Construct the task relationship matrix between the pose detection head and the keyhole detection head of the teacher model. and :
[0077] ,
[0078] ,
[0079] in Temperature coefficient;
[0080] Determine the dependencies between distillation tasks:
[0081] ;
[0082] in, For divergence;
[0083] The transfer teacher model must satisfy the following constraints:
[0084] ,
[0085] ,
[0086] in, and Depend on Analysis shows that, For the local coordinates of the keyhole, It is the keyhole index. Number of keyholes;
[0087] The overall training objective of the student model is to integrate task supervision and distillation loss.
[0088] ,
[0089] in, The total training loss for the student model. For the loss of original pose and coordinates, , , , These are adjustable weighting coefficients.
[0090] The beneficial effects of the technical solution of this invention are as follows:
[0091] This invention designs an end-to-end model architecture based on multimodal data input and a dual-branch decoder. It synchronously outputs the 6-DOF pose of a container and the 3D coordinates of a keyhole in a single inference iteration. An innovative keyhole coordinate decoding algorithm based on corner residuals is proposed, supporting keyhole open / closed state recognition. Through joint learning, it mines geometric constraints between tasks, addressing the shortcomings of error accumulation and low data transmission efficiency in multi-module algorithms, thus improving the overall accuracy and robustness of the algorithm. A multimodal adaptive fusion mechanism is designed, employing a hierarchical fusion strategy. It utilizes a self-attention mechanism to achieve cross-modal fusion at each layer of the feature pyramid, providing efficient and hierarchical feature representation for multimodal perception tasks. A novel model distillation framework is designed to adapt to multimodal input fusion, dual-task output coupling, and geometric constraint consistency, enabling efficient transfer of teacher model knowledge to a lightweight student model. The end-to-end model designed in this invention significantly reduces system complexity and design difficulty, improves system accuracy and reliability, ensures real-time performance, and possesses autonomous learning capabilities. Attached Figure Description
[0092] Figure 1 This invention provides a schematic flowchart of an automatic container identification and positioning method based on an end-to-end multi-task learning algorithm.
[0093] Figure 2 A schematic diagram of the multimodal feature pyramid backbone network structure provided by the present invention. Detailed Implementation
[0094] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0095] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0096] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0097] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0098] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0099] It should be emphasized here that the step markers mentioned below are not a limitation on the order of the steps, but should be understood as meaning that the steps can be executed in the order mentioned in the embodiments, or in a different order than in the embodiments, or several steps can be executed simultaneously.
[0100] This invention provides an automatic container identification and positioning method based on an end-to-end multi-task learning algorithm, see [link to relevant documentation]. Figure 1 It mainly includes the following steps:
[0101] S1. Acquire multimodal data and perform spatiotemporal synchronization; In this step, RGBD camera, LiDAR, IMU, and infrared thermal imaging sensor are used as multimodal data inputs. The multiple sensors complete spatiotemporal synchronization through time synchronization and spatial coordinate transformation.
[0102] S2. Input multimodal data into the feature pyramid backbone network, extract feature data of different modalities, and perform cross-modal feature fusion. In this step, the multimodal input data is preprocessed, the feature pyramid backbone is constructed, a multi-branch backbone network is designed, data of different modalities are processed separately, and then fused at different scales.
[0103] S3. Construct a dual-branch model to predict the 6-DOF pose of the container and the 3D coordinates of the keyhole, achieving end-to-end processing. The pose branch predicts the 6-DOF parameters of the container, including translation [tx, ty, tz] and rotation matrix R. The keyhole branch predicts the relative offset of the keyhole based on the container corner points, decoding the absolute coordinates using a predefined container geometry template. This step also includes introducing a geometric constraint loss to ensure mutual constraints between the two branches during training.
[0104] S4. Design a novel knowledge distillation transfer method that adapts to multimodal input fusion, dual-task output coupling, and geometric constraint consistency, to achieve efficient transfer from the above-mentioned dual-branch model (teacher model) to a lightweight student model.
[0105] In step S1 of the present invention, acquiring multimodal data and performing spatiotemporal synchronization specifically includes the following:
[0106] The system employs an RGBD camera, LiDAR, IMU, and infrared thermal imaging sensor as multimodal data inputs. The RGBD camera provides 3D geometric perception and captures texture information; the LiDAR point cloud fills visual blind spots and is unaffected by weather conditions, complementing the 3D camera; the IMU compensates for motion distortion of the LiDAR during vehicle movement and fuses with LiDAR data to suppress spreader vibration; the infrared thermal imaging sensor assists in detecting container corrosion by identifying surface temperature differences and detects keyhole opening / closing status through thermal conduction differences, and can replace visible light imaging in extreme environments. The sensor combination scheme designed in this invention can adapt to different environmental conditions such as strong light glare, rain, fog, vibration, and container corrosion through a multi-sensor collaborative mechanism. Furthermore, the sensor redundancy design improves the system's fault tolerance.
[0107] Multi-sensor time synchronization ensures consistent sampling times across all sensors through hardware triggering, and aligns asynchronous data using timestamp interpolation. Let the sensors... At any moment The collected data is ,in Indicates the sensor number, This represents the sampling time sequence number. For data points with time intervals, a linear interpolation algorithm is used. Assume we need to calculate the time at time... ( < < interpolated data The linear interpolation formula is:
[0108] ,
[0109] This formula allows for the estimation of missing data values within a time interval based on data points corresponding to adjacent timestamps. This enables precise alignment of asynchronous data in the time dimension, ensuring consistency of sensor data across time series and providing reliable time synchronization data for subsequent data fusion.
[0110] By establishing a mapping relationship between the sensor coordinate system and the global coordinate system through spatial coordinate transformation, sensor spatial synchronization is achieved. (LiDAR point cloud) Transform to world coordinate system W:
[0111] ,
[0112] in, It is data converted from lidar point cloud to the world coordinate system. It is the transformation matrix of the laser radar to the world coordinate system, including rotation and translation.
[0113] RGBD camera pixels ( Transform to camera coordinate system C:
[0114] ,
[0115] in, It is data converted from RGBD camera pixels to the camera coordinate system. For depth value, This is the camera intrinsic parameter matrix.
[0116] Transform from camera coordinate system to world coordinate system:
[0117] ,
[0118] This is the transformation matrix from the camera coordinate system to the world coordinate system.
[0119] Infrared image pixels Transforming to the world coordinate system requires a cascaded transformation from T to C to L to W:
[0120] ,
[0121] in, It is the extrinsic parameter matrix from the infrared sensor to the RGBD camera.
[0122] Through the coordinate transformation matrix operations described above, the mapping from the sensor coordinate system to the global coordinate system is realized, and sensor spatial synchronization is completed. This enables data collected by different sensors to be fused and analyzed under a unified spatial reference standard, thereby improving the accuracy of the system's spatial information judgment.
[0123] In step S2 of this invention, multimodal data is input into the feature pyramid backbone network, feature data of different modalities are extracted, and cross-modal feature fusion is performed, including the following:
[0124] Step S21: Preprocess the multimodal sensor data as follows:
[0125] Data augmentation processes such as HDR enhancement, adaptive histogram equalization, affine transformation, and noise addition are applied to RGB images to expand the dynamic range, improve detail, enhance contrast, and improve the algorithm's robustness to different environmental conditions. Bilateral filtering is applied to depth images to smooth them while preserving edge information.
[0126] Statistical filtering is used to remove outliers from the lidar point cloud, and voxel mesh filtering is performed. By dividing the point cloud into a three-dimensional voxel mesh, the centroid of each point in the voxel represents all points in that voxel, thereby reducing the point cloud density and the amount of data processing.
[0127] The temperature values of the infrared thermal imaging data are normalized to the [0,1] interval for easier subsequent processing. For temperature values... The normalization formula is:
[0128] ,
[0129] in, and These are the lowest and highest temperature values in the image, respectively.
[0130] Step S22: Construct a multimodal feature pyramid backbone network.
[0131] In multimodal perception tasks, due to significant differences in dimensionality, resolution, and semantic information among data from different sensors, constructing a multimodal feature fusion architecture is crucial for improving model performance. To address this issue, this invention designs a multi-branch backbone network that processes data from different modalities separately and performs feature fusion at different scales, achieving effective integration of multimodal data.
[0132] Multimodal feature pyramid backbone network structure such as Figure 2 As shown,
[0133] Branch 1 processes the RGB image (3 channels) and the depth map (1 channel), integrating them into a 4-channel input data X. RGB+D ∈R H×W×4Here, H and W represent the width and height of the image, respectively. An advanced CNN backbone network, ConvNeXt, is employed as the feature extractor. This network is based on a hierarchical pyramid architecture, progressively extracting multi-scale features by stacking multiple modules with different downsampling rates. Within each module, depthwise separable convolutions, layer normalization, and the GELU activation function are used to effectively reduce computational cost and enhance the network's non-linear expressive power.
[0134] Suppose that the ConvNeXt network contains The module, the first The output feature map of each module is ,in , and They represent the first Each module outputs the height, width, and number of channels of the feature map. The feature extraction process can be represented as:
[0135] ,
[0136] in, Through this hierarchical extraction, a multi-scale feature map set is finally obtained. These feature maps capture semantic and detail information at different levels.
[0137] Branch 2 focuses on processing LiDAR projection maps. LiDAR point cloud data is projected from the 3D space onto a plane aligned with the RGB image using a specific projection algorithm, forming a projection map with multiple feature channels. ,in This indicates the number of channels in the projection map, which may contain information such as the intensity and height of the point cloud.
[0138] To fully utilize the 3D spatial topological relationships of LiDAR data, sparse 3D convolution is employed for feature extraction. Sparse 3D convolution significantly reduces computation and memory usage by performing convolution operations only on non-zero elements, while effectively preserving the geometric structure information of the point cloud. Let the sparse 3D convolution kernel be... ,in , , These represent the dimensions of the convolution kernel in the x, y, and z directions, respectively. and These represent the number of input and output channels, respectively. For the input feature map... The sparse 3D convolution operation can be represented as:
[0139] ,
[0140] Where ⊙ represents element-wise multiplication. This is used to output feature maps. By embedding 3D spatial topological relationships into 2D feature maps in this way, the network's ability to understand scene geometry is significantly improved, resulting in feature maps. , and For feature map Height and width.
[0141] Branch 3 processes infrared thermal imaging (1 channel). Since the resolution of infrared images is typically lower than that of RGB images, directly extracting features from them independently may not fully utilize their information. Therefore, the infrared thermal imaging data is fused with a feature layer from the RGB branch that has a similar semantic level. Let the infrared thermal imaging data be... Select the first option in the RGB branch. Each feature layer The fusion process is performed using an element-by-element addition method, i.e.:
[0142] ,
[0143] in, () indicates an upsampling operation, which adjusts the infrared thermal imaging data to a level similar to that obtained using methods such as bilinear interpolation. Same size, This is the fused feature map. This fusion method can retain the high-resolution details of RGB images while introducing unique semantic information from infrared thermal imaging data, thus compensating for its insufficient resolution.
[0144] Step S23: Cross-modal feature fusion,
[0145] After completing feature extraction for RGB-D, LiDAR, and infrared thermal imaging, this step addresses the challenge of heterogeneous multimodal data dimensions by employing a Transformer encoder and a hierarchical fusion strategy to uniformly map different modal features to the feature pyramid space.
[0146] First, the outputs of the three feature extraction networks are converted into token sequences and input into the Transformer encoder.
[0147] Let the set of feature maps output by the RGB-D branch be... The lidar branch is Infrared branch is The superscript number corresponds to the feature pyramid level (e.g., 2 represents P2 level, which corresponds to 1 / 4 of the input size).
[0148] Flatten each feature map into a vector. Through linear projection Mapping to the same dimension ,Right now and add position encoding ( (where the length of the token sequence is) to obtain .
[0149] Employing a hierarchical fusion strategy, the Transformer achieves cross-modal fusion at each layer of the feature pyramid through a self-attention mechanism. Taking layer P4 as an example, it integrates the RGB-D branch... LiDAR branch Infrared branch (If the scales are different, adjust them to be consistent using methods such as bilinear interpolation) and then merge them.
[0150] Any of the following fusion methods can be used.
[0151] Fusion Method 1: Convolution after splicing
[0152] The feature maps of the three modalities are concatenated along the channel dimension to obtain... Through 1×1 convolution Compression fusion, the formula is:
[0153] ,
[0154] in, This represents the convolution operation. This is the activation function.
[0155] Fusion Method 2: Attention-Weighted Fusion
[0156] Taking the Squeeze-and-Excitation (SE) module as an example, the concatenated feature map... First, channel descriptors are obtained through global average pooling. Attention weights are then generated through two fully connected layers. :
[0157] ,
[0158] ,
[0159] in , ( (Compression ratio) It is the ReLU activation function. This is the Sigmoid function.
[0160] The final weighted fused feature map is as follows:
[0161] ,
[0162] Cross-modal attention mechanisms can dynamically allocate weights based on environmental conditions, and an attention weight matrix can be set up. The similarity of different modal features is calculated, and the weighted summation formula is as follows:
[0163] ,
[0164] in Represents mode, This is the corresponding modal feature map.
[0165] Step S24: Construct the feature pyramid.
[0166] The fused feature pyramid {P2, P3, P4, P5} will be used as input for subsequent detection heads (pose and keyhole coordinate prediction). Considering computational efficiency, fusion is preferentially performed at deeper semantic layers (such as P4, P5) because deeper features are more abstract and contain richer semantic information, resulting in better fusion performance.
[0167] Before fusion, features of each modality can be enhanced. Taking CBAM (Convolutional Block Attention Module) as an example, attention mechanisms are applied in both the spatial and channel dimensions. For feature map F, the spatial attention is calculated as follows:
[0168] ,
[0169] Channel attention is calculated as follows:
[0170] ,
[0171] The enhanced feature map is as follows:
[0172] ,
[0173] Another advantage of building a feature pyramid is that features at different resolutions can be specialized for specific tasks:
[0174] High-resolution P2 layer: with a resolution of 1 / 4 of the input, it focuses on fine texture recognition, such as the detection of minute features like box numbers and keyholes; Medium-resolution P3 layer: with a resolution of 1 / 8 of the input, it excels at geometric structure perception and can extract structural features such as corners and contours; Low-resolution P4 layer: with a resolution of 1 / 16 of the input, it focuses on spatial relationship reasoning and is suitable for tasks such as object stacking state analysis and pose estimation.
[0175] The multi-sensor feature pyramid architecture constructed through the above steps provides efficient and hierarchical feature representation for multimodal perception tasks.
[0176] In step S3 of this invention, a dual-branch decoder is constructed based on the multimodal feature pyramid to achieve container pose estimation and keyhole detection tasks, and the model accuracy and robustness are improved through geometric constraint optimization. The dual-branch decoder includes a pose detection head and a keyhole detection head, which share the multimodal features extracted by the backbone network. Through joint learning, the geometric constraint relationships between tasks are mined, effectively reducing error propagation.
[0177] The pose detection head is used to directly predict the 6-DOF parameters of a container, namely the three-dimensional translation vector. With rotation matrix To parameterize the rotation matrix, a Rodrigues transform is used to convert the rotation matrix into a rotation vector. This facilitates online learning. The output of the pose detection head is represented as... The high-level semantic features output from the feature pyramid are mapped using a multilayer perceptron (MLP). The computation process can be represented as follows:
[0178] ,
[0179] in, These are the P4 and P5 high-level feature maps in the feature pyramid. This is a multilayer perceptron network for a pose detection head.
[0180] The keyhole detection head is designed to predict the three-dimensional coordinates of the keyhole. and its open / closed state (0 indicates off, 1 indicates on). Let... The three-dimensional coordinates of the four corner points of the container predicted by the multi-task model, and the absolute three-dimensional coordinates of the keyhole. It can be calculated using the following formula:
[0181] ,
[0182] in, Corner point The weight, From the corner point The estimated rotation matrix, This represents the offset vector of the keyhole relative to a corner point in a predefined standard template. This represents the adjustment offset of the model prediction relative to the standard template. This indicates the adjusted keyhole offset.
[0183] The output of the keyhole detection head is expressed as follows: Similarly, features are processed using a multilayer perceptron:
[0184] ,
[0185] in, A multilayer perceptron network with keyhole branching.
[0186] To ensure that the pose and keyhole prediction results are physically consistent, a geometric constraint loss is introduced. An auxiliary consistency loss is constructed using the inherent geometric relationship between the predicted pose and the keyhole coordinates.
[0187] Known container pose and and keyhole coordinates According to the principle of rigid body transformation, the coordinates of the keyhole in the world coordinate system should satisfy the following relationship:
[0188] ,
[0189] The keyhole branch of the model directly predicts the coordinates of the same keyhole in the world coordinate system: .calculate and The difference between them is considered as loss:
[0190] ,
[0191] Among them, SmoothL1Loss is more robust to outliers than L2 Loss.
[0192] To address the issue of false detection of suspended containers, a suspension constraint term is introduced into the loss function by integrating a container stacking mechanics model. According to physical constraints, the lock holes of the lower box must be able to support the corner points of the upper box. Let the coordinates of the corner points of the upper box be... The coordinates of the lock hole in the lower box are: Then the dangling constraint term can be expressed as:
[0193] ,
[0194] in, The number of container pairs detected. and The first The lock holes of the lower and upper layers of the container are at the corners. coordinate, This is the safe distance threshold.
[0195] The model's total loss function Loss Prediction by Pose Keyhole Prediction Loss Geometric constraint loss and suspended constraint terms composition:
[0196] ,
[0197] in, and Hyperparameters are used to balance the weights of each loss term.
[0198] By designing this loss function, the model can guide the iterative update of model parameters during the training phase by quantifying the deviation between the predicted values and the actual target and physical constraints.
[0199] Through the above dual-branch decoder and geometric constraint optimization design, the model can effectively utilize multimodal features, automatically mine the geometric constraint relationship between tasks, and force the outputs of the two branches to be consistent in a physical sense, thereby reducing contradictory predictions and improving the overall accuracy and robustness of container end-to-end recognition.
[0200] In the task of jointly estimating container pose and keyhole coordinates, the aforementioned end-to-end bi-branch model (teacher model), while possessing high accuracy, suffers from computational complexity that makes it difficult to deploy to edge devices. Traditional distillation methods directly transfer feature responses from single-modal classification tasks, failing to adapt to the complexity of multimodal input fusion, dual-task output coupling, and geometric constraint consistency. This invention proposes a novel distillation framework to achieve efficient transfer of knowledge from the teacher model to a lightweight student model, the implementation process of which is as follows:
[0201] Step S41: Establish a Multi-modal Feature Adapter (MFA)
[0202] We construct a learnable cross-modal feature projection module to solve the problem of misalignment in multimodal feature space caused by architectural differences between teacher and student models.
[0203] Multimodal fusion feature map of teacher model Student Feature Map Dimension mismatch. Design a multimodal adapter for feature space projection:
[0204] ,
[0205] in These are learnable parameters. Define the feature-aligned distillation loss:
[0206] ,
[0207] in, To select the distillation layer, The Frobenius norm is used, and feature normalization eliminates scale differences.
[0208] Step S42: Dual-task Relation Distillation (DRD)
[0209] Construct a task correlation matrix between the pose branch and keyhole branch of the teacher model to maintain the synergy of the task decoupling structure of the student model.
[0210] Teacher model pose branch output With keyhole branch output There is a strong correlation. Construct a task relationship matrix:
[0211] ,
[0212] ,
[0213] in This is the temperature coefficient.
[0214] Dependencies between distillation tasks:
[0215] ,
[0216] KL divergence forces students to learn the statistical correlation of teacher task outputs.
[0217] Step S43: Geometric Constraint Transfer Loss (GCTL)
[0218] The implicitly learned geometric consistency knowledge of the teacher model is explicitly injected into student training to enhance the physical plausibility of the output. A differentiable geometric constraint function based on a 3D container model is defined:
[0219] ,
[0220] in, , Depend on Analysis shows that, For the local coordinates of the keyhole, This is the keyhole index. Transfer teacher constraint strength:
[0221] ,
[0222] By using geometric constraint transfer loss, students do not need to explicitly construct consistency loss; they can directly learn the geometric satisfaction of the teacher, thus avoiding the difficulty of student models autonomously learning complex constraints due to capacity limitations.
[0223] The overall training objective of the student model is to integrate task supervision and distillation loss.
[0224] ,
[0225] in, For the loss of original pose and coordinates, , , , These are adjustable weighting coefficients.
[0226] Based on the same inventive concept, this invention also provides an automatic container identification and positioning system based on an end-to-end multi-task learning algorithm, used to implement the above-mentioned automatic container identification and positioning method based on an end-to-end multi-task learning algorithm. The system includes:
[0227] The data processing module is used to acquire container multimodal data and perform spatiotemporal synchronization;
[0228] The multimodal feature extraction and fusion module is used to input the multimodal data into the multimodal feature pyramid backbone network, extract feature data of different modalities, and perform hierarchical cross-modal feature fusion to obtain hierarchical fused features;
[0229] The branch prediction module is used to construct a dual-branch model and predict the 6-DOF pose of the container and the 3D coordinates of the keyhole based on the hierarchical fusion features.
[0230] It is worth noting that the system embodiment corresponds to the above method embodiment. The implementation methods of the above method embodiments are all applicable to the system embodiment and can achieve the same or similar technical effects, so they will not be described in detail here.
[0231] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0232] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0233] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0234] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0235] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for automatic identification and positioning of containers based on an end-to-end multi-task learning algorithm, characterized in that, include: Acquire multimodal data of containers and perform spatiotemporal synchronization; The multimodal data is input into the multimodal feature pyramid backbone network to extract feature data of different modalities, and hierarchical cross-modal feature fusion is performed to obtain hierarchical fused features; A dual-branch model is constructed, and the 6-DOF pose of the container and the 3D coordinates of the keyhole are predicted based on the hierarchical fusion features.
2. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 1, characterized in that, The acquisition of container multimodal data and spatiotemporal synchronization includes: Multimodal data of the container were acquired using an RGBD camera, LiDAR, IMU, and infrared thermal imaging sensor, respectively. Hardware triggering ensures consistent sampling times for multimodal data, achieving time synchronization, and linear interpolation algorithms are used to align asynchronous data. Spatial synchronization is achieved by establishing a mapping relationship between the local coordinate system and the global coordinate system through spatial coordinate transformation.
3. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 1, characterized in that, The multimodal feature pyramid backbone network includes: The first branch is used to process RGB images and depth maps acquired by the RGBD camera and extract a set of multi-scale feature maps. The second branch is used to process the lidar projection map and extract feature maps. The third branch is used to process infrared thermal imaging data to obtain feature maps.
4. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 3, characterized in that, The first branch uses the CNN backbone network ConvNeXt as a feature extractor, the ConvNeXt network containing There are modules with different downsampling rates, and each module includes depthwise separable convolution, layer normalization, and GELU activation function; The first branch ultimately yields a multi-scale feature map set { } 5. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 3, characterized in that, The point cloud in 3D space collected by the lidar is projected onto a plane aligned with the RGB image to form a projection image with multiple feature channels. The second branch uses sparse 3D convolution to extract features from the projection map.
6. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 3, characterized in that, The third branch fuses infrared thermal imaging data with feature maps from the first branch that have similar semantic levels, including: Infrared thermal imaging data were adjusted to match the desired values using bilinear interpolation. Same size, For the first branch Each feature map; The adjusted infrared thermal imaging data was then added element by element. Integration.
7. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 3, characterized in that, The hierarchical cross-modal feature fusion includes: Flatten the feature maps extracted from the three branches into vectors respectively. Through linear projection Map to the same dimension, add position encoding, and input to the Transformer encoder; By adopting a hierarchical fusion strategy, the Transformer encoder achieves cross-modal feature fusion at each layer of the feature pyramid through a self-attention mechanism.
8. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 7, characterized in that, The dual-branch model includes a pose detection head and a keyhole detection head, both of which take the layered fusion features as input; The pose detection head is used to predict the 6-DOF parameters of the container, including the three-dimensional translation vector. With rotation matrix ; The keyhole detection head is used to predict the absolute three-dimensional coordinates of container keyholes. and open / closed state .
9. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 8, characterized in that, The pose detection head maps the high-level semantic features output from the feature pyramid using a multilayer perceptron, and outputs... ,in, To apply the Rodriguez transformation, the rotation matrix is... The resulting rotation vector; The keyhole detection head processes the high-level semantic features output from the feature pyramid using a multilayer perceptron, and the output is represented as follows: ,in The calculation is as follows: , in, The predicted three-dimensional coordinates of the four corner points of the container. Corner point The three-dimensional coordinates From the corner point The estimated rotation matrix, This represents the offset vector of the keyhole relative to a corner point in a predefined standard template. This represents the adjustment offset of the model prediction relative to the standard template. This indicates the adjusted keyhole offset.
10. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 9, characterized in that, In the prediction process of the dual-branch model, the following loss function is set: , , , in, For loss function, Loss for pose prediction To predict loss for keyholes, For geometric constraint loss, This is a dangling constraint term. and For hyperparameters, To determine the container pose based on the predicted position and and keyhole coordinates The coordinates of the keyhole in the world coordinate system are obtained based on the principle of rigid body transformation. This refers to the coordinates of the same keyhole in the world coordinate system predicted by the keyhole detection head. To smooth out the L1 loss, The number of container pairs detected. and The first The lock holes of the lower and upper layers of the container are at the corners. coordinate, This is the safe distance threshold.
11. The automatic container identification and positioning method based on an end-to-end multi-task learning algorithm according to claim 1, characterized in that, The method further includes: The knowledge distillation transfer method is used to transfer the teacher model to a lightweight student model, where the teacher model refers to the two-branch model. During the transfer process, A learnable cross-modal feature projection module is constructed to achieve dimensionality matching of the feature maps of the teacher model and the student model; the projection is represented as: , in, This represents the multimodal fusion feature map of the teacher model. for Projected feature map Indicates upsampling, For learnable parameters, the following loss function is used during the learning process. : , in, To select the distillation layer, It is the Frobenius norm. Feature map of student model; Construct the task relationship matrix between the pose detection head and the keyhole detection head of the teacher model. and : , , in Temperature coefficient; Determine the dependencies between distillation tasks: ; in, For divergence; The transfer teacher model must satisfy the following constraints: , , in, and Depend on Analysis shows that, For the local coordinates of the keyhole, It is the keyhole index. Number of keyholes; The overall training objective of the student model is to integrate task supervision and distillation loss. , in, The total training loss for the student model. For the loss of original pose and coordinates, , , , These are adjustable weighting coefficients.
12. An automatic container identification and positioning system based on an end-to-end multi-task learning algorithm, characterized in that, The system is used to implement the automatic container identification and positioning method based on the end-to-end multi-task learning algorithm according to any one of claims 1 to 11, the system comprising: The data processing module is used to acquire container multimodal data and perform spatiotemporal synchronization; The multimodal feature extraction and fusion module is used to input the multimodal data into the multimodal feature pyramid backbone network, extract feature data of different modalities, and perform hierarchical cross-modal feature fusion to obtain hierarchical fused features; The branch prediction module is used to construct a dual-branch model and predict the 6-DOF pose of the container and the 3D coordinates of the keyhole based on the hierarchical fusion features.
13. The container automatic identification and positioning system based on an end-to-end multi-task learning algorithm according to claim 12, characterized in that, The system also includes: A transfer model is used to transfer a teacher model to a lightweight student model using a knowledge distillation transfer method. The teacher model refers to the two-branch model. During the transfer process... A learnable cross-modal feature projection module is constructed to achieve dimensionality matching of the feature maps of the teacher model and the student model; the projection is represented as: , in, This represents the multimodal fusion feature map of the teacher model. for Projected feature map Indicates upsampling, For learnable parameters, the following loss function is used during the learning process. : , in, To select the distillation layer, It is the Frobenius norm. Feature map of student model; Construct the task relationship matrix between the pose detection head and the keyhole detection head of the teacher model. and : , , in Temperature coefficient; Determine the dependencies between distillation tasks: ; in, For divergence; The transfer teacher model must satisfy the following constraints: , , in, and Depend on Analysis shows that, For the local coordinates of the keyhole, It is the keyhole index. Number of keyholes; The overall training objective of the student model is to integrate task supervision and distillation loss. , in, The total training loss for the student model. For the loss of original pose and coordinates, , , , These are adjustable weighting coefficients.
Citation Information
Patent Citations
Workpiece 6D pose estimation method based on deep learning
CN111899301A
Container front hoisting control method and system, terminal and storage medium
CN118083809A
Bimodal fusion 6D pose estimation method based on Masked Point-Transform
CN118429421A