Empty container detection method and system based on multi-modal fusion
By using a multimodal fusion scheme of dual-sided 3D LiDAR and high-definition optical camera, a fused point cloud model with color information is generated, which solves the problems of high precision and easy deployment in empty container inspection, and realizes efficient and reliable detection of abandoned items and interlayers.
Patent Information
- Application Number
- CN202512009089.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-12-29
AI Technical Summary
Existing technologies cannot achieve high-precision 3D measurement and rich visual information for empty container inspection at low cost, and are not easily deployed at customs checkpoints, resulting in low inspection efficiency, insufficient accuracy, and safety hazards.
A multimodal fusion scheme using dual-sided 3D LiDAR and high-definition optical cameras is adopted. By synchronously acquiring, fusing and analyzing 3D point clouds and RGB images, a fused point cloud model with color information is generated. Combined with cluster analysis and size comparison, accurate detection of residues and interlayers is achieved.
It enables centimeter-level size measurement and detection of minute foreign objects, provides comprehensive empty container safety checks, provides intuitive and reliable test results, adapts to different lighting and weather conditions, is easy to deploy without modifying existing facilities, and improves on-site operation efficiency.
Smart Images

Figure CN121482501A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent customs and artificial intelligence technology, in particular to a container empty container detection method and system based on multi-modal fusion. BACKGROUND
[0002] In the customs clearance and station management process, quickly and accurately checking the declared empty container is a key link to prevent the concealment of prohibited goods and maintain the safety of trade business. When the container is declared empty, the external truck (container truck of external logistics company) leaves the port, and the customs needs to verify whether the container is empty at the gate. At present, the mainstream detection methods for whether the container is empty all have significant drawbacks:
[0003] Manual inspection: relying on customs officers to enter the container for sampling inspection, which has inherent defects such as low efficiency, high labor intensity, influence of subjective factors, and unguaranteed detection rate. It can cause congestion in peak period and may have safety hazards.
[0004] Technical detection:
[0005] 1) Ultrasonic detection: although there is no radiation, but the detection ability of residual material close to the wall is insufficient, and the precision is not enough.
[0006] 2) Radiographic imaging (X-ray): there is a risk of ionizing radiation, the cost of equipment purchase and maintenance is high, and the cost performance is low for empty container detection.
[0007] 3) Weight analysis: it is indirect detection, cannot locate residual material and identify sandwich, and is not sensitive to trace residual.
[0008] 4) Pure visual detection: affected by weather, no depth information (cannot identify sandwich), and poor identification effect for trace residual.
[0009] 5) Pure laser radar detection: usually 3D laser radar is deployed at the gate lifting rod. This method has specific requirements for the installation height of the radar to ensure that it can completely scan the internal structure of the container, so it often involves lifting the rod, reinforcing the structure, and modifying the line. In addition, the radar is installed on the vertical rod for long-term operation, which not only affects the overall detection accuracy of the system, but also affects the stress distribution of the lifting rod and the service life of the equipment due to continuous bearing and vibration. At the same time, this method lacks visual information support, cannot directly show the specific shape and appearance of the residual material in the container, and is difficult to store image evidence of illegal cases, which has certain limitations in responsibility tracing and data review.
[0010] In summary, the prior art fails to comprehensively utilize the precise three-dimensional geometric information and rich two-dimensional texture color information inside the container, and cannot realize a detection scheme with high-precision three-dimensional measurement, rich visual information, easy deployment and universality at low cost. SUMMARY
[0011] The present application aims to provide a container empty detection method and system based on multi-modal fusion, a storage medium and a computer program product to solve the problems raised in the background art.
[0012] The first aspect of the present application provides a container empty detection method based on multi-modal fusion, comprising the following steps:
[0013] Step S1, when the container door is opened, trigger two groups of sensor units to synchronously collect three-dimensional point cloud data and RGB images inside the container; wherein each group of sensor units comprises a three-dimensional laser radar and a high-definition optical camera;
[0014] Step S2, based on the pre-calibrated sensor parameters, fuse the three-dimensional point cloud data collected on both sides to the same coordinate system to form a complete three-dimensional point cloud inside the container;
[0015] Step S3, texture mapping the complete three-dimensional point cloud and the RGB image to generate a fusion point cloud model with color information;
[0016] Step S4, based on the fusion point cloud model, remove the container structure point cloud and perform clustering analysis on the remaining internal point cloud to detect whether there is residual material inside the container; and based on the fusion point cloud model, fit the inner surface of the container, calculate the actual internal dimensions and compare with the standard dimensions to determine whether there is a sandwich structure;
[0017] Step S5, output a detection report containing residual material information and sandwich detection results.
[0018] The second aspect of the present application provides a container empty detection system based on multi-modal fusion, comprising:
[0019] The data acquisition module is configured to: when the container door is opened, trigger two groups of sensor units to synchronously collect three-dimensional point cloud data and RGB images inside the container; wherein each group of sensor units comprises a three-dimensional laser radar and a high-definition optical camera;
[0020] The point cloud fusion module is configured to: based on the pre-calibrated sensor parameters, fuse the three-dimensional point cloud data collected on both sides to the same coordinate system to form a complete three-dimensional point cloud inside the container;
[0021] a texture mapping module, configured to texture map the complete three-dimensional point cloud with the RGB image to generate a fusion point cloud model with color information;
[0022] a detection analysis module, configured to detect whether there is residual object in the container based on the fusion point cloud model by removing the box structure point cloud and performing clustering analysis on the remaining internal point cloud, and fit the inner surface of the box based on the fusion point cloud model, calculate the actual internal size and compare it with the standard size to determine whether there is a sandwich structure;
[0023] a result output module, configured to output a detection report containing residual object information and sandwich detection results.
[0024] The third aspect of the present application provides a storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the method according to any one of the preceding aspects.
[0025] The fourth aspect of the present application provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps of the method according to any one of the preceding aspects.
[0026] Compared with the prior art, the present application has the following beneficial effects:
[0027] 1. High precision and high reliability: by fusing three-dimensional point cloud and two-dimensional image information, and combining point cloud fusion and fine calibration technology, the present application can realize centimeter-level size measurement and detection of small foreign objects, far exceeding other technical solutions.
[0028] 2. Comprehensive detection: simultaneously solving the two pain points of residual objects and sandwich, providing a comprehensive empty container safety verification solution.
[0029] 3. High efficiency and automation: the entire detection process is completed in tens of seconds without human intervention, and can realize all-weather and uninterrupted operation, greatly improving the efficiency of on-site operation.
[0030] 4. Intuitive and reliable detection results: the generated color point cloud model and two-dimensional image with accurate annotation make the detection results more intuitive, facilitating the customs officers to quickly verify and make decisions.
[0031] 5. Adaptability and robustness: combining the reliability of three-dimensional measurement and the intuitiveness of two-dimensional vision, the system can work stably under different light and weather conditions (radar is not affected by light), and the point cloud fusion and fine calibration process makes the system have self-learning and self-optimization ability, which can gradually adapt to changes in the field environment and sensor drift, ensuring the stability of long-term operation.
[0032] 6. Universality and generalizability: the proposed installation method (mounting on the stand on both sides of the customs checkpoint) is the typical layout of the customs checkpoint, without the need to modify the existing infrastructure, and the scheme is easy to replicate and deploy quickly at various customs ports. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A flowchart of a container empty container detection method based on multi-modal fusion disclosed in an embodiment of the present application is shown in the figure.
[0034] Figure 2 A schematic diagram of the layout of the sensor unit disclosed in an embodiment of the present application is shown in the figure.
[0035] Figure 3 A structural diagram of a container empty container detection system based on multi-modal fusion disclosed in an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0037] To solve the technical problems mentioned in the background art, the present application provides a container empty container detection scheme based on bilateral laser radar and camera multi-modal fusion. The core idea is to use a specific installation layout to obtain complementary point cloud data using bilateral laser radar, and to realize high-precision fusion and analysis of point cloud through a set of fine data processing procedures, thereby realizing accurate detection of container empty containers (leftover materials), interlayers, etc.
[0038] Referring to Figure 1 The present application provides a container empty container detection method based on multi-modal fusion, comprising the following steps:
[0039] Step S1, when the container door is opened, trigger two groups of sensor units to synchronously collect three-dimensional point cloud data and RGB images inside the container; wherein each group of sensor units comprises a three-dimensional laser radar and a high-definition optical camera;
[0040] In this step, when the vehicle arrives at the checkpoint passage (for example, whether the truck is parked in the preset detection area can be judged by video stream analysis), and it is determined that the container door is opened, the sensor units deployed on both sides of the checkpoint passage are automatically triggered to work synchronously. The left sensor unit and the right sensor unit independently and synchronously collect data on the interior space of the container.
[0041] The three-dimensional laser radar is responsible for acquiring point cloud data which is dense and has accurate three-dimensional coordinates inside the container, and the point cloud data completely reflects the spatial geometric structure of the inner wall and internal objects of the container. The high-definition optical camera cooperates with the laser radar to collect high-resolution RGB images covering the same field of view at the same time, and the RGB images provide rich visual appearance information such as texture and color of the scene inside the container.
[0042] In step S2, based on the pre-calibrated sensor parameters, the three-dimensional point cloud data collected on both sides is fused into the same coordinate system to form a complete three-dimensional point cloud inside the container.
[0043] In this step, the sensor internal and external parameters obtained by pre-calibration in the offline stage are called, mainly the external parameter matrix between the left and right sensor units. Based on the external parameter matrix, the three-dimensional point cloud data collected by the right sensor is uniformly converted into the coordinate system of the left sensor, for example, to realize the preliminary alignment and splicing of the point cloud data on both sides, and obtain a complete three-dimensional point cloud inside the container.
[0044] Further, in order to improve the fusion accuracy, the invention takes the parallelism and symmetry of the inner surface of the container (such as the left and right side walls) in the actual physical space as a constraint condition, constructs a nonlinear optimization problem, and iteratively optimizes and corrects the external parameter matrix of the left and right sensors used initially. After optimization, the optimized external parameters are used for point cloud transformation and splicing, and a plane segmentation algorithm such as RANSAC is used to accurately segment the left and right side walls, the top surface, the bottom surface and the innermost end surface of the container from the spliced point cloud, thereby forming a high-precision and complete three-dimensional point cloud model inside the container.
[0045] In step S3, the complete three-dimensional point cloud is texture-mapped with the RGB image to generate a fused point cloud model with color information.
[0046] In this step, for each three-dimensional point P(x, y, z) in the fused three-dimensional point cloud model generated in step S20, the corresponding camera-laser radar external parameter matrix and camera internal parameter are queried according to the source (left or right sensor unit) of the point. Using these parameters, the three-dimensional point P is accurately projected onto the RGB image plane collected by the corresponding high-definition optical camera, and its two-dimensional pixel coordinates (u, v) in the image are calculated. From the pixel coordinates (u, v), the RGB color value (including the brightness information of the red, green and blue channels) is extracted, and this color attribute is assigned to the three-dimensional point P.
[0047] By traversing all the three-dimensional points, a fused point cloud model with color attributes, high-precision three-dimensional geometric information and rich two-dimensional texture color is finally generated. It can be understood that the fused point cloud model has both the high-precision three-dimensional geometric information measured by the laser radar and the real visual appearance information obtained by the camera.
[0048] Step S4, based on the fused point cloud model, detect whether there is residual object inside the container by removing the box structure point cloud and clustering the remaining internal point cloud; and fit the inner surface of the box based on the fused point cloud model, calculate the actual internal dimensions and compare them with the standard dimensions to determine whether there is a sandwich structure;
[0049] In this step, based on the color fused point cloud model generated in step S3, the following two types of parallel detection analysis are performed:
[0050] Residual object detection: remove the box structure point cloud (each inner surface) segmented in step S2 from the fused point cloud model, and only keep the point cloud of the internal space of the box. Cluster analysis is performed on the internal point cloud to identify object clusters independent of the box wall. For each identified object cluster, its three-dimensional dimensions (length, width, height), spatial position coordinates in the box, and color information can be calculated and output, so as to comprehensively determine whether there is non-empty residual object.
[0051] Sandwich detection: based on the fused point cloud model, the plane equation of each inner surface is fitted again, and the intersection line and intersection point between these planes are calculated, so as to accurately calculate the actual length, width and height of the internal space of the container. Compare the actual calculated dimensions with the dynamic standard dimensions of the container type stored in the database through historical data self-learning and their statistical variance (such as 3σ range). If the deviation of a certain dimension (such as length) from the standard value exceeds a certain multiple (for example, 3 times variance) of its variance, or the parallelism of the relative surfaces of the box exceeds the set threshold, it is determined that the container has an illegally constructed sandwich structure.
[0052] Step S5, output the detection report containing residual object information and sandwich detection results.
[0053] In this step, the analysis results of step S4 are structured and packaged to generate a detection report. The detection report at least contains the overall state of empty / non-empty container, the determination result of whether there is a sandwich, and the detailed information list of the detected residual objects (such as number, three-dimensional dimensions and position of each residual object). At the same time, the detection results can also be visualized, for example, on the original RGB image, a prominent warning box is drawn around the detected residual objects. Finally, the structured detection report and the visualized image are reported to the customs business system for subsequent verification and decision making.
[0054] As an example, two groups of the sensor units are symmetrically deployed on both sides of the bay passage, their installation connection is perpendicular to the lane direction, and the installation spacing is greater than the maximum width of a standard container truck.
[0055] To realize the scanning of the interior of a standard container without dead angles and with high coverage, the application adopts a symmetrical and field engineering specification-compliant deployment scheme. Specifically, as shown in Figure 2 two groups of sensor units are rigidly mounted on fixed uprights on both sides of the customs checkpoint passage in a mirror-symmetrical manner. The connecting line direction of the two mounting points is designed to be substantially perpendicular to the driving direction of the container truck lane, thereby ensuring that the sensors observe from the front side of the container. In addition, the horizontal distance between the two mounting points is greater than the maximum overall width of the current standard container truck, thereby ensuring that even when the maximum size container truck is parked in the detection area, the sensors on both sides still have sufficient, unobstructed windows, enabling the laser beams and light rays emitted thereby to be shot from the side of the opened container door into the entire internal cavity space from the proximal container door edge to the distal container corner (i.e., the inner end surface close to the truck head), laying a physical foundation for subsequent acquisition of complete point cloud data.
[0056] As an example, in step S2, the extrinsic parameter matrix of the left and right point clouds is iteratively optimized based on the geometric feature constraints of the inner surface of the container to obtain the complete three-dimensional point cloud.
[0057] In the online detection phase, only the initial extrinsic parameter matrix obtained by offline calibration is used for point cloud splicing, which may introduce cumulative errors due to factors such as field vibration and temperature drift. To solve this problem, the application introduces a self-optimization process based on scene geometric characteristics after preliminary splicing. The core idea of the above self-optimization process is to use the container itself as a known calibration reference with regular geometric characteristics.
[0058] Specifically, first, the point cloud subsets of the inner surface of the container (especially the left and right side walls) are extracted from the preliminarily spliced point cloud. In an ideal case, these inner surfaces should be flat and parallel to each other. Based on this, the application constructs a nonlinear optimization problem with plane constraints and mirror constraints as the core. The plane constraint forces the same physical plane (such as the left side wall) fitted from the left and right point clouds to coincide after transformation; the mirror constraint utilizes the symmetry of the container structure.
[0059] The extrinsic parameter matrix (rotation matrix R and translation vector t) between the left and right sensors is adjusted in reverse through an iterative optimization algorithm (such as the Levenberg-Marquardt algorithm), so that the two side point clouds after transformation are optimally aligned under the condition of satisfying these strong geometric constraints.
[0060] Such processing can significantly correct subtle system errors and improve the point cloud splicing accuracy to the millimeter level, thereby ensuring the reliability of subsequent size measurement and detection of small foreign objects.
[0061] As an example, based on the geometric feature constraints of the inner surface of the container, the extrinsic parameter matrix of the left and right side point clouds is iteratively optimized to obtain the complete three-dimensional point cloud, comprising:
[0062] Geometric features capable of representing the rigid structure inside the container are extracted from the left and right point clouds, respectively, and based on the extracted geometric features, at least two different types of spatial correspondence constraints with complementary physical meanings are established between the left and right point clouds.
[0063] The at least two spatial correspondence constraints are fused to construct a multi-constraint coupled nonlinear optimization model, and the extrinsic parameter matrix is optimized and updated by iteratively solving the nonlinear optimization model to obtain a high-precision point cloud registration result that meets multiple geometric consistency requirements, i.e., the complete three-dimensional point cloud.
[0064] In this embodiment, the traditional method usually relies on fitting the inner surface of the container into an ideal plane as the basis for subsequent constraints, but in non-ideal cases (e.g., due to long-term use causing slight deformation of the container, presence of dirt or peeling of the coating on the inner wall, or having goods close to the wall when being detected), the inner surface point cloud cannot be accurately fitted into a standard plane model, or the fitted plane cannot represent the true wall position, which will cause the optimization based on the plane constraint to fail or introduce significant errors.
[0065] To solve this problem, the method of the present application no longer relies on the fitting of an idealized plane model. Specifically, geometric feature elements capable of representing the rigid main body structure inside the container are extracted from the point cloud data obtained by the left sensor unit and the point cloud data obtained by the right sensor unit. The geometric feature elements are different from the idealized plane model, but are local spatial information strongly related to the physical structure of the container body (such as reinforcing ribs and hinge seats) identified directly from the point cloud data. These structural features can maintain their basic geometric properties and relative positions even when the container body is slightly deformed or the surface is partially obscured.
[0066] Based on the extracted geometric feature elements, at least two types of spatial correspondence constraints with complementary characteristics in mathematical expression and physical meaning are constructed and established between the left and right point clouds. One type of constraint focuses on the alignment accuracy of the absolute coordinates of feature point pairs in three-dimensional space, providing accurate anchor points for optimization; the other type of constraint focuses on the relative consistency of structural line segments composed of feature points in length and direction, which has better tolerance to local deformation and noise, thereby providing effective supplementary constraints when plane features are unreliable.
[0067] Subsequently, at least two types of spatial correspondence constraints are fused to form a multi-constraint coupled nonlinear optimization mathematical model. By iteratively solving the mathematical model, the extrinsic matrix parameters between the left and right sensor units can be gradually adjusted and optimized. After this optimization process, a high-precision point cloud registration result that satisfies multiple geometric consistency criteria is obtained, which is the complete three-dimensional point cloud that accurately reflects the real geometric conditions of the container and is required for subsequent processes.
[0068] The present application can effectively overcome the technical limitations of traditional methods when the container geometry is not ideal through this strategy based on rigid structure features and multi-constraint fusion.
[0069] As an example, the establishment of at least two different types of spatial correspondence constraints includes establishing rigid constraints based on non-coplanar feature point pairs, including:
[0070] In the left and right point clouds, identify and extract the inherent, non-temporary physical structure feature regions of the container, including but not limited to: recessed regions formed by door hinge mounting seats, raised strip regions composed of inner wall reinforcing ribs, and local regions where floor fixing parts are located;
[0071] For each identified physical structure feature region, calculate its three-dimensional feature descriptor in the respective point cloud coordinate system and a local reference point representing the spatial position of the region;
[0072] By matching the three-dimensional feature descriptors in the two point clouds, two local reference points from the same physical structure and located in the left and right point clouds are established as a pair of virtual corresponding points;
[0073] A set of feature points distributed non-coplanarly in three-dimensional space is formed by multiple pairs of virtual corresponding points, and constraints are established based on this feature point set to minimize the sum of spatial distances between all corresponding point pairs after transformation by the extrinsic matrix.
[0074] In this embodiment, the spatial correspondence constraint includes rigid constraints based on non-coplanar feature point pairs, which solves the difficult problem of how to obtain stable and accurate absolute alignment benchmarks in complex scenarios.
[0075] Firstly, in the left and right point cloud data, through the point cloud local shape analysis and pattern recognition algorithm, the feature regions corresponding to the inherent and non-temporary physical structure of the container are automatically identified and extracted. The physical structure feature regions include but are not limited to the specific recessed regions formed by the door hinge mounting seat on the container body structure, the regular protruding strip regions presented by the reinforcing rib structure of the container inner wall, and the local regions where the locking devices for fixing goods on the container floor are located. These regions are reliable reference feature sources because of their stable structure, significant shape features and difficulty to be completely blocked by temporary goods.
[0076] For each successfully identified physical structure feature region, a three-dimensional feature descriptor in the local coordinate system of the side point cloud is calculated. The three-dimensional feature descriptor can effectively encode the local geometric characteristics such as the spatial distribution and normal vector variation of the region point cloud, and has certain rotation and translation invariance. At the same time, a local reference point for representing the spatial center position of the feature region is calculated and determined, for example, by calculating the centroid or feature peak point of the region point cloud.
[0077] Then, by performing similarity matching calculation (such as using the nearest neighbor search algorithm) on the three-dimensional feature descriptors of all extracted feature regions in the left and right point clouds, those feature regions with a similarity exceeding a preset threshold and considered to be derived from the same physical structure are paired. For each successfully paired region, the local reference points calculated in the left and right point clouds are established as a pair of virtual corresponding points.
[0078] Finally, a set of feature points distributed in a non-coplanar manner in three-dimensional space is formed by multiple pairs of virtual corresponding points. Based on this set of feature points, the mathematical expression of the rigid constraint is constructed: the optimization goal is to find an optimal external parameter transformation matrix, so that the sum of the squares of the three-dimensional Euclidean distances between the two corresponding points in all pairs of virtual corresponding points is minimized after the transformation. This constraint directly reflects the physical nature that the relative positions of feature points on the same object should remain unchanged under rigid transformation.
[0079] As an example, the establishment of at least two different types of spatial correspondence relationship constraints with complementary physical meanings also includes establishing a soft constraint based on internal structure topological consistency, including:
[0080] In the left and right point clouds, a group of local reference points are selected as nodes, and according to the actual physical connection relationship or spatial proximity relationship inside the container, the nodes are connected to form a spatial topological connection diagram;
[0081] In the spatial topology connection graph, pairs of line segments describing the same physical connecting component are identified, and a constraint is established for each identified pair of line segments, which requires that the lengths of the line segments in the pair should tend to be consistent and their three-dimensional spatial directions should tend to be parallel after the transformation by the external parameter matrix.
[0082] The constraints established based on all identified pairs of line segments are aggregated to collectively form the soft constraint of the internal structure topology consistency.
[0083] In this embodiment, the above-mentioned spatial correspondence constraint further includes a soft constraint based on the internal structure topology consistency, aiming to provide a relative geometric relationship correction mechanism with stronger robustness to local noise and non-rigid deformation.
[0084] Specifically, taking each of the local reference points extracted and calculated as a basis node, a connection is established between these nodes in the left and right point clouds respectively according to the prior knowledge of the internal structure of the container (such as the orientation of the reinforcing ribs, the relative position of the hinge seat) or the spatial proximity relationship between the nodes in the point cloud, thereby constructing a spatial topology connection graph for describing the internal skeleton framework of the container. The spatial topology connection graph is composed of a plurality of line segments connecting two nodes, and each line segment represents a stable spatial connection relationship between two feature areas.
[0085] Then, in the spatial topology connection graphs respectively constructed for the left and right point clouds, it is identified that those pairs of line segments describing the same physical connecting component by comparing whether the physical structure feature areas corresponding to the nodes at the two ends of the line segments match. For example, the line segment connecting two reinforcing rib protruding area reference points in the left point cloud and the line segment connecting the corresponding two reinforcing rib protruding area reference points in the right point cloud constitute a matched pair of line segments.
[0086] A specific soft constraint is established for each identified pair of line segments. The soft constraint contains two levels of requirements: (1) the lengths of the two line segments in the pair should tend to be consistent, i.e., the length difference should be minimized, after the transformation by the external parameter matrix; (2) the direction vectors of the two line segments in the pair in three-dimensional space should tend to be parallel, i.e., the direction angle should be minimized, after the transformation. The above two requirements collectively constitute the constraint of the consistency of the local size and orientation of the structure. It can be understood that the soft constraint has higher tolerance to the non-rigid deformation or point cloud noise of the container.
[0087] Finally, the constraints established based on all the identified matched line segments are summarized and integrated to form a soft constraint of internal structure topological consistency. Integrating this soft constraint with the aforementioned rigid constraints into the nonlinear optimization model can make the optimization process not only focus on the accurate alignment of feature points, but also force the overall shape correctness of the internal skeleton structure of the container, thereby significantly improving the accuracy and stability of point cloud registration in complex real scenarios.
[0088] As an example, in the step S4, the detection of the residual object is based on its three-dimensional geometric size, color and texture information; and the determination of the sandwich structure is achieved by comparing the actual internal size of the box body with the dynamically updated standard size of the box type and its allowable deviation range.
[0089] For residual object detection, a multi-modal information fusion criterion is used. Specifically, not only is the three-dimensional bounding box size of the clustered point cloud used to determine whether an object exists and its size (geometric criterion), but more importantly, the RGB color and texture information attached to the object point cloud are also synchronously fused (visual criterion). For example, a small dark block object, if its color is significantly different from the floor and the texture is abnormal, it will be given a higher confidence to determine as a suspicious residual object, rather than a stain or shadow on the floor. It can be understood that this comprehensive criterion can greatly reduce false positives due to changes in lighting and dust, etc.
[0090] For sandwich detection, the core is a dynamic and adaptive size comparison standard. Specifically, instead of using a fixed theoretical container size as a threshold, a continuously self-learning and updating database is maintained, which dynamically maintains a set of standard sizes and corresponding statistical variances for each common box type (such as 20GP, 40HQ). Among them, the standard size is calculated by moving average of historical detection data, and the variance represents the size fluctuation range of the box type under normal circumstances. During detection, the real-time calculated internal diameter is compared with the corresponding standard value in the database. Only when the deviation exceeds a certain multiple (such as 3σ) of its own variance, it is determined to be abnormal. It can be understood that this method can effectively distinguish between manufacturing tolerances, measurement noise and real illegal sandwich, and improve the robustness and adaptability of the system to different box types.
[0091] As an example, before the step S1, an offline calibration step is further included, which is used to obtain the internal parameters and external pose relationship of each sensor unit, and to establish the spatial transformation relationship between the left and right sensor groups.
[0092] The offline preparation stage is an essential and indispensable precise calibration process that enables the normal operation of the online detection method. This process is divided into two levels:
[0093] The first level is the participation in the external parameter calibration in a single sensor unit. For each laser radar-camera pair, the camera internal parameter (focal length, principal point, distortion coefficient) calibration and the rigid transformation relationship (external parameter) calibration of the laser radar to the camera are needed. This establishes the accurate projection mapping relationship from the three-dimensional point cloud to the two-dimensional pixel, which is the mathematical basis for the subsequent realization of point cloud coloring or texture mapping.
[0094] The second level is the joint calibration between sensor units. After the left and right two sensor units are installed in place, a specific calibration object (such as a calibration board with high-reflective marker points) needs to be placed in the public field of view area, and the spatial transformation relationship between the two sensor unit coordinate systems, i.e., the initial external parameter matrix, is calculated by simultaneously collecting data.
[0095] It can be understood that the accuracy of offline calibration directly determines the upper limit of the final detection performance.
[0096] Referring to FIG. 1, Figure 3 The embodiment of the present application also discloses a container empty container detection system 200 based on multi-modal fusion, which comprises:
[0097] A data acquisition module 201 is configured to trigger two groups of sensor units to synchronously collect three-dimensional point cloud data and RGB images inside the container when the container door is opened; wherein each group of sensor units comprises a three-dimensional laser radar and a high-definition optical camera;
[0098] A point cloud fusion module 202 is configured to fuse the three-dimensional point cloud data collected on both sides to the same coordinate system based on the pre-calibrated sensor parameters, to form a complete three-dimensional point cloud inside the container;
[0099] A texture mapping module 203 is configured to perform texture mapping on the complete three-dimensional point cloud and the RGB image, to generate a fusion point cloud model with color information;
[0100] A detection analysis module 204 is configured to detect whether there is a residual object in the container based on the fusion point cloud model by removing the container structure point cloud and performing clustering analysis on the remaining internal point cloud, and fit the inner surface of the container based on the fusion point cloud model, calculate the actual internal size and compare it with the standard size to determine whether there is a sandwich structure;
[0101] A result output module 205 is configured to output a detection report containing residual object information and sandwich detection results. The embodiment of the present application also discloses a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor realizes the steps of the method according to any one of the preceding embodiments when executing the computer program.
[0102] The embodiment of the present application further discloses a storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the method according to any one of the preceding embodiments.
[0103] The embodiment of the present application further discloses a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the steps of the method according to any one of the preceding embodiments.
[0104] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative effort shall fall within the protection scope of the present application.
Claims
1. A method for detecting empty containers based on multimodal fusion, characterized in that, Includes the following steps: Step S1: When the container door is opened, two sets of sensor units are triggered to synchronously collect three-dimensional point cloud data and RGB images inside the container; each set of sensor units includes a three-dimensional lidar and a high-definition optical camera. Step S2: Based on the pre-calibrated sensor parameters, the three-dimensional point cloud data collected from both sides are fused into the same coordinate system to form a complete three-dimensional point cloud inside the container. Step S3: Map the complete 3D point cloud to the RGB image to generate a fused point cloud model with color information; Step S4: Based on the fused point cloud model, the container structure point cloud is removed and the remaining internal point cloud is clustered to detect whether there are any remaining objects inside the container; and the internal surface of the container is fitted based on the fused point cloud model to calculate the actual internal dimensions and compare them with the standard dimensions to determine whether there is a sandwich structure. Step S5: Output an inspection report containing information on the remaining objects and the results of the interlayer inspection.
2. The method for detecting empty containers based on multimodal fusion according to claim 1, characterized in that: The two sets of sensor units are symmetrically deployed on both sides of the checkpoint channel, with their installation lines perpendicular to the lane direction and the installation spacing greater than the maximum width of a standard container truck.
3. The method for detecting empty containers based on multimodal fusion according to claim 1, characterized in that: In step S2, based on the geometric feature constraints of the inner surface of the container, the extrinsic parameter matrices of the point clouds on the left and right sides are iteratively optimized to obtain the complete three-dimensional point cloud.
4. The method for detecting empty containers based on multimodal fusion according to claim 3, characterized in that: Based on the geometric feature constraints of the inner surface of the container, the extrinsic parameter matrices of the point clouds on the left and right sides are iteratively optimized to obtain the complete 3D point cloud, including: Geometric features that characterize the rigid structure inside the container are extracted from the left and right point clouds respectively. Based on the extracted geometric features, at least two different types of spatial correspondence constraints with complementary physical meanings are established between the left and right point clouds. By fusing at least two of the aforementioned spatial correspondence constraints, a multi-constraint coupled nonlinear optimization model is constructed. The extrinsic parameter matrix is then optimized and updated by iteratively solving the nonlinear optimization model to obtain a high-precision point cloud registration result that satisfies multiple geometric consistency requirements, thereby obtaining the complete 3D point cloud.
5. The method for detecting empty containers based on multimodal fusion according to claim 4, characterized in that: The establishment of at least two different types of spatial correspondence constraints with complementary physical meanings includes establishing rigid constraints based on non-coplanar feature point pairs, including: In the left and right point clouds, identify and extract the inherent, non-temporary physical structural feature regions of the container, including but not limited to: the recessed area formed by the door hinge mounting base, the raised strip area formed by the inner wall reinforcing ribs, and the local area where the floor fasteners are located. For each identified physical structure feature region, calculate its three-dimensional feature descriptor in its respective point cloud coordinate system and a local reference point representing the spatial location of the region; By matching the three-dimensional feature descriptors in the point clouds on both sides, two local reference points that originate from the same physical structure and are located in the left and right point clouds respectively are established as a pair of virtual corresponding points. A set of feature points that are not coplanarly distributed in three-dimensional space is formed by multiple pairs of virtual corresponding points, and constraints are established based on this set of feature points to minimize the sum of spatial distances between all corresponding point pairs after the transformation of the extrinsic parameter matrix.
6. The method for detecting empty containers based on multimodal fusion according to claim 5, characterized in that: The establishment of at least two different types of spatial correspondence constraints with complementary physical meanings also includes establishing soft constraints based on the topological consistency of the internal structure, including: In the left and right point clouds, a set of local reference points are selected as nodes, and a spatial topology connection diagram is formed between the nodes according to the actual physical connection relationship or spatial proximity relationship inside the container. In the spatial topology connection diagram, line segment pairs describing the same physical connection component are identified, and constraints are established for each identified line segment pair. The constraints require that after the transformation of the extrinsic parameter matrix, the length of the line segment pair should tend to be consistent, and their three-dimensional spatial directions should tend to be parallel. The constraints established based on all identified line segments are summarized to form the soft constraints for the topological consistency of the internal structure.
7. The method for detecting empty containers based on multimodal fusion according to claim 1, characterized in that: In step S4, the detection of the remnant is based on its three-dimensional geometric dimensions, color, and texture information; the determination of the sandwich structure is achieved by comparing the actual internal dimensions of the box with the dynamically updated standard dimensions of the box type and its allowable deviation range.
8. The method for detecting empty containers based on multimodal fusion according to claim 1, characterized in that: Before step S1, an offline calibration step is also included, which is used to obtain the internal parameters and external pose relationship of each group of sensor units and to establish the spatial transformation relationship between the left and right sensor groups.
9. A container empty container inspection system based on multimodal fusion, characterized in that, include: The data acquisition module is configured to: trigger two sets of sensor units to synchronously acquire three-dimensional point cloud data and RGB images inside the container when the container door is opened; each set of sensor units includes a three-dimensional lidar and a high-definition optical camera; The point cloud fusion module is configured to: fuse the three-dimensional point cloud data collected from both sides into the same coordinate system based on pre-calibrated sensor parameters to form a complete three-dimensional point cloud inside the container; The texture mapping module is configured to: perform texture mapping between a complete 3D point cloud and an RGB image to generate a fused point cloud model with color information; The detection and analysis module is configured to: detect whether there are any remaining objects inside the container by removing the container structure point cloud and performing cluster analysis on the remaining internal point cloud based on the fused point cloud model; and calculate the actual internal dimensions by fitting the inner surface of the container based on the fused point cloud model and comparing them with standard dimensions to determine whether there is a sandwich structure. The results output module is configured to output a test report containing information on the remains and the results of the interlayer detection.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Corn plant height detection method based on binocular image and ground-based radar fusion point cloud
CN116883480A
Empty box detection method and device, electronic equipment and storage medium
CN117075139A
Empty container interlayer judgment method and system based on 3D laser point cloud
CN118505613A
Empty box identification method, device and equipment and storage medium
CN120388358A
Building design and construction method and system based on machine vision
CN120747410A