Airline luggage stack type identification method, device and equipment, medium and product
Through the depth camera and improved YOLOv8-seg model combined with point cloud technology, the environmental perception and adaptability problems in aviation luggage stack recognition are solved, high-precision luggage type and size recognition is achieved, recognition efficiency and accuracy are improved, and key data support is provided for the automated sorting system.
Patent Information
- Application Number
- CN202510627227.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-26
AI Technical Summary
The existing aeronautical luggage stack recognition technology is insufficient in terms of environmental perception and adaptability, making it difficult to accurately obtain three-dimensional space information of luggage. Especially when dealing with abnormal placement positions and irregular shapes, the recognition efficiency and accuracy are not ideal, which affects the efficiency and safety of luggage transportation.
The RGB and depth stacking images are synchronously collected by depth cameras, and baggage recognition is performed through the improved YOLOv8-seg model, luggage point cloud data is generated, and luggage rotation angle and dimension information is obtained by combining principal component analysis and axial enclosure box calculation to achieve high-precision stacking recognition.
In dense stacking scenarios, efficient and accurate identification of luggage type, rotation angle and size information is achieved, which improves identification accuracy and robustness, provides a reliable data foundation for automated sorting systems, and reduces computing complexity.
Smart Images

Figure CN120544145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment, medium and product for identifying the stack type of airline luggage. Background Art
[0002] Currently, compared to other stages of the departure baggage process, the handling and loading of airline baggage still relies primarily on manual labor. Loading is a complex, repetitive process with high workload. Long hours of manual work inevitably lead to decreased efficiency due to fatigue and damage to baggage caused by careless handling. To improve loading efficiency and ensure baggage safety, robots are being introduced to replace manual labor in these complex, repetitive tasks. Robots can perform the collection, handling, and stacking of baggage according to pre-set programs and instructions. However, robots cannot autonomously acquire external information and require the assistance of devices such as vision sensors to obtain more information about the environment and the operation. Therefore, the use of vision-guided robots to automatically handle baggage has become necessary.
[0003] Existing technologies for identifying the stack shape of airline baggage still have numerous shortcomings. Existing recognition systems have significant deficiencies in environmental perception, making it difficult to accurately capture three-dimensional spatial information about baggage, directly impacting the accuracy of subsequent sorting and stacking operations. Furthermore, existing technologies have limited ability to identify the physical characteristics of baggage, particularly in special cases such as unusual placement and irregular shapes, resulting in suboptimal recognition efficiency and accuracy. Furthermore, faced with the complex operating environment characterized by the wide variety of baggage sizes and types, existing systems lack the intelligence and adaptability to adapt. These technical shortcomings not only severely impact the efficiency and safety of baggage transportation but also become a key factor hindering the advancement of automation in aviation logistics. Summary of the Invention
[0004] The present invention provides a method, device, equipment, medium, and product for identifying the stack type of airline baggage. These methods can efficiently and accurately identify the type, rotation angle, and size of each piece of luggage in a densely stacked baggage stack, providing reliable technical support for intelligent sorting systems.
[0005] According to one aspect of an embodiment of the present invention, a method for identifying the stack type of airline baggage is provided, the method comprising:
[0006] Use a depth camera to capture RGB stacking images and depth stacking images of the airline baggage stacking area;
[0007] The RGB stack image is input into the baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag in the airline baggage stack area. The baggage recognition model is obtained by training the improved YOLOv8-seg model.
[0008] Generate baggage point cloud data corresponding to each baggage according to the mapping position of each baggage sub-image in the depth stacking image;
[0009] Based on the baggage point cloud data of each bag, calculate the baggage rotation angle of each bag in the airline baggage stacking area and the baggage size information of each bag;
[0010] The baggage type, baggage rotation angle, and baggage size information of each baggage are determined as the stack type recognition result of the aviation baggage stacking area.
[0011] According to another aspect of an embodiment of the present invention, there is also provided a device for identifying the stack type of airline luggage, the device comprising:
[0012] A data acquisition module, used to collect RGB stacking images and depth stacking images of the aviation baggage stacking area through a depth camera;
[0013] The baggage recognition module is used to input the RGB stack image into the baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag. The baggage recognition model is obtained by training the improved YOLOv8-seg model.
[0014] A point cloud generation module is used to generate baggage point cloud data corresponding to each bag according to the mapping position of each baggage sub-image in the depth stacking image;
[0015] The feature analysis module is used to calculate the luggage rotation angle of each bag in the airline baggage stacking area and the luggage size information of each bag based on the baggage point cloud data of each bag;
[0016] The information output module is used to determine the baggage type, baggage rotation angle and baggage size information of each baggage as the stack type recognition result of the aviation baggage stacking area.
[0017] According to another aspect of an embodiment of the present invention, an electronic device is provided, the electronic device comprising:
[0018] at least one processor; and
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the method for identifying the stack type of airline baggage according to any embodiment of the present invention.
[0021] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for identifying the stack type of airline luggage according to any embodiment of the present invention when executed.
[0022] According to another aspect of an embodiment of the present invention, a computer program product is provided, comprising computer instructions, which implement the steps of the method according to any embodiment of the present invention when executed by a processor.
[0023] The technical solution of the embodiment of the present invention uses a depth camera to simultaneously capture visual and depth information of the stacking area and trains color images based on a deep learning algorithm to accurately segment each individual baggage unit and identify its category, significantly improving the ability to distinguish targets in densely stacked scenarios. The system then intelligently matches the two-dimensional recognition results with the depth information to construct a three-dimensional point cloud model with dimensional information, effectively overcoming the limitations of traditional methods in spatial positioning. Based on the point cloud data of each piece of luggage, the system can accurately calculate the spatial orientation and physical dimensions of each piece of luggage, maintaining stable measurement accuracy even in complex stacking situations. The final output structured recognition result fully presents key parameters such as the luggage's type characteristics, placement posture, and external dimensions, providing a reliable data foundation for automated operation processes. This solution cleverly integrates deep learning models and three-dimensional vision technology to achieve a complete closed loop from two-dimensional perception to three-dimensional measurement. While improving recognition accuracy, it also significantly reduces computational complexity, fully meeting the stringent requirements of modern aviation logistics for efficient and precise operations.
[0024] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1This is a flow chart of a method for identifying the stack type of airline luggage provided in accordance with the first embodiment of the present invention;
[0027] Figure 2 This is a structural diagram of an efficient multi-scale attention mechanism network in a method for identifying the stack type of airline luggage applicable to an embodiment of the present invention;
[0028] Figure 3 This is a flow chart of a method for identifying the stack type of airline luggage provided in accordance with a second embodiment of the present invention;
[0029] Figure 4 This is a flowchart of specific steps for identifying the stack type of airline baggage in a method for identifying the stack type of airline baggage provided in accordance with the third embodiment of the present invention;
[0030] Figure 5 This is a flowchart of specific steps for measuring baggage in a method for identifying the stack type of airline baggage provided in accordance with the third embodiment of the present invention;
[0031] Figure 6 2 is a schematic structural diagram of a device for identifying the stack type of airline luggage provided in accordance with a fourth embodiment of the present invention;
[0032] Figure 7 It is a structural diagram of an electronic device for implementing a method for identifying the stack type of airline luggage provided in the fifth embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] Example 1
[0036] Figure 1 This is a flow chart of a method for identifying the stack type of airline baggage provided in a first embodiment of the present invention. This embodiment is applicable to situations where the stack type of airline baggage is identified in a stacking area. The method can be performed by an airline baggage stack type identification device. The airline baggage stack type identification device can be implemented in the form of hardware and / or software and can generally be configured in an electronic device, such as a desktop computer or a server.
[0037] Correspondingly, such as Figure 1 As shown, the method includes:
[0038] S110 : Capturing an RGB stacking image and a depth stacking image of the airline baggage stacking area through a depth camera.
[0039] In an embodiment of the present invention, a high-precision depth camera based on structured light or time-of-flight (ToF) principles can be used to synchronously acquire dual-modal image data of the baggage stacking area. Structured light technology calculates depth information by projecting a specifically coded light spot pattern and analyzing its deformation, while ToF technology acquires distance data by measuring the time-of-flight of infrared light. This system ensures pixel-level alignment of RGB and depth images through multi-sensor spatiotemporal synchronization, incorporates adaptive exposure adjustment to adapt to the complex lighting environment of airports, and employs a real-time data verification mechanism to ensure acquisition quality. Ultimately, this system achieves high-quality, simultaneous dual-modal data acquisition, providing a reliable input source with both rich color information and precise geometric data for subsequent processing.
[0040] The RGB stacking image refers to a standard color image of the luggage stacking scene, typically at a high resolution, such as 1920×1080. This RGB stacking image preserves visual features such as the luggage's surface texture and color. The depth stacking image is a single-channel matrix aligned with the RGB image pixels, with each pixel value representing the actual distance from the corresponding point to the camera. For example, when three suitcases are stacked on a conveyor belt, the system simultaneously collects a set of data, including a color image showing the luggage's appearance (i.e., the RGB stacking image) and a depth image recording the spatial position of each piece of luggage (i.e., the depth stacking image).
[0041] Understandably, this dual-modal acquisition approach offers multiple technical advantages: First, the simultaneously acquired depth camera data provides complete two-dimensional and three-dimensional information for subsequent processing; second, the active ranging feature of the depth camera makes it unaffected by ambient lighting, enabling stable operation in a variety of lighting conditions; and finally, hardware-level data alignment greatly simplifies the subsequent image registration process. For example, in airport sorting scenarios, the system can reliably acquire high-quality depth information even in dimly lit baggage sorting areas. This specialized image acquisition solution not only addresses the depth perception challenges of traditional monocular vision but, more importantly, lays a high-quality data foundation for the entire system. Careful camera selection, installation layout, and acquisition parameter optimization ensure accurate and reliable input data for subsequent processing, a primary guarantee for achieving high-precision pallet recognition.
[0042] S120. Input the RGB stacking image into a baggage recognition model to obtain a baggage subimage of each bag in the aviation baggage stacking area in the RGB stacking image, as well as the baggage type of each bag. The baggage recognition model is obtained by training an improved YOLOv8-seg model.
[0043] In this embodiment of the present invention, an RGB image of the baggage stacking area is first captured and then fed into a specially optimized baggage recognition model. This model, based on a modified YOLOv8-seg architecture and trained on massive amounts of airline baggage data, possesses powerful image analysis capabilities. During processing, the model simultaneously performs two key functions: first, accurately segmenting each individual bag in the image and generating a corresponding sub-image; second, accurately identifying and classifying various types of baggage (e.g., suitcases, backpacks, etc.).
[0044] A sub-image refers to an image region that is segmented from the original stack image and contains only a single complete piece of luggage. Specifically, the model first generates a binary mask of the same size as the original image. This is a matrix composed of 0s and 1s, where 1 (white) represents the pixel area belonging to the target luggage, and 0 (black) represents the background or other luggage areas. Take a real-world scenario as an example: when the input image contains a red suitcase and a blue backpack stacked together, the model generates two independent binary masks, one corresponding to each of the two pieces of luggage. Then, by performing a pixel-by-pixel multiplication operation on the mask and the original image, two sub-images are finally obtained: one showing only the red suitcase (with a black background), and the other showing only the blue backpack (with a black background).
[0045] This processing method offers three advantages: First, the binary mask precisely defines the physical boundaries of the baggage, avoiding interference from adjacent bags. Second, the generated sub-images preserve the baggage's complete visual features (such as color and texture), providing high-quality input for subsequent classification. Finally, each sub-image contains the baggage's spatial position within the original image, which is crucial for subsequent 3D coordinate calculations. For example, in airport applications, the system can accurately calculate the specific location of the baggage on the conveyor belt by analyzing the pixel distribution in the luggage image, guiding the robotic arm for precise grasping.
[0046] Notably, this refined image processing technology not only improves the accuracy of individual baggage recognition but, more importantly, solves the challenge of object separation in densely stacked baggage scenarios. By combining deep learning and computer vision technologies, the system intelligently handles a variety of complex situations, including partial occlusion, changing lighting, and other common real-world issues. This dual recognition mechanism ensures that the system can reliably identify and distinguish each piece of luggage, even in complex, densely stacked baggage scenarios, providing an accurate data foundation for subsequent processing.
[0047] S130 : Generate luggage point cloud data corresponding to each piece of luggage according to the mapping position of each luggage sub-image in the depth stacking image.
[0048] In this embodiment of the present invention, the system intelligently matches 2D baggage sub-images with 3D depth information through precise coordinate mapping technology. This process, based on the calibration parameters of the depth camera and the principles of geometric projection, enables accurate conversion from 2D pixels to 3D space. During this process, the system performs two key operations: first, extracting the corresponding depth data from the depth image based on the sub-image position; second, converting this depth data into a true-to-size 3D point cloud.
[0049] Among them, luggage point cloud data refers to a set of points that represent the geometric shape of the luggage surface, generated by three-dimensional reconstruction technology. Specifically, the system will first establish a conversion relationship between the pixel coordinate system and the three-dimensional camera coordinate system based on the camera's intrinsic parameters (including focal length, principal point coordinates, etc.). Take actual operation as an example: when the sub-image coordinates of a suitcase in the RGB image are (x1, y1) to (x2, y2), the system will extract the corresponding depth value matrix at the same position in the depth image, and then convert each pixel point (u, v) and its depth value d into three-dimensional coordinates (X, Y, Z) through the perspective projection formula, and finally generate a suitcase point cloud model composed of thousands of three-dimensional points. This three-dimensional reconstruction method ensures the integrity and accuracy of the point cloud data through sub-image-based depth data extraction, while retaining the original texture information. It not only effectively avoids background noise interference, but also provides accurate posture reference for robotic arm grasping.
[0050] Notably, this subgraph-guided point cloud generation technology effectively solves the segmentation challenges faced by traditional methods in densely packed scenes. By combining deep learning segmentation with 3D vision technology, the system can intelligently handle a variety of complex situations, including touching luggage and partial occlusion. This innovative data processing pipeline provides a reliable geometric data foundation for subsequent luggage dimensional measurement and spatial positioning, significantly improving the measurement accuracy and stability of the entire system.
[0051] S140: Calculate the luggage rotation angle of each luggage in the airline luggage stacking area and luggage size information of each luggage based on the luggage point cloud data of each luggage.
[0052] In the embodiment of the present invention, the system uses advanced three-dimensional point cloud analysis technology to accurately process the point cloud data of each independent baggage.
[0053] Optionally, an improved Principal Component Analysis (PCA) algorithm and Axis-Aligned Bounding Box (AABB) calculation method can be used. Through training and optimization with massive point cloud samples, the system possesses powerful 3D feature analysis capabilities. During processing, the system simultaneously performs two key functions: accurately calculating the spatial rotation angle of luggage and precisely measuring the luggage's 3D physical dimensions.
[0054] The luggage rotation angle can be understood as the luggage's rotation angle relative to the designated coordinate axes in the depth camera's camera coordinate system, indicating the luggage's rotational shape within the designated camera plane. Alternatively, since the camera coordinate system and the spatial coordinate system have a predefined coordinate transformation relationship, the luggage rotation angle can also be understood as the luggage's rotation angle relative to the designated coordinate axes in the spatial coordinate system, indicating the luggage's rotational shape within the designated spatial plane. The luggage dimension information can be understood as the length, width, and height of the luggage. It is understood that if a piece of luggage is not a standard rectangular shape, the luggage dimension information can be the length, width, and height of the minimum circumscribed rectangle of the luggage.
[0055] Specifically, the system generally first denoises and downsamples the point cloud data, then calculates the covariance matrix and performs eigenvalue decomposition to obtain three eigenvectors that represent the main directions of the point cloud distribution. Take a real-world scenario as an example: when the point cloud data of a standard suitcase is input into the system, the algorithm determines its main direction vector (that is, the direction corresponding to the largest eigenvalue), calculates the precise rotation angle (such as a 15° angle with the horizontal direction) based on the angle between this vector and the original coordinate system (typically, the camera coordinate system), and simultaneously calculates the length, width, and height dimensions (such as 68cm×42cm×25cm) based on the difference in the coordinates of the extreme points along the three axes.
[0056] It is understandable that this three-dimensional feature analysis method ensures the accuracy of posture calculation through principal component analysis, and combines the axial bounding box to achieve accurate size measurement, which not only effectively overcomes the influence of point cloud noise, but also provides a complete posture and size reference for robotic arm grasping. For example, in the airport sorting system, through this high-precision three-dimensional analysis, the robotic arm can accurately identify tilted suitcases and adjust the grasping angle to complete precise operations. This innovative three-dimensional data processing technology effectively solves the measurement difficulties of traditional methods in densely stacked scenarios. By combining point cloud processing and geometric calculation technology, the system can intelligently handle various complex situations, including common problems in actual scenarios such as partial point cloud missing and irregular shapes, providing reliable three-dimensional data support for automated sorting systems.
[0057] S150: Determine the baggage type, baggage rotation angle, and baggage size information of each baggage as a stack type recognition result of the aviation baggage stacking area.
[0058] In this embodiment of the present invention, the system intelligently integrates and outputs key feature data obtained through early processing using multi-dimensional information fusion technology. This process, based on a structured data fusion algorithm, systematically transforms raw images into complete stacking information through a pre-set standardized output interface. During this process, the system primarily performs two core functions: first, standardizing and packaging various recognition results; second, generating stacking description files that can be directly used by downstream systems.
[0059] The stack type recognition result refers to a structured data set containing complete feature information for each piece of luggage. Specifically, the system associates and integrates the luggage type label (such as "trolley case") output by the recognition model, the rotation angle (such as "15° from the horizontal") calculated by the 3D analysis module, and the precise size data (such as "68cm×42cm×25cm"). For example, when processing a stacking area containing five pieces of luggage, the system generates JSON-formatted data containing five records. Each record fully describes the type, 3D position, and physical dimensions of a piece of luggage, while preserving the relative positional relationships between the pieces of luggage.
[0060] It is understandable that this method of describing the stack type ensures information integrity through standardized data structures and uses lightweight data formats to improve transmission efficiency. It not only achieves efficient use of recognition results, but also provides a comprehensive basis for decision-making for the automated sorting system. For example, in the airport intelligent sorting center, the stack type recognition results can be directly imported into the robotic arm control system to guide it to automatically plan the optimal grabbing sequence and stacking plan based on the size and angle of the luggage. This end-to-end information integration technology effectively solves the problem of information islands in traditional sorting systems. Through a unified data interface and intelligent information fusion, the system can adapt to various complex operational requirements, including advanced application scenarios such as multi-robotic arm collaborative operation and dynamic path planning, providing a complete data solution for modern aviation logistics automation.
[0061] The technical solution of the embodiment of the present invention uses a depth camera to synchronously collect RGB stacking images and depth stacking images of the aviation baggage stacking area. This multimodal data acquisition method effectively overcomes the limitations of single sensor information and provides a more comprehensive data basis for subsequent processing; the RGB stacking image is input into the baggage recognition model constructed by the improved YOLOv8-seg model. Through the powerful image analysis capability of the model, the precise sub-image of each baggage and its type information can be output simultaneously, which realizes high-precision target segmentation in dense stacking scenarios and significantly improves the accuracy and robustness of baggage recognition; according to the mapping position of each baggage sub-image in the depth stacking image, the system generates point cloud data corresponding to each baggage. This process ensures the consistency between the two-dimensional recognition results and the three-dimensional recognition results through a precise coordinate conversion algorithm. The seamless integration of depth information provides reliable geometric data support for subsequent measurements. Based on the point cloud data of each piece of luggage, the system uses principal component analysis and axial bounding box calculation technology to accurately obtain the rotation angle and size information of the luggage. This technological breakthrough enables the system to overcome the limitations of traditional measurement methods in complex scenarios and achieve sub-centimeter measurement accuracy. Finally, the system integrates the type, rotation angle and size information of each piece of luggage into a structured stack recognition result. This standardized data output format is not only convenient for direct use by downstream systems, but more importantly, it establishes a complete data link from perception to decision-making, providing key technical support for the intelligent upgrade of the automated sorting system. Through innovative design in multiple links, the entire technical solution has achieved a significant improvement in the efficiency and accuracy of airline baggage handling.
[0062] Optionally, based on the above embodiments, before inputting the RGB stack image into the baggage recognition model, the following steps may be further included:
[0063] Collect various types of luggage images and construct a luggage image set, where the luggage image set contains RGB luggage images actually captured by a depth camera;
[0064] Perform image augmentation processing on each luggage image in the luggage dataset to obtain an expanded luggage image set;
[0065] Based on the expanded baggage image set, a training sample set is constructed, wherein each training sample is pre-labeled with a standard baggage area frame and baggage type;
[0066] The improved YOLOv8-seg model is trained multiple times using the training sample set to obtain a baggage recognition model.
[0067] Among them, in the improved YOLOv8-seg model, the Efficient Multi-Scale Attention Mechanism (EMA) network is used to replace the C2F (Context-aware Cross-level Fusion) network in the backbone network layer and feature enhancement layer of the standard YOLOv8-seg model, and during the model training process, the model parameters are adjusted based on the improved WIoU (Weighted Intersection over Unio) loss function.
[0068] Generally speaking, to more accurately and efficiently identify each baggage type within a stack, a professional, multi-category airline baggage dataset must be constructed. This dataset is constructed using a dual-track parallel strategy: First, standard RGB images of various baggage types are collected extensively through internet channels; second, high-precision depth cameras are used to capture images of baggage in real-world scenarios, ensuring the diversity and authenticity of the data sources. After systematic collection and screening, the resulting dataset contains 5,497 high-quality sample images across five categories, including 1,575 backpacks, 1,324 corrugated boxes, 1,215 travel bags, 1,093 woven bags, and 290 other special baggage types. This comprehensive dataset covers common baggage types in air transportation scenarios and lays a solid data foundation for model training.
[0069] Generally speaking, to further improve the quality of the dataset and the generalization ability of the model, a series of enhancement operations are performed on the original data, including rotation, scaling, contrast adjustment, and image effects in darkness and different light intensities. This multi-level, comprehensive enhancement strategy not only significantly improves the diversity and representativeness of the dataset, but also enables the trained model to easily cope with various complex situations in actual application scenarios, greatly improving the practical value of the system.
[0070] Generally speaking, the YOLOv8-seg instance segmentation model's network architecture consists of four main components: the input, the backbone network (also called the Backbone), the feature enhancement layer (also called the Neck layer), and the detection head layer (also called the Head layer). The Backbone is responsible for extracting features from the image and generally consists of a convolutional module, a Convolutional Block with Shortcut (CBS), a C2f network, and a Spatial Pyramid Pooling-Fast (SPPF) module. CBS consists of a convolutional layer, a batch processing layer, and a SiLU (Sigmoid Linear Unit) activation function layer; the C2f network mainly consists of an hourglass bottleneck convolution structure, which is the main module for learning residual features. It can have rich gradient flow information while ensuring lightweight; the SPPF module can convert feature maps of any size into feature vectors of fixed size; it mainly consists of a CBS convolutional layer and three serial Maxpooling layers (also called maximum pooling layers), which splices the feature map that has not been Maxpooled with the feature map obtained by each Maxpooling operation to achieve feature fusion.
[0071] The main function of the neck layer is to fuse multi-scale features to form a feature pyramid. It consists of two parts: FPN (Feature Pyramid Network) and PAN (Path Aggregation Network). FPN first extracts feature maps from a convolutional neural network and constructs a feature pyramid. It then uses a top-down approach to fuse the upsampled feature maps with coarser-grained feature maps, achieving the fusion of features at different levels. However, FPN alone may lack object location information. PAN, as a complement to FPN, adopts a bottom-up structure, fusing feature maps from different levels through a single convolutional layer, accurately preserving spatial information. The combination of FPN and PAN effectively integrates information flows from upstream and downstream of the network, thereby improving the network's detection performance. The head layer obtains the category and location information of target objects of different sizes based on feature maps of different sizes. Pixel-level masks are generated, each corresponding to a segmented piece of luggage.
[0072] Among them, in the feature extraction stage, Figure 2The highly efficient multi-scale attention mechanism network shown in the figure is added to the standard YOLOv8-seg backbone network, fusing shallow and deep feature maps. This allows the model to leverage the high-level semantic information of deep features while also preserving the spatial details of shallow features, significantly improving the network's feature extraction performance. By processing corresponding features through different branches of the EMA network, richer semantic feature information can be obtained, facilitating subsequent feature fusion and more accurate localization of key points.
[0073] Specifically, for any given input feature map, the EMA network divides the input into G sub-features through the Groups operation to learn different semantics. The EMA network adopts three parallel routes to extract the attention weight descriptors of the grouped feature maps. Among them, two parallel routes are located in the 1x1 branch and the third route is located in the 3x3 branch. In order to capture dependencies across all channels and reduce the computational budget, two 1D global average pooling operations are used in the 1x1 branch to encode the channels along the two spatial directions respectively, and only a single 3x3 kernel is stacked in the 3x3 branch to capture multi-scale feature representations. Then 2D global average pooling is used to encode the global spatial information in the output of the 1x1 branch.
[0074] Specifically, the formula for 2D global average pooling can be expressed as:
[0075] Among them, H and W are the tensors of feature x in two dimensions, i is the spatial index (horizontal coordinate) of the width direction (W) of the feature map, j is the spatial index (vertical coordinate) of the height direction (H) of the feature map, x c is the c-th channel of the input feature map.
[0076] Furthermore, based on the above embodiments, the improved YOLOv8-seg model is trained for multiple rounds using the training sample set to obtain a luggage recognition model, which may also include:
[0077] During each round of training, the current training sample is input into the improved YOLOv8-seg model in turn, and the prediction box marked in the current training sample output by the improved YOLOv8-seg model is obtained;
[0078] According to the predicted box and the standard luggage area box pre-labeled in the current training sample, the original WIoU loss function L is calculated based on the following formula: WIoUv1 ;
[0079]
[0080] Among them, W i 、H i are the width and height of the prediction box, W g、H g are the width and height of the standard luggage area frame, S u is the intersection area between the prediction box and the standard luggage area box, (x gt ,y gt ), (x, y) are the center coordinates of the standard luggage area frame and the prediction frame respectively; according to the formula After calculating the abnormality β, Indicates that the average LIoU is calculated for all images in the training batch. According to the following formula, the improved WIoU loss function L is calculated. WIoUv3 ;
[0081]
[0082] Among them, δ and α are model training parameters, which are used to control the gradient gain used in the back propagation process. Each training data is divided into several small batches, and each batch contains n number of sample data;
[0083] According to L WIoUv3 After adjusting the model parameters of the improved YOLOv8-seg model, the process returns to executing the operation of sequentially obtaining the current training samples and inputting them into the improved YOLOv8-seg model during each round of training until each round of training is completed.
[0084] After multiple rounds of training, obtain the model parameter weight files corresponding to each round of training, and obtain the target weight file from the multiple model parameter weight files;
[0085] Use the target weight file to configure the improved YOLOv8-seg model to obtain the luggage recognition model.
[0086] In this embodiment of the present invention, each training round of the model training process feeds the current batch of labeled samples into the improved YOLOv8-seg model for forward computation. This optimized model architecture, by embedding an EMA network and a modified loss function, outputs predictions including bounding box coordinates, class probabilities, confidence scores, and instance segmentation masks.
[0087] In a specific example, when fed an image of a complex scene containing a 28-inch large checked suitcase, a 20-inch carry-on suitcase, and a 16-inch backpack stacked together, the model demonstrated excellent performance: For the closely packed 28-inch and 20-inch suitcases, it clearly distinguished the edge contours of each case; for a 16-inch backpack partially obscured by the 28-inch suitcase, it fully restored the shape of the obscured portion; and for a 20-inch metal carry-on suitcase with reflective strips on its surface, it effectively overcame reflective interference and accurately identified the case structure. During training, the system dynamically compares the predicted boxes with the ground-truth annotations, calculates gradients based on a modified loss function, and updates the network parameters. This end-to-end training approach enables the model to gradually grasp the characteristics of luggage of various sizes, from small 18-inch carry-on suitcases to oversized 32-inch checked suitcases. Through a carefully designed network structure and optimization strategy, the entire training mechanism ensures the model's stable performance in a variety of complex scenarios, providing reliable technical support for automated airport baggage handling.
[0088] Accordingly, the improved YOLOv8-seg model's head layer uses the WIoU loss function instead of the original CIoU (Complete Intersection over Union) loss function. WIoU is a loss function based on a dynamic non-monotonic focusing mechanism that uses "outlier" to evaluate the quality of predicted boxes, further improving the model's generalization capabilities.
[0089] Specifically, the WIoUv3 loss function adopted in this method utilizes a non-monotonic focusing coefficient based on the WIoUv1 loss function. That is, the outlier degree β dynamically allocates the most suitable gradient gain according to the training situation. The gradient gain is negatively correlated with the outlier degree. The prediction box with a larger outlier degree is allocated a smaller gradient gain.
[0090] The calculation formula of WIoUv1 loss function is: Among them, W i 、H i Represents the width and height of the prediction box, W g 、H g Represent the width and height of the real box respectively, S u Represents the intersection area between the predicted box and the real box, which is used to calculate the ratio of the overlapping area to the predicted box area, (x gt ,y gt ), (x, y) represent the coordinates of the center point of the real box and the predicted box respectively. WIoUv3 defines the concept of outlier to characterize the quality of the predicted box. The smaller the outlier, the higher the quality of the predicted box. The calculation formula of outlier β is: In order to make bounding box regression focus on prediction boxes of average quality, a gradient gain is added to the loss function. The gradient gain is negatively correlated with the outlier degree, and prediction boxes with larger outliers are assigned smaller gradient gains. Applying the constructed non-monotonic focusing coefficient r to WIoUv1 can dynamically assign the most suitable gradient gain based on the training situation. The final WIoUv3 formula is: Here, α is an adjustment factor used to control the slope change; when β = δ, r = 1. Since LIoU changes dynamically, the quality classification standard of the anchor box is also dynamically adjusted, allowing WIoUv3 to make the gradient gain that best suits the current situation.
[0091] In the improved YOLOv8-seg model training process, the model optimization process forms a complete iterative loop. During training, the current batch of sample data is input into the architecturally improved YOLOv8-seg model. This model, with its optimized network structure, can effectively handle the various complex scenarios encountered in airline baggage recognition. After processing, the model outputs predictions, which are compared and analyzed with the actual annotations to drive automatic adjustment of model parameters. This training process repeats until all pre-set training rounds are completed. During this process, the model gradually learns the characteristics of various types of baggage, from large checked luggage to small carry-on luggage, demonstrating excellent adaptability to complex scenarios such as dense stacking and partial occlusion. Through this systematic training mechanism, the model ultimately achieves excellent recognition performance, accurately distinguishing closely stacked suitcases of different sizes and maintaining stable detection results for objects such as small backpacks. The design of the entire training process ensures the model's reliable performance in real-world applications, providing solid technical support for automated baggage handling.
[0092] Furthermore, during model training, the system saves the model parameter weight files generated after each epoch. Typically, these files are stored in .pt format and fully record the parameter states of each network layer, including key information such as convolution kernel weights and bias terms. The target weight file is selected based on a comprehensive performance evaluation on the validation set, using a combination of early stopping and model checkpointing strategies. During training, key indicators such as the average precision (mAP) and loss value of the model on an independent validation set are monitored in real time. When performance reaches a preset threshold or stops improving, the system automatically marks the current optimal weight file as the target weight. The technical advantage of this mechanism is that by retaining the intermediate states during training, it can avoid the risk of overfitting while ensuring that the model version with the best generalization ability is obtained.
[0093] In a specific example, during the model training process, the system will intelligently save and filter the optimal weight files. Taking the training of a dataset containing 5,000 airline luggage images as an example, the system will periodically save 50 intermediate weight files as checkpoints during the completion of 300 rounds of training. By comprehensively evaluating the performance of these weight files on the validation set, the system ultimately determined that the weight file generated by the 270th round of training was the optimal deployment version. This choice was based on the excellent performance demonstrated by this version: it can not only accurately identify the complex stacking of 28-inch large suitcases and 20-inch carry-on luggage, but also effectively reduce the missed detection rate of small-sized backpacks. This dynamic weight selection mechanism ensures that the final deployed model has both high recognition accuracy and good generalization ability by continuously monitoring model performance, perfectly balancing the various needs in practical applications.
[0094] Furthermore, during the model deployment phase, configuring the improved YOLOv8-seg model using a target weight file is a crucial step in model hardening. The target weight file, as the optimal parameter set selected during training, contains fully optimized weights for each network layer. These parameters capture the model's learned ability to recognize airline baggage features during training. Technically, the system loads these pre-trained weights to initialize the improved model architecture, which incorporates several innovative designs, including an attention mechanism for enhanced multi-scale feature extraction, an optimized segmentation prediction branch, and an intelligently weighted training strategy. By leveraging the model loading capabilities of the deep learning framework, the system accurately transfers the knowledge gained during training to the inference model, ensuring that the resulting baggage recognition model is both robust enough to handle complex scenarios and capable of meeting real-time requirements. This weight configuration mechanism ensures that the model fully utilizes its trained recognition capabilities during deployment, providing reliable technical support for airport automated baggage handling systems.
[0095] Furthermore, based on the above embodiments, obtaining a target weight file from a plurality of model parameter weight files may further include:
[0096] Extracting multiple types of target baggage images from the expanded baggage image set, and extracting baggage augmented images of each target baggage image under multiple geometric transformation scenarios and multiple illumination transformation scenarios;
[0097] Construct a verification dataset based on each target baggage image and each baggage augmented image;
[0098] The improved YOLOv8-seg model is configured using each of the model parameter weight files to obtain multiple alternative recognition models;
[0099] The validation data set is input into each alternative recognition model for model validation, and the model parameter weight file used by the alternative recognition model with the highest recognition accuracy is determined as the target weight file.
[0100] Specifically, during the model optimization process, the system first augments the original luggage image set using data augmentation techniques. This step involves applying various geometric transformations (rotation, scaling, etc.) and illumination transformations (brightness adjustment, contrast changes, etc.) to various target luggage images (such as suitcases and backpacks) to generate diverse augmented images. This data augmentation method effectively simulates various imaging conditions in real-world scenarios, providing sufficient test samples for subsequent model validation.
[0101] Furthermore, based on the enhanced image data, the system constructs a dedicated validation dataset. This dataset contains both the original samples and their corresponding transformed versions, ensuring coverage of all possible recognition scenarios. The validation set follows a strict sampling strategy to ensure a balanced distribution of samples across categories while maintaining data independence between the training and validation sets to avoid information leakage.
[0102] Furthermore, during the model validation phase, the system loads multiple intermediate weight files saved during training and configures them into the modified YOLOv8-seg model architecture to generate multiple candidate recognition models. While these candidate models share the same network architecture (including core components such as the EMA attention module and optimized feature pyramid), their parameters differ significantly due to their origins at different stages of the training process. For example, weights saved early on may perform better for luggage of common sizes, while weights saved later on may excel with luggage made of unusual materials or small sizes. Each candidate model is thoroughly tested on a unified validation set, which includes not only images of standard luggage but also challenging examples such as a 28-inch metal box stacked with a 20-inch woven bag, a dark backpack in low light, and unusually shaped luggage with complex surface patterns. The evaluation process comprehensively examines multiple performance metrics, including but not limited to the ability to distinguish overlapping objects, stability under varying lighting conditions, and adaptability to varying surface materials.
[0103] Through this systematic verification process, the weight file corresponding to the best-performing candidate model is ultimately determined as the target weight file. For example, the system may find that the weights saved in the 150th round perform well when processing regular suitcases, but the accuracy is insufficient when identifying backpacks smaller than 16 inches; while the weights saved in the 220th round improve the recognition rate of small backpacks by 15%, the processing of reflective metal surfaces has declined; and finally, the weights in the 180th round are selected as the target weights because they maintain stable performance in various test scenarios. This selection process is not a simple pursuit of the highest score for a single indicator, but through the design of a multi-dimensional evaluation scheme, it ensures that the selected model can work reliably in the actual airport environment, whether facing densely stacked luggage during the morning rush hour or special material bags under low light conditions at night. The entire verification mechanism provides a strong guarantee for the quality of the final deployed model by establishing a standardized stress testing environment.
[0104] Based on the above embodiments, while determining the baggage type, baggage rotation angle, and baggage size information of each baggage as the stack type recognition result of the airline baggage stacking area, the following further comprises:
[0105] Based on the three-dimensional point cloud data of each piece of luggage, the position description information of each piece of luggage in free space is determined, and the position description information is added to the stack type recognition result.
[0106] Furthermore, before each piece of luggage is conveyed to the stacking area via a conveyor belt for automated stacking, a pressure sensor installed on the conveyor belt can be used to measure the weight of each piece of luggage. Furthermore, after obtaining the stacking area's stack type recognition results, the center of gravity coordinates of each piece of luggage can be calculated based on this recognition result and the weight of each piece of luggage. Furthermore, by combining the center of gravity coordinates of each piece of luggage in the stacking area, the center of gravity coordinates of the entire stacking area can be calculated.
[0107] Specifically, the location description information of each piece of luggage, luggage size, luggage type, luggage rotation angle, and luggage weight can be input into a pre-trained center of gravity recognition model, and the center of gravity coordinates of each piece of luggage can be obtained through the center of gravity recognition model.
[0108] Afterwards, you can use the formula: Calculate the coordinates of the center of gravity of the entire stacking area (x c ,y c , z c ).
[0109] After obtaining the coordinates of the center of gravity of the entire stacking area, the entire stacking area can be monitored in real time for tipping risks. For example, the center of gravity can be detected to determine whether it exceeds the projection of the stacking area's bottom area, and whether the center of gravity height exceeds a preset height threshold. This allows for effective tipping risk warnings. Furthermore, the real-time coordinates of the center of gravity of the stacking area can be combined to assist in determining the stacking position of new luggage being conveyed on the conveyor belt within that stacking area.
[0110] Example 2
[0111] Figure 3 This is a flowchart for identifying airline baggage stack types, provided in the second embodiment of the present invention. This embodiment is an optimization of the previous embodiments. This embodiment specifically details the operations of "generating baggage point cloud data corresponding to each bag based on the mapping position of each baggage sub-image in the depth stacking image" and "calculating each baggage's rotation angle in the airline baggage stacking area and its dimensions based on the baggage point cloud data."
[0112] Correspondingly, such as Figure 3 As shown, the method includes:
[0113] S310 : Capture an RGB stacking image and a depth stacking image of the airline baggage stacking area using a depth camera.
[0114] S320: Input the RGB stacking image into a baggage recognition model to obtain a baggage sub-image of each bag in the aviation baggage stacking area in the RGB stacking image, as well as the baggage type of each bag. The baggage recognition model is obtained by training an improved YOLOv8-seg model.
[0115] S330 : Convert each pixel in the depth stacked image to the camera coordinate system according to the internal parameters of the depth camera, and generate three-dimensional point cloud data corresponding to the spatial area captured by the depth stacked image.
[0116] In the embodiments of the present invention, converting a depth image to a point cloud involves converting the depth value of each pixel into three-dimensional spatial coordinates. This conversion requires the intrinsic parameters of the depth camera, such as focal length and principal point coordinates, as well as possible extrinsic parameters such as the camera pose, i.e., rotation and translation matrices.
[0117] Specifically, there are four independent coordinate systems involved in the image processing process. If P w =(x w ,y w , z w ) represents a point in the world coordinate system, whose coordinates in the camera coordinate system are P c =(x c ,y c , z c ), P i =(x, y) represents the coordinates of the corresponding point in the image coordinate system, and p =(u, v) represents the coordinates in the pixel coordinate system. The transformation from the world coordinate system to the camera coordinate system is obtained by the homogeneous transformation matrix to obtain the coordinates of point P in the camera coordinate system. The specific formula is: Where R represents a 3×3 rotation matrix and t represents a three-dimensional translation vector. Then, the point in the three-dimensional space of the camera coordinate system is projected onto a plane in the two-dimensional space to form a two-dimensional image. The coordinates of a point in the camera coordinate system in the image coordinate system can be expressed as: Where f represents the focal length of the camera; (u0, v0) represents the coordinates of the origin of the imaging plane. The relationship between the point and the pixel coordinate system in the image coordinate system is: Therefore, a point in the pixel coordinate system can be expressed as: According to the intrinsic parameters of the depth camera, each pixel point p = (u, v) in the depth map can be converted to the position (x c ,y c , z c ), generates a 3D point cloud with the origin at the depth camera position, and its converted coordinates are P(X, Y, Z), accurately obtaining the point cloud data of the target.
[0118] S340 , based on the pairing relationship between each pixel in the depth stacking image and each pixel in the RGB stacking image, crop the cropped point cloud data corresponding to each baggage in the three-dimensional point cloud data according to each baggage sub-image.
[0119] In the embodiment of the present invention, in order to accurately correspond points at the same position in color and depth in space, the RGB image and the depth image are registered to obtain an RGB-D image. Figure 1 The conversion relationship between point mapping and color map is: rgb p =R·depth p +T, where R represents the rotation matrix, T represents the translation matrix, rgb p Represents the color camera coordinate system point, depth p Represents a point in the depth camera coordinate system.
[0120] Based on this registration relationship, the system implements a complete processing chain from 2D recognition to 3D segmentation: First, the trained YOLOv8-seg model is used to locate the precise boundaries of each piece of luggage in the RGB image, generating a luggage subimage marked with a binary mask. These 2D segmentation results are then mapped to a 3D point cloud using the aforementioned coordinate transformation relationship, extracting the 3D point set corresponding to each piece of luggage subimage from the point cloud data. In specific implementation, the system traverses all foreground pixels in the luggage subimage mask, finds their corresponding 3D coordinates in the point cloud using an R / T transformation, and finally aggregates these spatial points into independent cropped point cloud data. This semantically guided 3D segmentation method retains the advantages of deep learning in 2D image recognition while incorporating the precision of geometric data. This enables the system to accurately separate touching pieces of luggage, providing a reliable 3D data foundation for subsequent size measurement and pose estimation.
[0121] S350: Add the color information of each pixel in each luggage sub-image to the matched cropped point cloud data to generate luggage point cloud data corresponding to each luggage.
[0122] In an embodiment of the present invention, the system fuses the color information in the RGB image into the three-dimensional point cloud data through precise coordinate mapping, generating a color point cloud with complete (x, y, z, r, g, b) attributes.
[0123] Optionally, the process can leverage pre-calibrated camera parameters to establish correspondences between pixels and 3D points, employing a kd-tree spatial search and bilinear interpolation algorithm for efficient and accurate color matching. Color consistency verification and an adaptive enhancement module also ensure shading quality. For example, using a red suitcase as an example, the system not only reconstructs its 3D geometry but also preserves detailed features like surface logos, significantly increasing the visual information content of the point cloud data.
[0124] Specifically, this color point cloud generation technology, through the organic fusion of geometric and visual features, brings multiple advantages to the system: it maintains the accuracy of three-dimensional measurements while enhancing the expressiveness of texture details, enabling the system to better handle baggage items with complex surfaces. In actual sorting scenarios, this multimodal data not only supports precise dimensional measurement but also assists in visual feature verification, effectively improving recognition accuracy and providing a richer and more reliable data foundation for automated baggage handling. The entire colorization process uses intelligent quality control to ensure the precise matching of color and geometric structure, significantly improving the efficiency of subsequent processing links.
[0125] S360: Calculate the luggage rotation angle of each luggage in the airline luggage stacking area and luggage size information of each luggage based on the luggage point cloud data of each luggage.
[0126] S370: Determine the baggage type, baggage rotation angle, and baggage size information of each baggage as a stack type recognition result of the aviation baggage stacking area.
[0127] The technical solution of the embodiments of the present invention uses a depth camera to simultaneously capture RGB and depth images of the airline baggage stacking area. This multimodal data acquisition method effectively overcomes the limitations of single-sensor information. The RGB images are fed into a modified YOLOv8-seg model, achieving high-precision object segmentation and baggage type recognition in densely stacked scenarios. In the core processing phase, the system first converts the depth image into a precise 3D point cloud using the depth camera's intrinsic parameters. Then, based on RGB-D registration relationships, the system crops the 3D point sets corresponding to each piece of luggage from the point cloud data. Finally, it fuses color information to generate a color point cloud with complete attributes. This series of innovative processes forms a complete technical chain from 2D recognition to 3D reconstruction. The coordinate transformation algorithm ensures millimeter-level geometric accuracy, semantically guided segmentation effectively solves the problem of separating densely packed objects, and multimodal fusion significantly enhances the application value of point cloud data. Based on this high-quality point cloud data, the system accurately determines the spatial pose and physical dimensions of each piece of luggage through principal component analysis and axial bounding box calculation, ultimately outputting structured stack type recognition results. Through innovative design in multiple links, the entire technical solution has established a complete processing flow from data collection to intelligent decision-making, providing reliable technical support for the automated airline baggage sorting system and significantly improving the accuracy and efficiency of baggage handling.
[0128] Optionally, based on the above embodiments, calculating the luggage rotation angle of each luggage in the airline luggage stacking area and the luggage size information of each luggage based on the luggage point cloud data of each luggage may also include:
[0129] Based on the baggage type of each bag, a random sampling consensus algorithm is used to fit the maximum baggage plane corresponding to each baggage point cloud data. The matching baggage point cloud data is then denoised based on each maximum baggage plane.
[0130] Based on the denoised point cloud data of each piece of luggage, the principal component analysis method is used to calculate the luggage rotation angle of each piece of luggage;
[0131] According to the rotation angle of each bag, the baggage size information of each bag is calculated by fitting the minimum axis-aligned bounding box.
[0132] In an optional implementation of this embodiment, a random sampling consensus algorithm is used according to the baggage type of each bag to fit the maximum baggage plane corresponding to each baggage point cloud data, which may include:
[0133] Obtain target baggage point cloud data and target baggage type corresponding to the target baggage, and determine a distance threshold based on the target baggage type;
[0134] Randomly selecting a set number of sampling points from the target baggage point cloud data, and fitting the current iteration plane based on each sampling point;
[0135] Calculate the distance between each point in the target baggage point cloud data and the current iteration plane, take the point cloud points with distance values less than the distance threshold as the inliers of the current iteration plane, and count the total number of inliers in the current iteration plane;
[0136] If it is determined that the total number of the current inliers is greater than the total number of the inliers of the currently stored optimal iteration plane, the current iteration plane is updated to the optimal iteration plane;
[0137] Returning to the execution, randomly selecting a set number of sampling points in the target baggage point cloud data, and fitting the current iteration plane based on each sampling point until the current number of iterations reaches a preset iteration threshold;
[0138] The optimal iteration plane stored at the end of the iteration is determined as the maximum luggage plane corresponding to the target luggage.
[0139] In this embodiment, a distance threshold is dynamically determined based on the target baggage type. This distance threshold is used to classify point cloud points as either in-plane or out-of-plane. The inventors further considered that for luggage with high rigidity, such as suitcases, a relatively flat surface is typically present in the luggage point cloud data. Therefore, for this type of luggage, a smaller distance threshold can be set to effectively filter out out-of-plane noise points. For luggage with low rigidity, such as backpacks, the surface defined in the luggage point cloud data is not very flat. Therefore, a larger distance threshold can be set for this type of luggage to maximize the likelihood of filtering out luggage points as noise points.
[0140] In this embodiment, a mapping relationship between different baggage types and different distance thresholds can be pre-established. Alternatively, the baggage material can be determined through image recognition. Furthermore, the distance threshold corresponding to different baggage point cloud data can be calculated based on the baggage type and material using a pre-set distance calculation formula.
[0141] In the point cloud processing phase of this invention, the segmentation masks generated by the YOLOv8-seg model contain a certain degree of error, as well as the presence of noise points in the mapped point cloud data. The system uses the RANSAC (Random Sample Consensus) algorithm for robust fitting of the luggage's maximum plane. In specific implementation, the algorithm sets four key parameters: the point cloud to be processed D, the number of iterations Iter, the distance threshold Thresh, and the number of sample points Sample. In each iteration, the algorithm randomly selects Sample points to fit a candidate plane. It then calculates the distance from all points in point cloud D to this plane, marking points with a distance less than Thresh as inliers. By comparing the number of inliers in different candidate planes, the algorithm continuously updates the current optimal plane and its corresponding set of inliers. After Iter iterations, the final optimal plane is output as the luggage's maximum plane.
[0142] This plane fitting method has demonstrated strong adaptability in practical applications. Taking the processing of a locally deformed suitcase with a label on its surface as an example, although the point cloud data contains anomalies caused by the protrusion of the label and interference points from adjacent luggage, the RANSAC algorithm can still accurately identify the true maximum surface plane of the suitcase through multiple random sampling and verification. The advantage of this algorithm is that it does not rely on the integrity and accuracy of the point cloud data, but instead extracts the most reliable geometric features from noisy data through the principle of statistical consistency. This feature enables the system to effectively overcome the impact of segmentation errors and provide an accurate basic plane reference for subsequent size measurement and pose estimation, greatly improving the robustness and reliability of the entire processing flow.
[0143] Furthermore, in the 3D feature analysis phase of the present invention, the system first uses the PCA method to determine the main direction of the luggage based on the maximum plane point cloud data extracted by the RANSAC algorithm. The specific implementation process is: for m 3D point cloud data, X 3×m =[x1, x2, ..., x m ], each x is a 3-dimensional column vector; by Perform a decentralized operation (make the data mean 0) and calculate the covariance matrix After performing eigenvalue decomposition on the covariance matrix C to obtain the characteristic matrix (arranged in columns from large to small according to the eigenvalue), the first k columns of the characteristic matrix are taken to form the matrix P 3×k .
[0144] Among them, P is equivalent to a coordinate system, and each column in P is a coordinate axis; projecting the original data into the P coordinate system will obtain the reduced-dimensional data. The angle between the new coordinate system P and the original coordinate system is the rotation angle of the luggage.
[0145] Taking a standard rectangular suitcase as an example, even if the suitcase is tilted on the conveyor belt, the PCA method can accurately determine the direction of the long axis of the suitcase through the direction of the first principal component, thereby determining its rotation angle.
[0146] Furthermore, after obtaining the main direction of the luggage, the system further calculates the precise size of the luggage by constructing an AABB. The technical advantage of this step is that the main direction determined by PCA provides the optimal reference coordinate system for size measurement, so that the subsequent AABB calculation can accurately reflect the actual physical size of the luggage. For example, when processing a tilted suitcase, the system will first convert the point cloud data into the main direction coordinate system determined by PCA, and then search for extreme points along each coordinate axis. The final length, width and height measurement results are not affected by the angle of the luggage placement. This size measurement method based on main direction analysis effectively overcomes the error problem of traditional methods when measuring tilted objects, and provides accurate and reliable size data for the automated sorting system. The entire processing flow achieves accurate analysis of the spatial posture and physical size of the luggage through the organic combination of RANSAC, PCA and AABB.
[0147] Example 3
[0148] For ease of understanding, the specific application scenarios to which each embodiment of the invention is applicable are described. In this specific application scenario, in order to make the stack type identification of airline luggage more accurate, the embodiment of the present invention designs a complete stack type identification solution for airline luggage.
[0149] In this embodiment, two key steps are used to ensure accurate identification of the baggage stack type:
[0150] Specifically, in Figure 4 FIG. 4 is a flowchart showing the specific steps of the baggage stack type recognition module in a method for recognizing the stack type of airline baggage according to an embodiment of the present invention. Figure 4 As shown, this embodiment uses a depth camera to simultaneously capture RGB and depth images as input, and implements multi-stage processing using an improved YOLOv8-seg model. First, during the feature extraction stage, an EMA attention mechanism is embedded to enhance key region features. Subsequently, a FPN+PAN architecture is used to achieve multi-scale feature fusion. Finally, a detection head optimized with WioUv3 is used to output accurate segmentation masks and classification results. Throughout the entire process, the EMA module dynamically adjusts the weight distribution of features at different scales through a grouped multi-branch structure, while the WioUv3 loss function adaptively balances the training intensity between difficult and easy samples through an outlier assessment mechanism. After obtaining the binary mask, the system accurately aligns the segmentation result with the depth image and establishes a mapping relationship between 2D pixels and 3D space using the depth camera's calibration parameters. For each baggage area identified by the mask, the system extracts the corresponding depth information and converts it into 3D point cloud data with RGB color information. This conversion process strictly adheres to the principles of camera imaging geometry, ensuring that the 2D recognition results are accurately mapped to 3D space. This complete technology chain, from 2D semantic segmentation to 3D geometric reconstruction, combines the advantages of deep learning in object recognition with the precision of 3D visual measurement. Through point cloud generation guided by binary masks, the system effectively separates touching baggage items, providing a data foundation with both semantic information and geometric accuracy for subsequent dimensional measurement and pose estimation, fully meeting the high standards of automated airline baggage handling.
[0151] Further, in Figure 5 FIG. 4 is a flowchart showing the specific steps of the baggage measurement module in the stack type identification of airline baggage according to an embodiment of the present invention. Figure 5 This embodiment uses 3D point cloud data processing technology to achieve accurate baggage measurement. First, based on the color point cloud data output by the recognition module, the RANSAC algorithm is used to fit the maximum surface plane of the baggage, effectively filtering out noise points and outliers. Then, principal component analysis (PCA) is used to determine the principal orientation of the baggage in 3D space and calculate its placement angle. Finally, an axially aligned bounding box (AABB) is constructed to measure the length, width, and height of the baggage along the principal orientation coordinate system. Throughout the entire process, the RANSAC algorithm ensures robust plane fitting through iterative optimization, PCA analysis accurately extracts the main eigenvectors of the point cloud distribution, and AABB measurement obtains precise physical dimensions in the optimal coordinate system. This progressive "plane extraction - orientation analysis - dimension measurement" process achieves a complete conversion from raw point cloud to structured measurement results, providing reliable pose and dimension information for automated sorting systems.
[0152] In addition, combined again Figure 2 It can be seen that in Figure 2 The article details the structural design of the efficient multi-scale attention mechanism module employed in an embodiment of the present invention. This module utilizes an innovative multi-branch parallel architecture and consists of three core processing paths: the first path extracts local features through 1×1 convolutions, complemented by bidirectional horizontal and vertical global average pooling to capture long-range spatial dependencies; the second path utilizes 3×3 convolutions to acquire multi-scale contextual information; and the third path uses a grouping mechanism to divide the feature map into multiple sub-features for independent processing. The outputs of these three paths undergo feature recombination and weighted fusion to ultimately generate enhanced features with multi-scale awareness. Particularly noteworthy is the way this architecture is embedded in the C2F module, enabling the network to dynamically adjust its focus on features of different scales during forward propagation. For example, when processing a mixed stack of 28-inch suitcases and 20-inch carry-on luggage, it maintains equal multi-scale awareness of both the overall structure of the large luggage and local features such as small handles. This design significantly improves the model's feature extraction capabilities in complex airline luggage scenarios, laying a solid foundation for subsequent accurate segmentation and measurement.
[0153] Furthermore, through the ingenious combination of the above steps, accurate identification of the stack type of airline luggage can be achieved, achieving the following effective results:
[0154] (1) Significantly improve recognition capabilities in densely stacked scenarios. Through the synergistic effect of multi-scale feature fusion and attention mechanism, the system can effectively distinguish closely adjacent luggage items, solving the recognition difficulties of traditional methods in complex stacking scenarios.
[0155] (2) Achieve high-precision three-dimensional dimension measurement. The three-dimensional reconstruction technology based on depth information combined with intelligent geometric analysis methods ensures accurate measurement of the dimensions of various types of luggage, meeting the strict requirements of aviation logistics.
[0156] (3) Enhance the stability of the system in complex environments. The dynamically optimized loss function design enables the system to adapt to different lighting conditions and occlusion situations, maintaining reliable recognition performance in various practical application scenarios.
[0157] (4) Build a complete intelligent processing system. The integrated solution from image acquisition to 3D measurement has achieved intelligent upgrades for the entire process of air baggage handling and provided comprehensive technical support for the automated sorting system.
[0158] Example 4
[0159] Figure 6 This is a schematic diagram of the structure of a device for identifying the stack type of airline luggage provided by the fourth embodiment of the present invention. Figure 6As shown, the device includes: a data acquisition module 610, a luggage identification module 620, a point cloud generation module 630, a feature analysis module 640, and an information output module 650, wherein:
[0160] The data acquisition module 610 is configured to acquire an RGB stacking image and a depth stacking image of the airline baggage stacking area through a depth camera;
[0161] The baggage recognition module 620 is configured to input the RGB stack image into a baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag in the airline baggage stack area. The baggage recognition model is obtained by training an improved YOLOv8-seg model.
[0162] The point cloud generation module 630 is used to generate baggage point cloud data corresponding to each baggage according to the mapping position of each baggage sub-image in the depth stacking image;
[0163] A feature analysis module 640 is configured to calculate the luggage rotation angle of each bag in the airline baggage stacking area and the luggage size information of each bag based on the baggage point cloud data of each bag;
[0164] The information output module 650 is used to determine the baggage type, baggage rotation angle and baggage size information of each baggage as the stack type recognition result of the aviation baggage stacking area.
[0165] The technical solution of the embodiment of the present invention uses a depth camera to synchronously collect RGB stacking images and depth stacking images of the aviation baggage stacking area. This multimodal data acquisition method effectively overcomes the limitations of single sensor information and provides a more comprehensive data basis for subsequent processing; the RGB stacking image is input into the baggage recognition model constructed by the improved YOLOv8-seg model. Through the powerful image analysis capability of the model, the precise sub-image of each baggage and its type information can be output simultaneously, which realizes high-precision target segmentation in dense stacking scenarios and significantly improves the accuracy and robustness of baggage recognition; according to the mapping position of each baggage sub-image in the depth stacking image, the system generates point cloud data corresponding to each baggage. This process ensures the consistency between the two-dimensional recognition results and the three-dimensional recognition results through a precise coordinate conversion algorithm. The seamless integration of depth information provides reliable geometric data support for subsequent measurements. Based on the point cloud data of each piece of luggage, the system uses principal component analysis and axial bounding box calculation technology to accurately obtain the rotation angle and size information of the luggage. This technological breakthrough enables the system to overcome the limitations of traditional measurement methods in complex scenarios and achieve sub-centimeter measurement accuracy. Finally, the system integrates the type, rotation angle and size information of each piece of luggage into a structured stack recognition result. This standardized data output format is not only convenient for direct use by downstream systems, but more importantly, it establishes a complete data link from perception to decision-making, providing key technical support for the intelligent upgrade of the automated sorting system. Through innovative design in multiple links, the entire technical solution has achieved a significant improvement in the efficiency and accuracy of airline baggage handling.
[0166] Furthermore, based on the above embodiments, the apparatus for identifying the stack type of airline luggage further includes:
[0167] An image set construction module is used to collect multiple types of baggage images and construct a baggage image set before inputting the RGB stacked images into the baggage recognition model. The baggage image set includes RGB baggage images actually captured by the depth camera.
[0168] An image set expansion module is used to perform image augmentation processing on each baggage image in the baggage dataset to obtain an expanded baggage image set;
[0169] A training sample set construction module is used to construct a training sample set based on the expanded baggage image set, wherein each training sample is pre-labeled with a standard baggage area frame and a baggage type;
[0170] The model training module is used to perform multiple rounds of training on the improved YOLOv8-seg model using the training sample set to obtain a baggage recognition model;
[0171] In the improved YOLOv8-seg model, the EMA network is used to replace the upper C2F network in the backbone network layer and feature enhancement layer of the standard YOLOv8-seg model, and during the model training process, the model parameters are adjusted based on the improved WIoU loss function.
[0172] Furthermore, based on the above embodiments, the model training module can be specifically used to:
[0173] During each round of training, the current training sample is input into the improved YOLOv8-seg model in turn, and the prediction box marked in the current training sample output by the improved YOLOv8-seg model is obtained;
[0174] According to the predicted box and the standard luggage area box pre-labeled in the current training sample, the original WIoU loss function L is calculated based on the following formula: WIoUv1 ;
[0175]
[0176] Among them, W i 、H i are the width and height of the prediction box, W g 、H g are the width and height of the standard luggage area frame, S u is the intersection area between the prediction box and the standard luggage area box, (x gt ,y gt ), (x, y) are the center coordinates of the standard baggage area frame and the predicted frame respectively;
[0177] According to the formula After calculating the abnormality β, the improved WIoU loss function L is calculated according to the following formula: WIoUv3 ;
[0178]
[0179] Among them, δ and α are model training parameters, which are used to control the gradient gain used in the back propagation process. Each training data is divided into several small batches, and each batch contains n number of sample data;
[0180] According to L WIoUv3 After adjusting the model parameters of the improved YOLOv8-seg model, the process returns to executing the operation of sequentially obtaining the current training samples and inputting them into the improved YOLOv8-seg model during each round of training until each round of training is completed.
[0181] After multiple rounds of training, obtain the model parameter weight files corresponding to each round of training, and obtain the target weight file from the multiple model parameter weight files;
[0182] Use the target weight file to configure the improved YOLOv8-seg model to obtain the luggage recognition model.
[0183] Furthermore, based on the above embodiments, the model training module can be further used to:
[0184] Extracting multiple types of target baggage images from the expanded baggage image set, and extracting baggage augmented images of each target baggage image under multiple geometric transformation scenarios and multiple illumination transformation scenarios;
[0185] Construct a verification dataset based on each target baggage image and each baggage augmented image;
[0186] The improved YOLOv8-seg model is configured using each of the model parameter weight files to obtain multiple alternative recognition models;
[0187] The validation data set is input into each alternative recognition model for model validation, and the model parameter weight file used by the alternative recognition model with the highest recognition accuracy is determined as the target weight file.
[0188] Based on the above embodiments, the point cloud generation module 630 may further include:
[0189] According to the intrinsic parameters of the depth camera, each pixel in the depth stacked image is converted to the camera coordinate system to generate 3D point cloud data corresponding to the spatial area captured by the depth stacked image;
[0190] Based on the pairing relationship between each pixel in the depth stacking image and each pixel in the RGB stacking image, the cropped point cloud data corresponding to each baggage is cropped from the 3D point cloud data according to each baggage sub-image;
[0191] The color information of each pixel in each luggage sub-image is added to the matching cropped point cloud data to generate luggage point cloud data corresponding to each luggage.
[0192] Based on the above embodiments, the feature analysis module 640 may further include:
[0193] A denoising processing unit is used to fit the maximum luggage plane corresponding to each piece of luggage point cloud data using a random sampling consensus algorithm according to the luggage type of each piece of luggage, and to perform denoising processing on the matching luggage point cloud data according to each maximum luggage plane;
[0194] A rotation angle determination unit, configured to calculate the luggage rotation angle of each luggage piece by using a principal component analysis method based on the denoised point cloud data of each luggage piece;
[0195] The luggage size information calculation unit is used to calculate the luggage size information of each luggage by fitting a minimum axis-aligned bounding box according to the rotation angle of each luggage.
[0196] Furthermore, based on the above embodiments, the denoising processing unit may be further configured to: obtain target baggage point cloud data and target baggage type corresponding to the target baggage, and determine a distance threshold according to the target baggage type;
[0197] Randomly selecting a set number of sampling points from the target baggage point cloud data, and fitting the current iteration plane based on each sampling point;
[0198] Calculate the distance between each point in the target baggage point cloud data and the current iteration plane, take the point cloud points with distance values less than the distance threshold as the inliers of the current iteration plane, and count the total number of inliers in the current iteration plane;
[0199] If it is determined that the total number of the current inliers is greater than the total number of the inliers of the currently stored optimal iteration plane, the current iteration plane is updated to the optimal iteration plane;
[0200] Returning to the execution, randomly selecting a set number of sampling points in the target baggage point cloud data, and fitting the current iteration plane based on each sampling point until the current number of iterations reaches a preset iteration threshold;
[0201] The optimal iteration plane stored at the end of the iteration is determined as the maximum luggage plane corresponding to the target luggage.
[0202] An apparatus for identifying the stack type of airline baggage provided in an embodiment of the present invention can execute a method for identifying the stack type of airline baggage provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.
[0203] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0204] Example 5
[0205] Figure 7A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0206] As shown in FIG. X , the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, that is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0207] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0208] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a method for identifying the stack type of airline baggage as described in any embodiment of the present invention.
[0209] That is:
[0210] Use a depth camera to capture RGB stacking images and depth stacking images of the airline baggage stacking area;
[0211] The RGB stack image is input into the baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag in the airline baggage stack area. The baggage recognition model is obtained by training the improved YOLOv8-seg model.
[0212] Generate baggage point cloud data corresponding to each baggage according to the mapping position of each baggage sub-image in the depth stacking image;
[0213] Based on the baggage point cloud data of each bag, calculate the baggage rotation angle of each bag in the airline baggage stacking area and the baggage size information of each bag;
[0214] The baggage type, baggage rotation angle, and baggage size information of each baggage are determined as the stack type recognition result of the aviation baggage stacking area.
[0215] In some embodiments, a method for identifying the stack type of airline baggage may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps for identifying the stack type of airline baggage described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform stack type identification of airline baggage via any other suitable means (e.g., via firmware).
[0216] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0217] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0218] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0220] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0221] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0222] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0223] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for identifying the stack type of airline luggage, characterized in that: include: Use a depth camera to capture RGB stacking images and depth stacking images of the airline baggage stacking area; The RGB stack image is input into the baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag in the airline baggage stack area. The baggage recognition model is obtained by training the improved YOLOv8-seg model. Generate baggage point cloud data corresponding to each baggage according to the mapping position of each baggage sub-image in the depth stacking image; Based on the baggage point cloud data of each bag, calculate the baggage rotation angle of each bag in the airline baggage stacking area and the baggage size information of each bag; The baggage type, baggage rotation angle, and baggage size information of each baggage are determined as the stack type recognition result of the aviation baggage stacking area.
2. The method according to claim 1, characterized in that Before feeding the RGB stack image into the baggage recognition model, the following steps are also included: Collect various types of luggage images and construct a luggage image set, where the luggage image set contains RGB luggage images actually captured by a depth camera; Perform image augmentation processing on each luggage image in the luggage dataset to obtain an expanded luggage image set; Based on the expanded baggage image set, a training sample set is constructed, wherein each training sample is pre-labeled with a standard baggage area frame and baggage type; The improved YOLOv8-seg model is trained multiple times using the training sample set to obtain a baggage recognition model. Among them, in the improved YOLOv8-seg model, the efficient multi-scale attention mechanism EMA network is used to replace the context-aware cross-level fusion C2F network in the backbone network layer and feature enhancement layer of the standard YOLOv8-seg model, and during the model training process, the model parameters are adjusted based on the improved weighted intersection-over-union (WIoU) loss function.
3. The method according to claim 2, characterized in that The improved YOLOv8-seg model is trained multiple times using the training sample set to obtain a baggage recognition model, including: During each round of training, the current training sample is input into the improved YOLOv8-seg model in turn, and the prediction box marked in the current training sample output by the improved YOLOv8-seg model is obtained; According to the predicted box and the standard luggage area box pre-labeled in the current training sample, the original WIoU loss function L is calculated based on the following formula: WIoUv1 ; Among them, W i 、H i are the width and height of the prediction box, W g 、H g are the width and height of the standard luggage area frame, S u is the intersection area between the prediction box and the standard luggage area box, (x gt ,y gt ), (x, y) are the center coordinates of the standard baggage area frame and the predicted frame respectively; According to the formula After calculating the abnormality β, the improved WIoU loss function L is calculated according to the following formula: WIoUv3 ; Among them, δ and α are model training parameters, which are used to control the gradient gain used in the back propagation process. Each training data is divided into several small batches, and each batch contains n number of sample data; According to L WIoUv3 After adjusting the model parameters of the improved YOLOv8-seg model, the process returns to executing the operation of sequentially obtaining the current training samples and inputting them into the improved YOLOv8-seg model during each round of training until each round of training is completed. After multiple rounds of training, obtain the model parameter weight files corresponding to each round of training, and obtain the target weight file from the multiple model parameter weight files; Use the target weight file to configure the improved YOLOv8-seg model to obtain the luggage recognition model.
4. The method according to claim 3, characterized in that Get the target weight file from multiple model parameter weight files, including: Extracting multiple types of target baggage images from the expanded baggage image set, and extracting baggage augmented images of each target baggage image under multiple geometric transformation scenarios and multiple illumination transformation scenarios; Construct a verification dataset based on each target baggage image and each baggage augmented image; The improved YOLOv8-seg model is configured using each of the model parameter weight files to obtain multiple alternative recognition models; The validation data set is input into each alternative recognition model for model validation, and the model parameter weight file used by the alternative recognition model with the highest recognition accuracy is determined as the target weight file.
5. The method according to any one of claims 1 to 4, wherein: According to the mapping position of each baggage sub-image in the depth stacking image, baggage point cloud data corresponding to each baggage is generated, including: According to the intrinsic parameters of the depth camera, each pixel in the depth stacked image is converted to the camera coordinate system to generate 3D point cloud data corresponding to the spatial area captured by the depth stacked image; Based on the pairing relationship between each pixel in the depth stacking image and each pixel in the RGB stacking image, the cropped point cloud data corresponding to each baggage is cropped from the 3D point cloud data according to each baggage sub-image; The color information of each pixel in each luggage sub-image is added to the matching cropped point cloud data to generate luggage point cloud data corresponding to each luggage.
6. The method according to any one of claims 1 to 4, characterized in that Based on the baggage point cloud data of each bag, the baggage rotation angle of each bag in the airline baggage stacking area and the baggage size information of each bag are calculated, including: Based on the baggage type of each bag, a random sampling consensus algorithm is used to fit the maximum baggage plane corresponding to each baggage point cloud data. The matching baggage point cloud data is then denoised based on each maximum baggage plane. Based on the denoised point cloud data of each piece of luggage, the principal component analysis method is used to calculate the luggage rotation angle of each piece of luggage; According to the rotation angle of each bag, the baggage size information of each bag is calculated by fitting the minimum axis-aligned bounding box.
7. The method according to claim 6, characterized in that According to the baggage type of each bag, a random sampling consensus algorithm is used to fit the maximum baggage plane corresponding to each baggage point cloud data, including: Obtain target baggage point cloud data and target baggage type corresponding to the target baggage, and determine a distance threshold based on the target baggage type; Randomly selecting a set number of sampling points from the target baggage point cloud data, and fitting the current iteration plane based on each sampling point; Calculate the distance between each point in the target baggage point cloud data and the current iteration plane, take the point cloud points with distance values less than the distance threshold as the inliers of the current iteration plane, and count the total number of inliers in the current iteration plane; If it is determined that the total number of the current inliers is greater than the total number of the inliers of the currently stored optimal iteration plane, the current iteration plane is updated to the optimal iteration plane; Returning to the execution, randomly selecting a set number of sampling points in the target baggage point cloud data, and fitting the current iteration plane based on each sampling point until the current number of iterations reaches a preset iteration threshold; The optimal iteration plane stored at the end of the iteration is determined as the maximum luggage plane corresponding to the target luggage.
8. A device for identifying the stack type of airline luggage, characterized in that: include: A data acquisition module, used to collect RGB stacking images and depth stacking images of the aviation baggage stacking area through a depth camera; The baggage recognition module is used to input the RGB stack image into the baggage recognition model to obtain a baggage sub-image of each bag in the RGB stack image and the baggage type of each bag. The baggage recognition model is obtained by training the improved YOLOv8-seg model. A point cloud generation module is used to generate baggage point cloud data corresponding to each bag according to the mapping position of each baggage sub-image in the depth stacking image; The feature analysis module is used to calculate the luggage rotation angle of each bag in the airline baggage stacking area and the luggage size information of each bag based on the baggage point cloud data of each bag; The information output module is used to determine the baggage type, baggage rotation angle and baggage size information of each baggage as the stack type recognition result of the aviation baggage stacking area.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the method for identifying the stack type of aviation baggage according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for identifying the stack type of aviation baggage according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Airport luggage carrying method and device based on simulation and control, and electronic equipment
CN121132705A
Luggage case visual stacking size calculation method based on depth map
CN122156291A