Logistics product anomaly detection and elimination system and method
By combining multimodal image information with the CNN-ViT hybrid model, the problems of low efficiency and insufficient robustness in anomaly detection of logistics products are solved, and high-precision, real-time anomaly detection and elimination are achieved on high-speed production lines. It is suitable for logistics bag products with variable shapes and soft materials, reduces dependence on the lighting environment, and improves detection accuracy and robustness.
Patent Information
- Application Number
- CN202510789236.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are inefficient and highly subjective in detecting anomalies in logistics products, and cannot meet the needs of high-speed and large-scale production. They also have poor detection effects on logistics bag products with variable shapes, soft materials, and complex defect types. They are not robust enough and are greatly affected by the external lighting environment. The complex deep learning model inference requires a large amount of computational complexity, making it difficult to deploy on high-speed assembly lines.
Multimodal image information is combined with the CNN-ViT hybrid model. Multimodal information is obtained through the image acquisition module. The defect data enhancement module is used to generate synthetic defect images. The real defect image samples are integrated to build a training data set. Edge computing is deployed to optimize the defect detection model. The intelligent rejection module is used to control the robotic arm to reject abnormal products.
It realizes the high-precision, high-real-time detection and accurate and stable elimination of various abnormal logistics bags on high-speed assembly lines, overcomes the limitations of single modality and single model, meets the real-time requirements of high-speed assembly lines, reduces dependence on lighting environment, and improves the generalization ability and adaptability of the model.
Smart Images

Figure CN120707494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of product detection technology, and in particular to a system and method for detecting and rejecting abnormalities in logistics products. Background Art
[0002] Traditionally, anomaly detection for logistics products (such as logistics bags) relies on manual visual inspection, with operators visually inspecting the bags on the production line. However, this method is inefficient, highly subjective, prone to fatigue and missed detections, and cannot meet the high-speed, high-volume requirements of modern logistics. Existing methods also employ machine vision or visible light deep learning, using fixed templates to compare with actual images to identify products with abnormal shapes and sizes. However, these methods exhibit poor detection performance for logistics bags with diverse shapes, soft materials, and complex defect types (such as fibers and minor tears). These methods also lack robustness and are significantly affected by external lighting conditions. Furthermore, complex deep learning models require high computational complexity for inference. Therefore, there is an urgent need for a method for detecting abnormal defects in logistics products, such as logistics bags, that is more accurate and robust, less susceptible to external lighting conditions, and suitable for high-speed production lines. Summary of the Invention
[0003] The purpose of the present invention is to provide a logistics product abnormality detection and rejection system and method, which can realize high-precision, high-real-time detection and accurate and stable rejection of various types of abnormal logistics bags on high-speed production lines.
[0004] The technical solutions provided by the present invention are as follows:
[0005] In a first aspect, the present application provides a logistics product anomaly detection and rejection system, comprising:
[0006] An image acquisition module, configured with at least two image acquisition devices, is used to synchronously acquire multimodal information of the products to be inspected on the assembly line at a preset acquisition frequency;
[0007] A defect data enhancement module is used to generate synthetic defect images based on at least two image synthesis models and fuse real defect image samples to construct a training dataset;
[0008] an edge computing deployment module, configured to optimize a defect detection model of a defect detection module according to the defect data enhancement module, and input the multimodal information transmitted by the image acquisition module into the defect detection model;
[0009] a defect detection module, configured to extract and fuse at least two features from the multimodal information, and input the fused features into the defect detection model to identify abnormal products;
[0010] The intelligent rejection module is used to reject abnormal products identified by the defect detection module by controlling the robotic arm.
[0011] In some embodiments, the image acquisition device includes a visible light camera, an infrared camera, and a depth camera, and the multimodal information includes a visible light image, an infrared thermal imaging image, and a depth image;
[0012] The image acquisition module includes:
[0013] an image preprocessing unit, connected to the visible light camera, the infrared camera, and the depth camera, respectively, for preprocessing the visible light image, the infrared thermal imaging image, and the depth image;
[0014] A synchronization control unit is used to synchronize the visible light camera, the infrared camera and the depth camera through a synchronization trigger.
[0015] In some embodiments, the image acquisition module is further configured to calculate the preset acquisition frequency based on pre-input assembly line speed and product spacing, so that the multimodal information is acquired at least once for each product to be inspected.
[0016] In some embodiments, the image synthesis model includes an image generation model and an adversarial network model;
[0017] The defect data enhancement module includes:
[0018] A model construction unit, configured to respectively construct the image generation model and the adversarial network model;
[0019] a first synthesis unit, configured to input the acquired normal product image and defect description text into the image generation model to generate a first synthesized defect image;
[0020] The adversarial network includes a generator and a discriminator. The generator is used to generate a second synthetic defect image based on the normal product image and input random noise; the discriminator is used to distinguish the second synthetic defect image from the real defect image and optimize the generator based on the distinction result.
[0021] In some embodiments, the defect detection module includes:
[0022] A first feature extraction unit, configured to extract local detail features of the multimodal information through a convolutional neural network and output a feature map;
[0023] a second feature extraction unit configured to segment the multimodal information into a plurality of image blocks using a visual transformer, capture long-range dependencies and global context information corresponding to each of the image blocks, identify widely distributed or irregular defects, and output a feature vector;
[0024] The feature fusion unit is used to fuse the feature map and the feature vector in the deep layer of the network to obtain a fused feature.
[0025] In some embodiments, further comprising:
[0026] A positioning and classification unit is used to determine the defect classification and positioning according to the fusion features.
[0027] In some embodiments, the edge computing deployment module includes a model quantization unit, an operator fusion unit, and a pruning unit, which are used to optimize the defect detection model.
[0028] In some embodiments, the intelligent rejection module includes:
[0029] a path planning unit, configured to plan an operating path of the robotic arm based on the abnormal products identified by the defect detection module and the speed of the assembly line;
[0030] A confirmation unit is used to confirm whether the abnormal product is removed from the assembly line based on the photoelectric switch.
[0031] In some embodiments, a negative pressure adsorption device and a pressure sensor are provided at the end of the robotic arm, and the intelligent rejection module is further used to adjust the adsorption negative pressure of the negative pressure adsorption device according to the type of abnormal product and the type of defect.
[0032] In a second aspect, the present application provides a method for detecting and eliminating abnormalities in logistics products, comprising the steps of:
[0033] Synchronously acquiring multimodal information of the product to be inspected on the production line using at least two image acquisition devices at a preset acquisition frequency;
[0034] Extracting at least two features and fusing the features on the multimodal information, and inputting the fused features into the defect detection model to identify abnormal products, wherein the defect detection model is configured to be trained using a training dataset constructed by fusing synthetic defect images generated by at least two image synthesis models with real defect image samples;
[0035] Abnormal products identified by the defect detection module are removed by controlling the robotic arm.
[0036] According to the present invention, a logistics product abnormality detection and rejection system and method can achieve high-precision, high-real-time detection and accurate, stable rejection of various types of abnormal logistics bags on high-speed assembly lines. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The preferred implementation scheme will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of this solution.
[0038] Figure 1 It is a schematic diagram of the overall structure of an embodiment of the present invention;
[0039] Figure 2 Schematic diagram of a system framework according to an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of an image acquisition module according to an embodiment of the present invention;
[0041] Figure 4 is a schematic diagram of a defect data enhancement module according to an embodiment of the present invention;
[0042] Figure 5 Schematic diagram of a defect detection module according to an embodiment of the present invention. DETAILED DESCRIPTION
[0043] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can derive other drawings and other embodiments based on these drawings without inventive effort.
[0044] To simplify the drawings, only portions relevant to the present invention are schematically depicted in each figure; these do not represent the actual structure of the product. Furthermore, to simplify the drawings and facilitate understanding, in some figures, only one component with the same structure or function is schematically depicted or labeled. As used herein, "one" not only means "only one" but also "more than one."
[0045] Traditionally, manual visual inspection has been used to detect anomalies in logistics products (such as logistics bags, which will be used as an example below). Operators visually inspect the bags on the production line. However, this method is inefficient, highly subjective, prone to fatigue and missed inspections, and cannot meet the high-speed, high-volume requirements of modern logistics. Existing technologies also employ machine vision for detection, using fixed templates to compare with actual images to identify products with abnormal shapes and sizes. However, this method has poor detection effectiveness and lacks robustness for logistics bags with variable shapes, soft materials, and complex defect types (such as fibers and minor tears). Existing technologies also include methods for detection through single-modal (visible light) deep learning. However, this method is highly data-dependent, and deep learning models require a large number of labeled defect samples for training. In actual production, samples of abnormal logistics bags (especially those with diverse defects) are often scarce, resulting in insufficient model generalization. This method also has low environmental sensitivity and, using only visible light, is susceptible to interference from environmental factors such as lighting changes, shadows, and reflections. Its ability to detect specific defects (such as subtle scratches on transparent films and adhesions of similar colors) is limited. Furthermore, complex deep learning models require a large amount of computational inference, and achieving low-latency (<50ms) detection on high-speed production lines (e.g., >1 m / s) presents technical bottlenecks, making it difficult to deploy at the edge. Therefore, there is an urgent need for a detection method for abnormal defects in logistics products such as logistics bags that is more accurate and robust, less affected by the external lighting environment, and suitable for high-speed production lines. This solution combines the advantages of multimodal image information with the CNN-ViT hybrid model, effectively overcoming the limitations of single modalities and model types, enabling higher accuracy in identifying various types of abnormal logistics bags on high-speed assembly lines. Through model optimization and edge computing deployment, the system's single-frame detection time is strictly controlled within 50ms, meeting the real-time requirements of high-speed logistics and packaging lines (e.g., ≥1.5 m / s). Furthermore, multimodal fusion reduces dependence on a single light source or environmental conditions. The rich dataset generated by the hybrid data augmentation strategy significantly improves the model's generalization and adaptability to lighting variations, diverse logistics bag shapes, and new types of defects, resulting in greater robustness. This solution is described in detail below with accompanying figures:
[0046] In one embodiment, the reference specification Figure 1 and 2 The present application provides a logistics product anomaly detection and rejection system, including an image acquisition module 10 , a defect data enhancement module 20 , an edge computing deployment module 30 , a defect detection module 40 and an intelligent rejection module 50 .
[0047] The image acquisition module 10 is equipped with at least two image acquisition devices and is used to synchronously acquire multimodal information of products to be inspected on the assembly line at a preset acquisition frequency. The defect data enhancement module 20 is used to generate synthetic defect images based on at least two image synthesis models and fuse real defect image samples to construct a training dataset. The edge computing deployment module 30 is used to optimize the defect detection model of the defect detection module 40 based on the defect data enhancement module 20 and input the multimodal information transmitted by the image acquisition module 10 into the defect detection model. The defect detection module 40 is used to extract at least two features from the multimodal information and perform feature fusion. The fused features are then input into the defect detection model to identify abnormal products. The intelligent rejection module 50 is used to control a robotic arm to reject abnormal products identified by the defect detection module 40.
[0048] This application does not limit the specific type of image acquisition device; the choice can be based on practical conditions. In one specific implementation, the image acquisition device includes a visible light camera (resolution ≥ 5 megapixels) and an infrared camera (detection band 8-14 μm, thermal sensitivity ≤ 50 mK). Accordingly, the multimodal information acquired includes visible light images and infrared thermal images. In another specific implementation, the image acquisition device includes a visible light camera, an infrared camera, and a depth camera. Accordingly, the multimodal information includes visible light images, infrared thermal images, and depth images. Visible light images can be used to capture structural defects (such as the edge contours of cracks and holes), while infrared images can enhance the detection of non-structural defects (such as loose fibers and adhesions) by leveraging temperature differences between defect areas and the background (such as fiber areas heated by friction or temperature differences caused by differences in air permeability).
[0049] In a specific implementation, Figure 3 As shown, the image acquisition module 10 includes: an image preprocessing unit, which is connected to the visible light camera, the infrared camera and the depth camera respectively, and is used to preprocess the visible light image, the infrared thermal imaging image and the depth image; a synchronization control unit, which is used to synchronize and control the visible light camera, the infrared camera and the depth camera through a synchronization trigger.
[0050] The image acquisition module 10 is also configured to calculate a preset acquisition frequency based on the pre-input line speed and product spacing, ensuring that multimodal information is captured at least once for each product to be inspected. Specifically, the image acquisition frequency f must meet the requirements of the line speed v and the bag spacing d, ensuring that each bag is imaged at least once. For example, if the line speed is 1.5 m / s and the bag spacing is 0.3 m, the acquisition frequency should be at least 5 Hz.
[0051] Specifically, to ensure that various image acquisition devices can simultaneously capture images of the same target, a synchronization trigger mechanism can be implemented to ensure that each device captures images simultaneously when a logistics bag passes through the inspection area. For example, the visible light camera uses a Basler acA2440-75um industrial camera with a resolution of 2448x2048 pixels and a frame rate of 75fps; the infrared camera uses a FLIR A615 thermal imager with a wavelength range of 7.5-14μm and a thermal sensitivity of <50mK. A hardware synchronization signal generator is used as the synchronization trigger to ensure a synchronization error of ≤1ms between the two cameras. The mounting brackets for the visible light camera and the infrared thermal imager are fixed above the assembly line, with a field of view covering the entire width of the assembly line (approximately 1.5m). The cameras and infrared thermal imager are connected via a synchronization trigger to ensure simultaneous capture of images of the logistics bag. A data cable transmits the images to an edge computing device (NVIDIA Jetson AGX Orin). Before image acquisition, the camera and infrared thermal imager were mounted above the assembly line, and the angles were adjusted to ensure field of view coverage. The acquisition frequency was set to 10 fps (assembly line speed 1.2 m / s, spacing 0.5 m, f = 1.2 / 0.5 = 2.4 Hz, with a margin of 10 fps). The visible light camera exposure time was configured to 1 ms to avoid motion blur (blur < 1 pixel). The synchronization error was tested to ensure it was ≤ 1 ms.
[0052] The present application does not limit the specific type of the image synthesis model, which can be selected according to actual conditions. In a specific implementation, the image synthesis model includes an image generation model and an adversarial network model.
[0053] In one specific implementation, the defect data enhancement module 20 includes: a model construction unit for constructing an image generation model and an adversarial network model; a first synthesis unit for inputting a captured normal product image and defect description text into the image generation model to generate a first synthesized defect image. The adversarial network includes a generator and a discriminator. The generator generates a second synthesized defect image based on the normal product image and input random noise; the discriminator distinguishes the second synthesized defect image from a real defect image and optimizes the generator based on the discrimination result.
[0054] In a specific implementation, Figure 4 As shown, this application adopts a hybrid defect data enhancement module, including:
[0055] Defect injection based on a large-scale generative model: This model uses a text-to-image generation model (a Flux.1 architecture fine-tuned for industrial scenarios) to input images of normal logistics bags and text describing defects to generate highly realistic synthetic defect images. Model pruning and quantization techniques are used to control the generation time of a single image to within an acceptable range for industrial applications (e.g., less than 30 seconds). The quality of the generated images is evaluated using metrics such as FID, striving for an FID close to 0.
[0056] Defect generation based on domain-specific GANs: A conditional generative adversarial network (cGAN) is designed, consisting of a generator G and a discriminator D. Generator G takes a normal logistics bag image and random noise as input to generate synthetic defect images. Discriminator D distinguishes between real defect images and synthetic defect images. Poisson fusion is used to optimize the edge transition between defects and background, and random noise is added to simulate sensor and environmental interference to enhance the realism of the synthesized data.
[0057] For example, 1000 normal logistics bag images and 200 real defect images were collected; the normal images and defect descriptions (such as "loose fibers") were input using fine-tuned Flux to generate synthetic defect images (generation time <20s, FID <50); the model was trained to generate 5000 synthetic defect images, adding Gaussian noise (σ=0.01); the data was integrated, and the total training set was 6000 images (1200 real + 4800 synthetic, with synthetic accounting for 80%).
[0058] The edge computing deployment module 30 adopts a cloud management platform and deploys edge node gateways for data management.
[0059] In one specific implementation, the edge computing deployment module 30 includes a model quantization unit, an operator fusion unit, and a pruning unit. These units optimize the defect detection model for the edge computing platform (NVIDIA Jetson Orin) through model quantization (reducing FP32 to INT8), operator fusion, and pruning. This reduces the computational effort by approximately four times, significantly improving inference speed.
[0060] This application does not limit the specific method of feature extraction, which can be selected according to actual conditions. In one specific implementation, the defect detection module 40 includes: a first feature extraction unit, which is used to extract local detail features of multimodal information through a convolutional neural network and output a feature map; a second feature extraction unit, which is used to segment the multimodal information into multiple image blocks through a visual transformer, and capture the long-range dependencies and global context information corresponding to each image block, identify widely distributed or irregular defects, and output a feature vector; and a feature fusion unit, which is used to fuse the feature map and feature vector deep in the network to obtain a fused feature.
[0061] In a specific implementation, Figure 5 As shown, the defect detection module 40 of the present application utilizes a multimodal fusion AI defect detection module. Feature extraction includes a CNN branch and a ViT branch. The CNN branch utilizes a lightweight CNN architecture to extract local detail features (such as tear edges, hole outlines, and fiber texture) and outputs a feature map. The ViT branch utilizes a visual transformer (such as a small version of the SwinTransformer) to segment the image into multiple image blocks, capturing long-range dependencies and global contextual information to identify widespread or irregular defects (such as large adhesions or scattered fibers). This outputs a feature vector. CNN local features and ViT global features are fused deep within the network, employing an attention-weighted fusion strategy to generate a robust multi-scale feature representation. Furthermore, the defect detection module 40 also includes a localization and classification unit, which determines defect classification and location based on the fused features. A detection head is designed based on the fused features to perform defect classification and location (outputting bounding boxes or masks). An adaptive threshold mechanism is introduced to dynamically adjust the confidence threshold based on the defect type. For example, a lower threshold is used to improve recall for subtle fiber defects, while a higher threshold is used to ensure precision for significant damage.
[0062] In a specific implementation, the intelligent rejection module 50 includes: a path planning unit, which is used to plan the operation path of the robot arm based on the abnormal products identified by the defect detection module 40 and the speed of the assembly line; and a confirmation unit, which is used to confirm whether the abnormal products are removed from the assembly line based on the photoelectric switch.
[0063] This application utilizes a high-speed, high-precision robotic arm to plan rejection paths in real time based on abnormality signals (defect location, type, and posture information) output by the detection module. A negative pressure suction device and pressure sensor are installed at the end of the robotic arm. The intelligent rejection module also adjusts the suction pressure of the negative pressure suction device based on the type of abnormal product and defect. After the rejection action is completed, a photoelectric switch confirms whether the abnormal product has been successfully removed from the production line. The results are fed back to the system for statistical analysis and process optimization.
[0064] This solution leverages the synergy of multimodal image acquisition, hybrid data enhancement, multimodal AI fusion detection, edge computing optimization, and an intelligent rejection module to achieve high-precision, real-time detection and precise removal of abnormal logistics bags (such as loose fibers, adhered flocs, tears, holes, and other defects) on the logistics packaging line. Specifically, the logistics product anomaly detection and rejection system provided by this invention has at least the following technical benefits:
[0065] 1) High-Precision Detection: This technology combines the advantages of multimodal image information with the CNN-ViT hybrid model to effectively overcome the limitations of single modality and single model type. It achieves an overall detection accuracy of over 98% for common logistics bag defects, including loose fibers, sticky flocs, tears, and holes, significantly outperforming existing technologies.
[0066] 2) High real-time performance and edge adaptability: Through model optimization and edge computing deployment, the system's single-frame detection time is strictly controlled within 50ms, meeting the real-time requirements of high-speed (e.g., ≥1.5 m / s) logistics and packaging lines.
[0067] 3) Strong robustness: Multimodal fusion reduces dependence on a single light source or environmental conditions; the rich dataset generated by the hybrid data augmentation strategy significantly improves the model's generalization and adaptability to lighting changes, logistics bag morphology diversity, and new defects.
[0068] 4) Intelligent and precise rejection: Dynamic path planning combined with adaptive negative pressure adsorption technology enables rapid (rejection action cycle <0.5 seconds), stable and precise removal of detected abnormal logistics bags, with a false rejection rate of less than 0.5%, while avoiding damage to normal products.
[0069] 5) Solved the problem of scarcity of defect samples: The innovative hybrid data augmentation method can efficiently generate a large amount of high-quality and diverse synthetic defect data, effectively alleviating the pain point of insufficient real defect samples in industrial scenarios and lowering the data threshold for model training.
[0070] 6) High system integration and improved automation: The system tightly integrates high-precision detection with intelligent rejection to form an end-to-end automated solution, significantly reducing manual intervention and improving the efficiency and intelligence level of the entire logistics and packaging process, with significant economic and social benefits.
[0071] In one embodiment, the present application provides a method for detecting and eliminating abnormalities in logistics products, comprising the steps of:
[0072] S100, synchronously acquiring multimodal information of a product to be inspected on an assembly line using at least two image acquisition devices at a preset acquisition frequency;
[0073] S200: Extract and fuse at least two features from the multimodal information, and input the fused features into a defect detection model to identify abnormal products. The defect detection model is configured to be trained using a training dataset constructed by fusing synthetic defect images generated by the at least two image synthesis models with real defect image samples.
[0074] S300: Controlling the robotic arm to remove abnormal products identified by the defect detection module.
[0075] The technical concept of the logistics product anomaly detection and rejection method is consistent with the logistics product anomaly detection and rejection system of the aforementioned embodiment, and will not be described in detail here.
[0076] It should be noted that the above embodiments can be freely combined as needed. The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A logistics product abnormality detection and rejection system, characterized in that: include: An image acquisition module, configured with at least two image acquisition devices, is used to synchronously acquire multimodal information of the products to be inspected on the assembly line at a preset acquisition frequency; A defect data enhancement module is used to generate synthetic defect images based on at least two image synthesis models and fuse real defect image samples to construct a training dataset; an edge computing deployment module, configured to optimize a defect detection model of a defect detection module according to the defect data enhancement module, and input the multimodal information transmitted by the image acquisition module into the defect detection model; a defect detection module, configured to extract and fuse at least two features from the multimodal information, and input the fused features into the defect detection model to identify abnormal products; The intelligent rejection module is used to reject abnormal products identified by the defect detection module by controlling the robotic arm.
2. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The image acquisition device includes a visible light camera, an infrared camera and a depth camera, and the multimodal information includes a visible light image, an infrared thermal imaging image and a depth image; The image acquisition module includes: an image preprocessing unit, connected to the visible light camera, the infrared camera, and the depth camera, respectively, for preprocessing the visible light image, the infrared thermal imaging image, and the depth image; A synchronization control unit is used to synchronize the visible light camera, the infrared camera and the depth camera through a synchronization trigger.
3. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The image acquisition module is further configured to calculate the preset acquisition frequency based on the pre-input assembly line speed and product spacing, so that the multimodal information is acquired at least once for each product to be inspected.
4. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The image synthesis model includes an image generation model and an adversarial network model; The defect data enhancement module includes: A model construction unit, configured to construct the image generation model and the adversarial network model respectively; A first synthesis unit is configured to input the acquired normal product image and defect description text into the image generation model to generate a first synthesized defect image; The adversarial network includes a generator and a discriminator. The generator is used to generate a second synthetic defect image based on the normal product image and input random noise; the discriminator is used to distinguish the second synthetic defect image from the real defect image, and optimize the generator based on the distinction result.
5. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The defect detection module includes: A first feature extraction unit, configured to extract local detail features of the multimodal information through a convolutional neural network and output a feature map; a second feature extraction unit, configured to segment the multimodal information into a plurality of image blocks through a visual transformer, capture long-range dependencies and global context information corresponding to each of the image blocks, identify widely distributed or irregular defects, and output a feature vector; The feature fusion unit is used to fuse the feature map and the feature vector in the deep layer of the network to obtain a fused feature.
6. A logistics product abnormality detection and rejection system according to claim 5, characterized in that: Also includes: The positioning and classification unit is used to determine the defect classification and positioning according to the fusion features.
7. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The edge computing deployment module includes a model quantization unit, an operator fusion unit and a pruning unit, which are used to optimize the defect detection model.
8. A logistics product abnormality detection and rejection system according to claim 1, characterized in that: The intelligent rejection module includes: a path planning unit, configured to plan an operating path of the robotic arm based on the abnormal products identified by the defect detection module and the speed of the assembly line; A confirmation unit is used to confirm whether the abnormal product is removed from the assembly line based on the photoelectric switch.
9. A logistics product abnormality detection and rejection system according to claim 8, characterized in that: A negative pressure adsorption device and a pressure sensor are provided at the end of the robotic arm, and the intelligent rejection module is further used to adjust the adsorption negative pressure of the negative pressure adsorption device according to the type of abnormal product and the type of defect.
10. A method for detecting and eliminating abnormalities in logistics products, characterized in that: Including steps: Synchronously acquiring multimodal information of the product to be inspected on the production line using at least two image acquisition devices at a preset acquisition frequency; Extracting at least two features and fusing the features on the multimodal information, and inputting the fused features into the defect detection model to identify abnormal products, wherein the defect detection model is configured to be trained using a training dataset constructed by fusing synthetic defect images generated by at least two image synthesis models with real defect image samples; Abnormal products identified by the defect detection module are removed by controlling the robotic arm.