Method for processing visual data for object detection
By using a detection model to generate visual tokens with CLIP and ConvNext-CLIP models, the method addresses object detection challenges in automated warehouses, enhancing accuracy and efficiency while managing anomalies and improving stock control.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2026-03-12
AI Technical Summary
Existing object detection methods in automated warehouses face challenges such as variations in object appearance, sensitivity to noise and blur, complex backgrounds, scalability issues, computational intensity, domain adaptation problems, and the presence of anomalies like packaging debris, leading to inaccurate and inefficient object detection.
The method involves using a detection model to identify objects, creating cropped images, reducing image resolution, and processing these with a CLIP model and/or ConvNext-CLIP model to generate visual tokens, followed by combining these with text tokens to create response or action tokens for improved object detection and handling.
This approach enhances object detection accuracy and versatility, reduces computational requirements, and enables real-time processing, effectively handling anomalies and improving stock control in automated warehouses.
Smart Images

Figure EP2025074794_12032026_PF_FP_ABST
Abstract
Description
[0001] Method for processing visual data for object detection
[0002] The present invention relates to a method for processing visual data for object detection using Al techniques. The method provides an improved detection of objects in a processed image and may particularly be used in the context of robotic grippers and picking totes in an automated warehouse. The present invention further relates to a computer program product and a computer system.
[0003] Background
[0004] In automated warehouses, robotic picking grippers are essential components that help in handling and moving objects. These grippers are designed to grasp and manipulate various items, ensuring efficient and accurate picking in warehouse operations.
[0005] An integral part of the robotic picking process is the detection of objects in the totes. Computer vision techniques, often powered by machine learning algorithms, are used to identify and locate items within the tote. This allows the gripper to adjust its movements and grasp the object securely, minimizing damage and increasing productivity.
[0006] Typically, object manipulation involves several consecutive steps. The first step is the image acquisition. A camera or other imaging device captures an image of the tote and its contents. The captured image may then be preprocessed in an intermediate step. To improve the accuracy of the object detection process, the image is enhanced and normalized. As the next step follows the object detection. Using machine learning algorithms, the system identifies and locates objects within the image. For instructing the robotic gripper to pick a certain object, a grasping point calculation is required. Based on the detected object's location and size, the gripper calculates the optimal grasping point. The final step is then the motion planning and manipulation action. The gripper plans its movements to reach the grasping point and pick up the object and executes the action.
[0007] By combining advanced robotic grippers with accurate object detection techniques, automated warehouses can significantly improve their efficiency and reduce errors in the picking process. However, object detection, while beneficial in many applications, faces several challenges and problems that can impact its performance. One of them are variations in object appearance. Objects can appear different due to changes in lighting, angles, or occlusions, making it difficult for the system to recognize them consistently. Object detection models can be sensitive to noise, blur, or other image artifacts, leading to false positives or negatives. Complex backgrounds with
[0008] 1 . September 2025 1 / 20 similar colors or textures to the objects can confuse object detection algorithms, leading to incorrect detections. Another problem is the scalability. Object detection models may struggle to scale and maintain accuracy when dealing with a large number of object categories or when the number of objects in the scene increases. Particularly when a real-time performance is required, processing images and detecting objects can be computationally intensive, requiring powerful hardware and optimized algorithms. Moreover, object detection systems can be fooled by adversarial examples, where subtle changes to the input image cause the model to misclassify objects. Objects can also be partially or fully occluded by other objects, making it difficult for the system to detect and locate them accurately. Finally, the domain adaptation of the object detection models is an important problem. Models trained on one dataset may not generalize well to new datasets or environments, requiring additional fine-tuning or retraining.
[0009] These problems apply to object detection in general and are not limited to automated warehouses but are common to all robotic picking processes which do not rely on preprogrammed picking positions and predefined objects in always the same orientation.
[0010] Hence, the common object detection processes in general still need to be improved to better address these problems. Regarding automated warehouse applications, this concerns particularly more accuracy and versatility. Today's object detection methods are often limited to objects or context to which they are specifically trained for, and their retraining or additional training can be quite time consuming. With a view to possible frequent changes in stocked items, this is of course a great disadvantage.
[0011] A specific additional problem with the operation of automated warehouses is the fact that, every now and then, the picking totes will not only contain the desired items but also garbage, such as parts of the packaging or adhesive tapes, items erroneously picked from stock, or damaged items with torn or scuffed packaging or deformations. These "anomalies" pose an additional demand on the object detection.
[0012] Further, intralogistics suffers from consistently poor process data that can make it harder to control stock and fulfill customer orders on time. This can be due to conflicting or inaccurate inventory data or the anomalies described above. Only an effective monitoring and detection can make it possible to create a precise warehouse stock database and design proper optimization processes. Hence, also for this reason, an improved object detection is desirable.
[0013] Summary of this disclosure
[0014] In an aspect, the present disclosure relates to a method for processing visual data for object
[0015] 1 . September 2025 2 / 20 detection, in particular of objects in a warehouse, comprising the steps of: a) providing an image by means of an image capturing device and / or on a data storage; b) processing the image of step a) with a detection model for detecting objects and creating cropped images of each detected object; and c) reducing the resolution of the image of step a) and processing the created image with reduced resolution and the cropped images of step b) with a CLIP model and / or a ConvNext- CLIP model for creating visual tokens.
[0016] In another aspect, the present disclosure relates to a computer program product comprising computer program code which, when executed on a computing device having one or more processors, causes the one or more processors to perform a method according to the aforementioned aspect.
[0017] In another aspect, the present disclosure relates to a computer system comprising a processor and a data storage device operatively coupled to the processor, the data storage device containing portions of program code which, when executed by the processor, configure the processor to perform a method according to the aforementioned aspect.
[0018] Embodiments, features and effects disclosed hereinafter in the context of a method for processing visual data for object detection apply respectively to the computer program product and / or the computer system. In embodiments, the computer system can comprise some of all of the physical components, configurations and functions that are described in the context of the method. For example, the computer system can comprise an image capturing device, such as a camera, for providing an image.
[0019] Short of the Fi
[0020] Figure 1 shows a schematic view of method steps according to the invention.
[0021] Details of this disclosure
[0022] The details of this disclosure relate to the aspect described in the summary of this disclosure.
[0023] The method will be described with reference to the scheme shown in Figure 1 which also contains further optional steps.
[0024] The present method for processing visual data for detecting objects is especially useful for detecting items in an automated warehouse, for example, in a picking tote. There, the generated
[0025] 1 . September 2025 3 / 20 tokens may advantageously be used for improved item handling and stock control. However, its application is not limited to this particular field of application. The improved object detection will likewise be beneficial for any process involving the evaluation of an image as input for the creation of a (textual) response or action.
[0026] The first step a) is the provision of an image 1 by means of an image capturing device and / or on a data storage. The visual data are not necessarily limited to images taken in the visual range but may also comprise images taken in the IR or UV range if the field of application requires it. Also, combinations of such spectral ranges are possible. In the warehousing context, a combination of IR and visual images can, for example, be used for the detection of scratched or damaged surfaces of transparent or glossy items. The images may either be still images or frames from a video clip. For example, cameras, scanners, or video cameras can be used as image capturing devices. As an alternative to this "live" input, the respective images can also be provided as files on a data storage medium. The latter may be useful in cases where an additional preprocessing of the acquired images is performed to enhance the image quality or to normalize the image.
[0027] In Figure 1 , as an example, the provision of an image 1 by a camera is indicated on the lower left-hand side of the scheme. The exemplary image 1 contains three different objects symbolized by the cross, bolt of lightning, and heart icons. In the second step b), this image 1 is processed with a detection model 2 for detecting the objects and creating cropped images 3 of each detected object. The detection model 2 should be able to return a bounding box for the detected objects which almost all common models are capable of. The cropped images 3 of the detected objects are essentially images cropped to the content of the bounding boxes, where possible, containing only the single object in full. If there are two or more objects overlapping each other, the resulting cropped images 3 will differ according to the bounding boxes. In Figure 1 , this is indicated by the three small images containing only the cross, bolt of lightning, and heart icon, respectively.
[0028] In the final step c), the resolution of the image 1 is reduced and then the reduced image is processed together with the cropped images 3 of the objects with a CLIP model 4 and / or a ConvNext-CLIP model 4 for creating visual tokens 5.
[0029] Object detection models are designed to identify and locate objects within images or video frames. These models vary in architecture, speed, and accuracy, making them suitable for different applications in object detection across various domains. Most modern object detection models are designed to return bounding boxes in image coordinates.
[0030] 1 . September 2025 4 / 20 In embodiments, the detection model 2 in step b) may be selected from the group comprising YOLO, Faster R-CNN, SSD, RetinaNet, Mask R-CNN, EfficientDet, CenterNet, DETR, Cascade R-CNN, YOLOv5, PP-YOLO, and Swin Transformer. While other models may also be used, these have been found to perform particularly well in the present method.
[0031] YOLO (You Only Look Once) is a real-time object detection system that predicts bounding boxes and class probabilities directly from full images in a single evaluation. It outputs bounding boxes directly in image coordinates as part of its predictions.
[0032] Faster R-CNN is an extension of the R-CNN family that uses a Region Proposal Network (RPN) to generate region proposals, which are then classified and refined. It returns bounding boxes for detected objects in image coordinates after processing through the Region Proposal Network (RPN) and the detection head.
[0033] SSD (Single Shot MultiBox Detector) is a model that detects objects in images using a single deep neural network, predicting bounding boxes and class scores for multiple object categories. It provides bounding box predictions in image coordinates as part of its single-shot detection process.
[0034] RetinaNet introduces the Focal Loss function to address the class imbalance in object detection tasks, allowing it to focus more on hard-to-detect objects. It outputs bounding boxes in image coordinates along with class scores for detected objects.
[0035] Mask R-CNN is an extension of Faster R-CNN that adds a branch for predicting segmentation masks on each Region of Interest (Rol), enabling instance segmentation. In addition to segmentation masks, it also returns bounding boxes in image coordinates for each detected instance.
[0036] EfficientDet is a family of object detection models that balance accuracy and efficiency, using a compound scaling method to optimize the model size and performance. It returns bounding boxes in image coordinates as part of its object detection output.
[0037] CenterNet is a keypoint-based object detection model that predicts the center points of objects and their sizes, allowing for efficient detection. It predicts the center points of objects and their sizes, which can be converted to bounding boxes in image coordinates.
[0038] DETR (Detection Transformer) is a novel approach that uses transformers for object detection, treating the detection task as a direct set prediction problem. It outputs bounding boxes in image coordinates as part of its set prediction framework.
[0039] Cascade R-CNN is an extension of Faster R-CNN that uses a multi-stage object detection
[0040] 1 . September 2025 5 / 20 framework to improve detection performance at different Intersection over Union (loU) thresholds. It returns bounding boxes in image coordinates through its multi-stage detection process.
[0041] YOLOv5 is a popular and efficient version of the YOLO model, known for its ease of use and strong performance in real-time object detection tasks. Similar to its predecessors, it outputs bounding boxes in image coordinates.
[0042] PP-YOLO is an improved version of YOLO that incorporates various enhancements to boost performance while maintaining real-time capabilities. It also provides bounding box predictions in image coordinates.
[0043] Swin Transformer is a transformer-based model that has been adapted for object detection tasks, leveraging hierarchical feature maps for improved accuracy. When adapted for object detection tasks, it outputs bounding boxes in image coordinates.
[0044] The method of this disclosure achieves its improved performance by focusing the computing power onto the objects in the provided image. In a first step, the objects are detected and corresponding object images with the original resolution are generated. Then the full image is reduced in its resolution before being processed together with the high-resolution object images in a CLIP model 4 and / or a ConvNext-CLIP model 4 for creating the visual tokens 5. This way, the models work in full detail on the interesting objects and only with reduced accuracy on the full image. As a result, the produced visual tokens 5 will be producable either with the same quality as before but in shorter time or in the same time but with higher quality.
[0045] A CLIP model (Contrastive Language-Image Pre-Training) is an Al model that combines text and image information into a single model to handle a variety of image and language processing tasks. CLIP is a powerful and versatile model that combines the capabilities of image and speech processing in a single system, enabling a wide range of applications. It has a multimodal learning architecture wherein it is trained to understand and associate image-text pairs. It uses contrastive learning methods to maximize the correspondence between images and text descriptions. The model can be trained with large amounts of image-text pairs from the Internet. Through contrastive training, it learns which text descriptions belong to which images and which do not. CLIP can handle a variety of tasks without special adaptation or additional training. For example, it can be used to classify images, generate captions, recognize objects in images or find text descriptions for images.
[0046] Two of its important advantages are that it can classify images into different categories based on text descriptions without the need for specific labels for the categories and zero-shot learning, which means that it can solve new tasks without having been specifically trained for these
[0047] 1 . September 2025 6 / 20 tasks. It uses the knowledge acquired during training to find generalized solutions. The architecture of CLIP consists of two parts: an image encoder and a text encoder. The image encoder processes images and converts them into vectors, while the text encoder does the same for text descriptions. The two vectors are then compared to determine how well the image and text match.
[0048] CLIP models can be categorized based on various criteria, including their architecture, training methods, and specific applications. A first general type of CLIP models which are suitable for the present disclosure is the Standard CUP model, which is the original model developed by OpenAI. It uses a transformer architecture to learn joint representations of images and text. Other types are Vision-Language Models, which are variants that focus on improving the interaction between visual and textual modalities, often incorporating additional techniques like attention mechanisms. Multimodal CLIP models are models that extend CLIP to handle more than two modalities, such as audio or video, in addition to images and text. Domain-Specific CLIP models are adaptations of the CLIP model trained on specific datasets tailored for particular domains, such as medical imaging, fashion, or wildlife. Zero-Shot CLIP are models that leverage the zero-shot capabilities of CLIP to perform tasks without additional training on specific datasets, relying on the learned representations. Fine-Tuned CLIP are variants of the original model that have been fine-tuned on specific tasks or datasets to improve performance in targeted applications. CLIP with Enhanced Architectures are models that incorporate advanced architectures, such as Vision Transformers (ViTs) or convolutional neural networks (CNNs), to improve feature extraction. CLIP with Contrastive Learning Enhancements are variants that utilize improved contrastive learning techniques to enhance the model's ability to differentiate between similar images and texts. CLIP for Image Generation are models that combine CLIP with generative adversarial networks (GANs) or diffusion models to create images from textual descriptions. CLIP for Retrieval Tasks are specialized models optimized for tasks like image retrieval or text retrieval, focusing on improving the efficiency and accuracy of searching through large datasets. These types of CLIP models highlight the flexibility and adaptability of the original architecture to various tasks and domains, allowing developers to tailor the model to their specific needs.
[0049] In embodiments of the method for processing visual data, the CLIP model 4 in step c) may be selected from the group comprising CLIP, CLIP-ViT, CLIP-ResNet, OpenCLIP, CLIP+GAN, CLIPDraw, SigLIP, DALL-E, and CLIP-Adapter.
[0050] These models may be chosen depending on the specific use case and the models used for further processing of their output. They refer to the following specific models. CLIP (Contrastive Language-Image Pretraining) is the original model developed by OpenAI. CLIP-ViT is a version
[0051] 1 . September 2025 7 / 20 of CLIP that uses Vision Transformers (ViTs) for image processing. CLIP-ResNet is a variant that employs ResNet architectures for image feature extraction. OpenCUP is an open-source implementation of CLIP that allows for further experimentation and adaptation. CLIP+GAN are models that combine CLIP with Generative Adversarial Networks for text-to-image generation tasks. CLIPDraw is a model that uses CLIP to generate images based on textual prompts, often used for artistic purposes. SigLIP (Sigmoid loss for Language-Image Pre-training) is a contrastive image-text zero-shot image classification model which operates solely on image-text pairs and does not require a global view of the pairwise similarities for normalization. DALL-E is primarily a text-to-image generation model, but it incorporates CLIP-like mechanisms for understanding and generating images from text. CLIP-Adapter is a model that integrates adapters into the CLIP architecture for improved performance on specific tasks. All of these models represent a mix of direct adaptations of the CLIP model and related models that utilize similar principles for multimodal learning.
[0052] A ConvNext-CLIP model is a combination of the two powerful architectures ConvNext and CLIP in the field of image and text processing. ConvNext is a Convolutional Neural Network (CNN) architecture that can be viewed as an evolution of traditional CNNs. ConvNext is designed to take advantage of modern vision transformer architectures while maintaining the efficiency and scalability of CNNs. ConvNext improves the structure and performance of traditional CNNs through various architectural optimizations. The CLIP model has already been described above. The ConvNext-CLIP model combines the strengths of ConvNext and CLIP to create a powerful, multimodal Al system. The combined model can provide an enhanced image processing. ConvNext brings in the benefits of advanced CNN techniques, resulting in better image visualization and processing. This includes higher accuracy and efficiency in the extraction of image features. The combined model further provides multimodal learning. CLIP enables the integration of text and image information by processing and linking these modalities simultaneously. This enables tasks such as image description, image search based on text and text generation based on images. By combining ConvNext and CLIP, the model can handle a wide range of tasks, including image classification, caption generation, object recognition and more. Using the ConvNext architecture can improve CLIP'S image processing performance by extracting more detailed and accurate image features. This leads to better overall performance of the combined model in multimodal tasks.
[0053] Like with the CLIP models discussed above, in principle almost any ConvNeXT-CLIP model can be used in the present disclosure. They should be chosen with a view to the specific task to be achieved with the object detection and the models used for further processing their output. A first common general type of ConvNeXT-CLIP models that can be used is the Base Model or
[0054] 1. September 2025 8 / 20 Standard Model which is the standard ConvNeXT-CLIP model that serves as a foundational architecture for multimodal tasks. The Large Model is a larger version of the ConvNeXT-CLIP model, typically with more parameters and layers, designed for improved performance on complex tasks. Fine-tuned Models comprise variants of ConvNeXT-CLIP that have been fine-tuned on specific datasets or tasks, such as image classification, object detection, or specific domain applications. Multimodal Variants are models specifically designed to handle different types of input data, such as audio-visual or text-image pairs, leveraging the ConvNeXT architecture for both modalities. Task-Specific Models include configurations of ConvNeXT-CLIP tailored for particular applications, such as zero-shot learning, retrieval tasks, or generative tasks. Distilled Models are smaller, more efficient versions of the ConvNeXT-CLIP model that have been distilled from larger models to maintain performance while reducing computational requirements. Finally, Hybrid Models are models that integrate ConvNeXT-CLIP with other architectures or techniques, such as transformers or recurrent neural networks, to enhance performance on specific tasks.
[0055] In embodiments of the method for processing visual data, the ConvNext-CLIP model 4 in step c) may be selected from the group comprising ConvNeXT -CLIP Standard, ConvNeXT-CLIP XL, DINOv2, ConvNeXT-CLIP Fine-tuned ImageNet, ConvNeXT-CLIP Fine-tuned COCO, Con- vNeXT-CLIP Zero-Shot, ConvNeXT-DistillCLIP, and ConvNeXT-CLIP RNN Enhanced. These models may be chosen depending on the specific use case and the models used for further processing of their output. They refer to the following specific models. ConvNeXT-CLIP Standard and ConvNeXT-CLIP XL are the above mentioned standard and large versions of the basic model architecture. DINOv2 (Self-Distillation with No Labels) is an advanced self-supervised learning model for vision tasks, building upon the original DINO framework. Developed by researchers at Facebook Al Research (FAIR), DINOv2 aims to improve the representation learning of visual data without the need for labeled datasets. This is particularly useful for tasks where labeled data is scarce or expensive to obtain. ConvNeXT-CLIP Fine-tuned ImageNet is a variant of the ConvNeXT model that has been fine-tuned on the ImageNet dataset using the CLIP objective. This means that the model has been trained to predict the text labels associated with images in the ImageNet dataset, in addition to the standard image classification task. ConvNeXT-CLIP Fine-tuned COCO is a variant of the ConvNeXT -CLIP model that has been fine-tuned on the COCO (Common Objects in Context) dataset using the CLIP objective. ConvNeXT-CLIP Zero-Shot is a variant of the ConvNeXT-CLIP model that has not been finetuned on any specific dataset but trained on a large, general-purpose dataset of images and text pairs using the CLIP objective. This training process allows the model to learn a general representation of images and text, without being specialized to a particular dataset or task. ConvNeXT-DistillCLIP is a variant of the ConvNeXT model that has been trained using a
[0056] 1 . September 2025 9 / 20 knowledge distillation approach, where the teacher model is a CLIP model and the student model the ConvNeXT model. The distillation process allows the model to inherit the knowledge and patterns learned by the larger CLIP model, while reducing the computational requirements. ConvNeXT-CLIP RNN Enhanced is a variant of the ConvNeXT-CLIP model that incorporates a Recurrent Neural Network (RNN) component to enhance its performance.
[0057] In embodiments, the method of this disclosure may comprise the further steps: d) receiving an input prompt 6; e) processing the input prompt 6 of step d) and / or a verbalization of coordinates of the cropped images 3 of step b) with a word embedding layer 7 for creating text tokens 8; and f) combining the visual tokens 5 of step c) with the text tokens 8 of step e) and processing them with a large language model 9 for creating response text tokens 10 and / or action tokens 11 as a response to the input prompt.
[0058] In Figure 1, the additional step d) of receiving an input prompt 6 is indicated on the lower righthand side of the scheme with the input window icon and the symbolic input prompt "abode". This prompt may be a human input or an input from another program and may comprise questions and instructions in relation to the provided image 1. In additional step e), this prompt and / or a verbalization of the coordinates of the cropped images 3 of step b) are processed with a word embedding layer 7 for creating the text tokens 8.
[0059] Word embedding layers are crucial components in natural language processing (NLP) models, as they convert words into dense vector representations. The word embedding layer 7 used within this disclosure is not particularly limited to a certain type. Among others, it may be chosen from any of the following exemplary list.
[0060] Word2Vec is a model that uses either Continuous Bag of Words (CBOW) or Skip-Gram techniques to create word embeddings based on context.
[0061] GloVe (Global Vectors for Word Representation) generates embeddings by leveraging global word co-occurrence statistics from a corpus.
[0062] FastText is an extension of Word2Vec which represents words as bags of character n-grams, allowing it to generate embeddings for out-of-vocabulary words.
[0063] ELMo (Embeddings from Language Models) generates context-sensitive embeddings by using deep bidirectional language models.
[0064] 1. September 2025 10 / 20 BERT (Bidirectional Encoder Representations from Transformers) provides contextual embeddings by using transformer architecture, allowing for dynamic word representations based on surrounding context.
[0065] GPT (Generative Pre-trained Transformer) can generate contextual embeddings but is designed primarily for text generation tasks.
[0066] Sentence-BERT (SBERT) is an adaptation of BERT that produces sentence embeddings, allowing for better performance in tasks like semantic similarity and clustering.
[0067] Universal Sentence Encoder (USE) provides embeddings for sentences and phrases, focusing on capturing semantic meaning.
[0068] Transformers-based Embeddings comprise various transformer models (like RoBERTa, Distil- BERT, etc.) which provide embeddings that can be used in downstream NLP tasks.
[0069] T5 (Text-to-Text Transfer Transformer) treats all NLP tasks as text-to-text problems and generates embeddings accordingly.
[0070] All of these embedding layers are widely used in various NLP applications, from sentiment analysis to machine translation, and they help capture the semantic relationships between words and phrases. Their selection should be made dependent on the required capabilities for the prompts and the desired output for further processing.
[0071] Finally, in additional step f), the visual tokens 5 of step c) are combined with the text tokens 8 of step e) which can be done by simple concatenation. Then, they are processed with a large language model 9 for creating response text tokens 10 and / or action tokens 11 as a response to the input prompt.
[0072] A Large Language Model (LLM) is typically a neural network model that is trained on vast amounts of text data to generate human-like language.
[0073] In embodiments, the large language model 9 in step f) may be selected from the group comprising Mixture-of-Experts Large Language Model, BERT, GPT-4, BART, T5, CLIP, DALL-E, and TinyBERT.
[0074] Mixture-of-Experts (MoE) Large Language Models are a type of neural network architecture that dynamically selects a subset of experts (sub-networks) to process input data, allowing for efficient scaling and improved performance.
[0075] BERT (Bidirectional Encoder Representations from Transformers) is designed for
[0076] 1 . September 2025 11 / 20 understanding the context of words in search queries.
[0077] GPT-4 (Generative Pre-trained Transformer) is the fourth version of the Generative Pre-trained Transformer model developed by OpenAI which focuses on text generation and understanding.
[0078] BART (Bidirectional and Auto-Regressive Transformers) is typically used for text generation and summarization.
[0079] T5 (Text-to-Text Transfer Transformer) treats all NLP tasks as text-to-text tasks and is also an encoder-decoder model that converts all tasks into a text-to-text format.
[0080] CUP (Contrastive Language-Image Pretraining) can connect images and text for various multimodal tasks.
[0081] DALL-E generates images from textual descriptions.
[0082] TinyBERT is a smaller, faster, and lighter distilled version of BERT aimed at efficiency that retains much of its performance.
[0083] In particularly preferred embodiments, the large language model 9 in step f) may be a Mixture- of-Experts Large Language Model selected from the group comprising Switch Transformer, GShard, MoE Transformer, Turing-NLG, DeepSpeed MoE, BERT with MoE, MoeBERT, Flamingo, Gato, and GLaM. These models leverage the Mixture-of-Experts architecture to achieve high performance on language tasks while managing computational resources effectively. In particular, they are extremely effective when used for creating text and action tokens in the context of automated warehousing.
[0084] Switch Transformer is a model using a Mixture-of-Experts approach where only a subset of the available experts is activated for each input, significantly increasing the model's capacity while maintaining efficiency.
[0085] GShard is a framework that enables the training of large-scale models using a Mixture-of- Experts. It allows for the efficient distribution of model parameters across multiple devices.
[0086] MoE Transformer is a variant of the transformer architecture that incorporates a Mixture-of- Experts, allowing for dynamic routing of inputs to different expert networks based on the input characteristics.
[0087] Turing-NLG is a model that incorporates a Mixture-of-Experts to enhance its language generation capabilities, allowing it to scale effectively.
[0088] 1. September 2025 12 / 20 DeepSpeed MoE is an extension of the DeepSpeed library, which provides support for training large Mixture-of-Experts models efficiently, enabling the use of a Mixture-of-Experts in various transformer architectures.
[0089] BERT with MoE are adaptations of BERT incorporating a Mixture-of-Experts to enhance their performance on specific tasks while maintaining efficiency.
[0090] MoeBERT is a specific implementation of BERT that integrates a Mixture-of-Experts to improve its capacity and performance on NLP tasks.
[0091] Flamingo is a multimodal large language model designed for few-shot learning across various modalities, including text and images. It uses a combination of vision and language models to understand and generate visual content, and can perform tasks such as image captioning, visual question answering, and text-to-image generation.
[0092] Gato is a multimodal large language model and a generalist agent that can perform a variety of tasks, including natural language processing and computer vision, across different modalities. It is a highly versatile model that can adapt to different tasks and domains.
[0093] GLaM (Generalist Language Model) is a Mixture-of-Experts model that uses a sparse activation mechanism to improve efficiency and performance on various language tasks.
[0094] The response text tokens 10 and action tokens 11 created by the large language model 9 as a response to the input prompt can be used for various purposes depending on the context of the application of the disclosed method.
[0095] In preferred embodiments, the image 1 of step a) is an image 1 of objects in a radius of action of a robotic gripper and the action tokens 11 of step f) instruct the robotic gripper to grip one or more of the detected objects.
[0096] In general, as the name already suggests, the action tokens 11 can be used for achieving desired actions. These actions may cover a broad range of applications. They can range from purely software related input into other programs for program control or data collection to real world physical actions by activating control systems and actuators. In the latter case, action tokens 11 can be utilized in the context of robotics and control systems to facilitate movement actions, such as directing a robot to perform specific tasks. Action tokens 11 can represent specific movement actions or commands that a robot can execute. For example, tokens might encode actions like "move forward", "turn left", "pick up an object", or "navigate to a location". In robotic manipulation applications, action tokens can guide robotic arms to perform tasks like
[0097] 1 . September 2025 13 / 20 assembling parts, picking and placing objects, or performing delicate operations.
[0098] Hence, when using the method of this disclosure to detect objects in a radius of action of a robotic gripper by processing an image 1 depicting this range, the robotic gripper can, for example, be instructed to pick a bolt and then the corresponding nut. By training models to understand the relationship between action tokens and the corresponding control commands, robots can interpret these tokens to execute physical movements. This mapping can be learned through supervised learning, reinforcement learning, or imitation learning. In scenarios where a robot needs to perform a series of actions, action tokens can help represent the sequence of movements. The model can generate a sequence of action tokens that dictate the robot's behavior over time, allowing for coordinated and complex movements.
[0099] In embodiments of the method, the image 1 of step a) may be an image 1 of a picking tote of an automated warehouse.
[0100] This represents the application of the method in the logistics domain, in particular in the intralogistics domain. Automated warehouses are highly mechanized and computerized facilities that use advanced technology to manage and process inventory, packaging, and shipping. The automation systems, including conveyor belts, robots, and automated storage and retrieval systems (AS / RS), are controlled by a warehouse management system (WMS). The WMS, inventory management software, and other control systems are implemented to manage and monitor the warehouse operations. Goods are received at the warehouse and often scanned, weighed, and measured using automated systems. The inventory is then stored in designated locations, often using AS / RS. When an order is received, the WMS directs robots or automated picking systems to retrieve the required items from storage. The items are then packed and labeled using automated systems. Robots and robotic gripper arms are used for tasks such as picking, packing, and palletizing. If packing of the items is done by humans, the WMS will instruct the AS / RS to pick the items belonging to an order from the racks and place them into a tub or box, also commonly known as the picking tote, for delivery to the human packer.
[0101] A very advantageous implementation of the method of this disclosure involves monitoring cameras taking images of the picking totes which are analyzed for detection of the objects within the picking totes. This may take place already before handover to the human to allow for correction of possible errors. For example, a robotic gripper can be instructed to grip the detected erroneous objects and restock them correctly or to discard detected foreign objects such as waste from outer packaging. For quality control purposes, the method can also be used for detection of non-pristine objects in the picking totes showing scuffing, tears, or deformation and instructing a robotic gripper to remove and replace them.
[0102] 1. September 2025 14 / 20 In embodiments of the method, the visual tokens 5 of step c) and / or the response text tokens 10 and / or action tokens 11 of step f) may be used for monitoring and / or correcting a warehouse stock database of the depicted objects.
[0103] The WMS in an automated warehouse comprises a warehouse stock database for keeping track of the stored items. There, often nominal-actual discrepancies occur for various reasons.
[0104] Items may have been misplaced in the racks or damaged or the system may have delivered the wrong item for dispatch. Even if the database may, at least to a certain extent, be corrected if these errors are noticed by the humans at the place of the picking totes, this will always be retroactive, i.e., the real-time database entries will suffer from incorrect data. The method the of the present disclosure can successfully be used for monitoring and correcting the warehouse stock database. To this end, images of a camera monitoring the picking totes can be analyzed for the contained items and particularly the presence of wrong or damaged items. Using one or more of the created tokens, the database can be corrected, a robotic gripper can be instructed to remove the respective items, and / or a textual or audio message can be produced to inform a human at the picking tote or monitoring the system. As this can be done "on-the-fly", the realtime database is corrected with a minimum of delay and always presents accurate data about the stock.
[0105] 1. September 2025 15 / 20 Reference numerals
[0106] 1 image
[0107] 2 detection model
[0108] 3 cropped image
[0109] 4 CLIP model I ConvNext-CLIP model
[0110] 5 visual tokens
[0111] 6 input prompt
[0112] 7 word embedding layer
[0113] 8 text tokens
[0114] 9 large language model
[0115] 10 response text tokens
[0116] 11 action tokens
[0117] 1. September 2025 16 / 20
Claims
Claims1. A method for processing visual data for object detection, in particular of objects in a warehouse, comprising the steps of: a) providing an image (1) by means of an image capturing device and / or on a data storage; b) processing the image (1) of step a) with a detection model (2) for detecting objects and creating cropped images (3) of each detected object; and c) reducing the resolution of the image (1) of step a) and processing the created image with reduced resolution and the cropped images (3) of step b) with a CLIP model (4) and / or a ConvNext-CLIP model (4) for creating visual tokens (5).
2. The method for processing visual data according to claim 1 , comprising the further steps: d) receiving an input prompt (6); e) processing the input prompt (6) of step d) and / or a verbalization of coordinates of the cropped images (3) of step b) with a word embedding layer (7) for creating text tokens (8); and f) combining the visual tokens (5) of step c) with the text tokens (8) of step e) and processing them with a large language model (9) for creating response text tokens (10) and / or action tokens (11) as a response to the input prompt.
3. The method for processing visual data according to claim 2, wherein the image (1) of step a) is an image (1) of objects in a radius of action of a robotic gripper and the action tokens (11) of step f) instruct the robotic gripper to grip one or more of the detected objects.
4. The method for processing visual data according to claim 2 or 3, wherein the image (1) of step a) is an image (1) of a picking tote of an automated warehouse.
5. The method of processing visual data according to one of the preceding claims, wherein the visual tokens (5) of step c) and / or the response text tokens (10) and / or action tokens (11) of step f) are used for monitoring and / or correcting a warehouse stock database of the depicted objects.
6. The method for processing visual data according to one of the preceding claims, wherein the detection model (2) in step b) is selected from the group comprising YOLO, Faster R-1. September 2025 17 / 20CNN, SSD, RetinaNet, Mask R-CNN, EfficientDet, CenterNet, DETR, Cascade R-CNN, YOLOv5, PP-YOLO, and Swin Transformer.
7. The method for processing visual data according to one of the preceding claims, wherein the CLIP model (4) in step c) is selected from the group comprising CLIP, CLIP-ViT, CLIP- ResNet, OpenCLIP, CLIP+GAN, CLIPDraw, SigLIP, DALL-E, and CLIP-Adapter.
8. The method for processing visual data according to one of the preceding claims, wherein the ConvNext-CLIP model (4) in step c) is selected from the group comprising ConvNeXT- CLIP Standard, ConvNeXT-CLIP XL, DINOv2, ConvNeXT-CLIP Fine-tuned ImageNet, Con- vNeXT-CLIP Fine-tuned COCO, ConvNeXT-CLIP Zero-Shot, ConvNeXT-DistillCLIP, and ConvNeXT-CLIP RNN Enhanced.
9. The method for processing visual data according to one of preceding claims 2 to 8, wherein the large language model (9) in step f) is selected from the group comprising Mixture-of- Experts Large Language Model, BERT, GPT-4, BART, T5, CLIP, DALL-E, and TinyBERT.
10. The method for processing visual data according to one of preceding claim 2 to 9, wherein the large language model (9) in step f) is a Mixture-of-Experts Large Language Model selected from the group comprising Switch Transformer, GShard, MoE Transformer, Turing- NLG, DeepSpeed MoE, BERT with MoE, MoeBERT, Flamingo, Gato, and GLaM.
11. A computer program product comprising computer program code which, when executed on a computing device having one or more processors, causes the one or more processors to perform the method according to one of the preceding claims.
12. A computer system comprising one or more processors and at least one data storage device operatively coupled to the processor, the at least one data storage device containing portions of program code which, when executed by the one or more processors, configure the one or more processors to perform a method according to one of claims 1 to 10.
1. September 2025 18 / 20