Multi-modal ai picking system

The multimodal AI picking system addresses the limitations of conventional systems by integrating text generation AI with image recognition and ensemble learning, offering a flexible, low-cost solution that improves efficiency, accuracy, and human-machine collaboration for diverse products and environments.

JP2025126098APending Publication Date: 2025-08-28中村义一
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024034402
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-16
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Conventional picking systems are costly, complex, lack flexibility, and fail to support effective human-machine collaboration, making them difficult for small to medium-sized businesses to adopt and adapt to diverse products and environments.

Method used

A multimodal AI picking system utilizing text generation AI with image recognition and generation capabilities, combined with ensemble learning, provides a flexible and low-cost solution that integrates multiple AI models for improved accuracy and communication through voice, text, and visual interfaces.

Benefits of technology

The system significantly enhances work efficiency, accuracy, and adaptability, reducing labor costs and errors while facilitating intuitive human-AI collaboration, enabling quick responses to market changes and improving overall productivity and quality assurance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126098000001_ABST
    Figure 2025126098000001_ABST
Patent Text Reader

Abstract

To solve a problem that picking operations in factories and warehouses often rely heavily on manual labor, resulting in high costs, time consumption, and human error issues, and in small businesses in particular, the high cost of implementing automated systems poses a significant burden, improvements in efficiency and accuracy being required.SOLUTION: A multi-modal AI picking system automates the picking process by utilizing image recognition and AI-driven text generation. Through real-time item identification and condition assessment, intuitive instructions are provided to workers via a multimodal interface, enhancing both speed and accuracy of the work. This technology is applicable across diverse industries from manufacturing to logistics and retail, enabling cost-effective operational efficiency improvements, particularly, for small businesses with limited resources.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to automation technology, and in particular to the automation of object identification and sorting processes using artificial intelligence (AI) systems with image recognition and generation capabilities.

[0002] More specifically, it relates to a multimodal AI picking system that aims to improve the efficiency of picking operations (item selection, classification, and transportation) in factory and warehouse environments. The system uses image recognition and generation technology to automatically assess the condition of items and instruct relevant parties on the necessary actions.

[0003] This technical field is the application of computer vision, machine learning, and especially text generation AI using deep learning algorithms, which involves the development of systems that can identify objects from still or moving images, understand their characteristics, and automatically derive appropriate actions. The purpose of this invention is to improve efficiency, accuracy, and cost reduction in industrial automation and logistics.

[0004] The present invention also relates to a communication method that uses a multimodal interface (e.g., voice, text messaging, visual display, etc.) to convey the system's decision results to human operators and other stakeholders, thereby increasing the transparency of AI system decisions and facilitating collaboration between humans and AI.

[0005] Furthermore, the present invention relates to a method for applying ensemble learning techniques to integrate the predictions of multiple AI models to improve the overall accuracy and reliability of the system, thereby reducing the risk of misjudgment by a single AI model and realizing a more robust picking system.

[0006] The present invention is particularly focused on providing a low-cost, highly flexible and adaptable solution for automated picking and quality inspection systems, especially for small to medium-sized manufacturers and logistics companies. [Background technology]

[0007] In recent years, improving the efficiency of logistics processes in factories and warehouses has become an important issue, and there has been particular interest in automating picking work. Picking work refers to the task of gathering specific items in a warehouse based on an order form, and improving the efficiency of this task directly contributes to reducing logistics costs and improving service quality.

[0008] In the past, various technologies, such as barcode scanners, RFID (radio frequency identification) technology, and automated picking robots, have been used to improve the efficiency of picking operations. These technologies automate the identification and tracking of items and can improve picking accuracy. However, these systems are generally expensive and require specialized knowledge for installation and maintenance, making their adoption difficult, especially for small to medium-sized businesses.

[0009] While there are also picking systems that use image recognition technology, these systems are specialized for specific types of items and have the problem of difficulty adapting to new types of items or different work environments. The lack of flexibility of such systems is a major issue in modern manufacturing, where high-mix, low-volume production is the norm.

[0010] Furthermore, while conventional picking systems aim to fully automate picking tasks, in reality there are many situations where a human operator is required, making collaborative work between humans and machines important. However, conventional technologies lack the functionality to support effective interaction between the human operator and the system.

[0011] The present invention was proposed to solve the problems of the conventional technology described above. In particular, there is a need for the development of a low-cost, highly flexible picking system that can be easily introduced by small to medium-sized businesses by utilizing text generation AI with image recognition and generation functions.

[0012] In particular, the main challenges facing traditional picking systems are high cost, complexity of implementation and operation, limited adaptability, and lack of human interaction. These challenges hinder the improvement of efficiency and flexibility of picking operations, especially in manufacturing industries that produce a wide variety of products in small quantities, and in warehouse operations that must respond to rapidly changing logistics needs.

[0013] Additionally, while traditional picking systems use image recognition technology to identify items, they fail to take advantage of the benefits of image generation and multimodal interaction powered by generative AI. Advances in generative AI have made it possible to not only recognize images, but also generate specific instructions when problems are discovered and effectively communicate with relevant human stakeholders. This allows for closer and more efficient collaboration between AI and human operators, improving the overall efficiency and accuracy of the picking process.

[0014] To address these challenges, this invention proposes a low-cost, highly flexible multimodal picking system that utilizes text generation AI with image recognition and generation functions. This system can be easily implemented by small to medium-sized businesses, and by promoting collaboration between humans and AI, it improves the efficiency and quality of picking work. [Prior art documents] [Non-patent literature]

[0015] In developing the present invention, prior art documents in the following technical fields were reviewed:

[0016] [Non-Patent Document 1] Technical literature on automated picking systems: Technologies related to the automatic sorting and transportation of goods in factories and warehouses. In particular, research on systems using robotic arms and automated guided vehicles (AGVs) and control methods for these systems was referenced.

[0017] [Non-patent document 2] Research literature on image recognition technology: Literature on image analysis methods using object recognition, pattern recognition, and machine learning techniques. The latest research on deep learning-based approaches and the development of algorithms for identifying specific objects was reviewed.

[0018] [Non-patent document 3] Technical literature on text generation AI: Research on AI systems that combine natural language processing (NLP) and text generation techniques. The literature was surveyed on how these systems generate useful information from text data and integrate it into user interfaces.

[0019] [Non-patent document 4] Research literature on multimodal interaction: Literature on the design of user interfaces that combine voice, text, and visual feedback. In particular, research focused on interaction techniques to improve human-machine interaction was considered.

[0020] [Non-Patent Document 5] Technical literature on ensemble learning: research into how to combine multiple machine learning models to improve the accuracy and reliability of predictions. Literature on improving the overall performance of a system by integrating different approaches was consulted.

[0021] These documents provide a foundation for understanding how the present invention advances existing technologies and provides new value. In particular, the present invention combines and improves these prior art technologies to realize a new level of automation and efficiency in picking operations in factories and warehouses.

[0022] This invention integrates the achievements of each technical field presented in the prior art literature and proposes an innovative approach to overcome the limitations of these technologies. In particular, the combination of image recognition and text generation AI significantly improves the flexibility and adaptability of the picking system. Furthermore, the use of multimodal interaction improves communication between the system and the operator, enabling more intuitive and effective work.

[0023] The introduction of ensemble learning further improves the accuracy and reliability of system decisions by combining the predictions of multiple AI models, offering significant advantages, especially in the field of industrial automation where minimizing the risk of false positives is crucial.

[0024] Through review of prior art literature, the present invention overcomes the functional limitations of existing picking systems and provides a highly flexible and efficient picking solution that is applicable to a wide range of businesses, from small to large. In this way, the present invention opens new avenues for the automation and optimization of operations in logistics and manufacturing. Summary of the Invention

[0025] This invention relates to a multimodal AI picking system that utilizes text generation AI with image recognition and generation capabilities to automate and streamline picking operations in factories and warehouses. This system automates item identification, condition assessment, and the provision of appropriate operating instructions, effectively supporting collaborative work with human operators.

[0026] The core of this invention is that it combines AI-based image recognition and text generation technologies to identify items and assess their condition in real time, and communicates the results to operators through a multimodal interface (e.g., voice guidance, text messages, visual feedback, etc.), thereby improving the accuracy and speed of picking operations and increasing work efficiency.

[0027] Furthermore, this invention employs ensemble learning to integrate the prediction results of multiple AI models to improve the accuracy and reliability of judgments. This approach reduces the risks of relying on a single AI model and realizes a more robust picking system.

[0028] One of the features of the present invention is its high customizability. The system can be easily adapted to suit the requirements of different types of items, various picking tasks, and specific work environments. This allows businesses of all sizes to implement and use the system to suit their needs.

[0029] The main benefits offered by this invention are reduced labor costs through the automation of work, improved accuracy and speed of work, and improved employee satisfaction through an improved work environment. In addition, the ease of system implementation and maintenance allows businesses to take advantage of the latest automation technology without the need for expensive initial investment or specialized knowledge.

[0030] By optimizing the interaction between humans and AI, this invention not only automates the picking process but also supports the decision-making process of workers. This human-centric approach aims to improve work efficiency as well as worker satisfaction and work quality.

[0031] This invention also facilitates the integration of the latest AI technology with existing logistics and manufacturing processes, providing businesses with the flexibility to quickly adapt to market changes and maintain their competitive edge. In this way, this invention is expected to contribute not only to technological innovation but also to business process innovation. [Problem to be solved by the invention]

[0032] The present invention aims to address several problems, including:

[0033] Reduced labor costs and improved work efficiency: Traditionally, picking work in factories and warehouses has been largely done manually, which has led to high labor costs and inefficiencies. The present invention aims to reduce labor costs and significantly improve work efficiency through the automation of picking work.

[0034] Improved work accuracy: Manual picking work can lead to reduced accuracy due to human error, such as selecting the wrong product or incorrect quantities. This invention uses advanced image recognition technology and AI-based judgment to improve the accuracy of picking work.

[0035] Lowering the barriers to technology adoption for small businesses: Conventional automated picking systems require large capital investments and specialized knowledge, making adoption particularly difficult for small businesses. This invention provides a system that can be easily installed and operated at low cost, allowing businesses of all sizes to benefit from automation technology.

[0036] Enhanced Flexibility and Adaptability: Flexibility and adaptability of a picking system are important to quickly respond to market needs and changes in production lines. The present invention provides a highly customizable system that can easily adapt to various types of products and different work environments.

[0037] Facilitating human-machine collaboration: Effective collaboration between human operators and machine systems is crucial when implementing automation technology. This invention facilitates human-AI interaction through a multimodal interface, enabling more intuitive and efficient work.

[0038] Our approach to these issues is achieved by utilizing text generation AI with image recognition and generation capabilities, and improving judgment accuracy through ensemble learning. This optimizes the entire picking process and provides the following specific solutions:

[0039] Significant improvement in work efficiency through automation: AI's image recognition capabilities enable accurate identification and classification of items, making picking faster and more accurate than manual work.

[0040] Reduced human error: AI-driven decisions are consistent and unaffected by fatigue and distraction, significantly reducing work errors and improving overall work quality.

[0041] Low-cost deployment: The invention is designed to be easily integrated into existing hardware and infrastructure, requiring no special capital investment and making it easy for even small businesses to deploy.

[0042] Customizable and adaptable: The system is easily customizable to accommodate different types of products and changing work environments, and can be tailored to an operator's specific needs.

[0043] Improved collaborative work: Through the multimodal interface, operators can intuitively understand feedback from AI and quickly take appropriate actions, which facilitates smooth interaction between humans and AI and improves overall work efficiency.

[0044] As described above, the present invention provides practical and effective solutions to several challenges facing modern factories and warehouses. By solving these challenges, businesses can increase productivity, improve the work environment, and increase employee satisfaction. Furthermore, the adoption of the present invention enables businesses to respond quickly to market fluctuations, maintaining and improving their competitive position. [Means for solving the problem]

[0045] The solution to the problem according to the present invention is realized by the following means:

[0046] Utilizing text-generation AI with image recognition capabilities: The core of this invention is the use of an AI system that combines advanced image recognition technology with text generation capabilities. This system analyzes image data acquired by cameras and sensors to identify specific items and assess their condition. The identified information is interpreted by the text-generation AI and provided to the operator as work instructions and feedback.

[0047] Implementing a Multimodal Interface: The present invention features a multimodal interface that provides information to operators through multiple communication channels, including voice, text messages, and visual indicators. This approach allows operators to intuitively understand instructions from the AI ​​and quickly take appropriate action.

[0048] Improving decision accuracy through ensemble learning: By adopting an ensemble learning approach that integrates the predictions of multiple AI models, the overall decision accuracy of the system is improved. By cross-validating the information provided by different AI models and highlighting matching results, the risk of false positives is minimized.

[0049] Customizable System Design: The system of the present invention can be easily adapted to suit different types of goods, work environments, and the needs of specific operators. This flexibility allows the system to be adapted to a wide range of applications, allowing operators to operate the system in a configuration that best suits their requirements.

[0050] Easy integration and low cost: The system is designed to be compatible with existing factory and warehouse infrastructure, and can be easily integrated without requiring special capital investment or a complex installation process, reducing installation costs and making it accessible to small businesses.

[0051] Through these measures, the present invention automates picking operations, improving efficiency and accuracy while simultaneously reducing worker burden and improving the work environment. Furthermore, the system's flexibility and customizability enable businesses of various industries and scales to quickly respond to market fluctuations and production line updates. The system also functions as an interface to promote collaboration between humans and AI, enabling a more intuitive and effective picking process.

[0052] Furthermore, the adoption of ensemble learning eliminates the risk of relying on a single AI model and makes it possible to maintain high judgment accuracy even under various conditions and abnormal situations, which contributes not only to improved productivity but also to strengthened quality assurance.

[0053] Ultimately, the solutions provided by this invention are expected to provide practical and innovative solutions to major challenges facing the modern manufacturing and logistics industries, contributing to the advancement of the industry as a whole. These solutions demonstrate the uniqueness and value of this invention in both technical details and practical applications. [Effects of the Invention]

[0054] The main advantages obtained by implementing the present invention are as follows:

[0055] Significant improvement in work efficiency: The introduction of text generation AI with image recognition and generation functions will enable the automation of picking work, which will significantly improve work efficiency and reduce time and costs compared to manual work.

[0056] Improved work accuracy: The use of advanced AI image recognition technology reduces human error and improves the accuracy of picking work, which reduces the number of incorrect products and quantity errors, directly contributing to improved customer satisfaction.

[0057] Increased flexibility and scalability: The system's customizability and adaptability allows it to quickly respond to different product types and changing market needs, allowing businesses to easily add new product lines or change production processes.

[0058] Facilitating human-machine collaboration: Multimodal interfaces allow operators to intuitively understand instructions and feedback from AI systems, enabling them to perform tasks faster and more accurately, facilitating effective collaboration between human operators and AI systems.

[0059] Cost reduction and improved ROI (return on investment): This invention can be introduced at low cost and contributes to reducing operational costs, which allows operators to expect a quick return on investment and improved profits in the long term.

[0060] Enhanced quality assurance: The improved judgment accuracy achieved through ensemble learning also contributes to product quality assurance. Accurate picking and inspection ensures consistent product quality and increases brand reliability.

[0061] As a result of these effects, the present invention will significantly advance automation and efficiency in the manufacturing and logistics industries, while also contributing to cost reductions and improved quality control in business operations. These benefits are important factors for businesses to maintain their competitiveness and increase customer satisfaction. [Brief explanation of the drawings]

[0062] [Figure 1] system configuration diagram DETAILED DESCRIPTION OF THE INVENTION

[0063] As one embodiment of the present invention, a multimodal AI picking system for automating picking tasks in factories and warehouses is proposed. This system includes the following main components and processes:

[0064] Image Acquisition Unit: Using cameras and other sensors, it captures images and data of the objects being picked. This unit provides high-resolution images that allow accurate identification of the object's shape, color, size, and other characteristics.

[0065] Image Recognition and Analysis Module: The acquired image data is analyzed using deep learning algorithms and other image recognition techniques. This module identifies objects and evaluates their condition (e.g., whether they are normal or defective).

[0066] Text generation AI engine: Based on information from the image recognition and analysis module, the text generation AI generates work instructions and feedback. This engine uses natural language processing technology to provide instructions that are easy for workers to understand.

[0067] Multimodal Interface: Instructions are provided to workers in multiple forms, including voice guidance, text messages, LED indicators, and touchscreen displays. The interface can be customized to suit worker preferences and the work environment.

[0068] Ensemble learning system: Combining multiple AI models to improve judgment accuracy. Predictions from different approaches are integrated, and the result with the highest confidence is used as the basis for picking decisions.

[0069] User Interface and Control Panel: The interface through which the system operator adjusts the system settings and monitors and manages operations. This panel is used to change the settings of operation parameters, monitor the system status, and troubleshoot abnormalities.

[0070] In this embodiment, the system offers high flexibility and adaptability for specific picking tasks. For example, if a new product line needs to be added or the picking process needs to be changed, the system can be easily reconfigured to accommodate new requirements with minimal downtime. The image acquisition unit and image recognition and analysis module quickly learn the characteristics of new objects, and the text generation AI engine instantly generates new work instructions. This allows operators to flexibly respond to changing market trends and production requirements and maintain production efficiency.

[0071] Multimodal interfaces also allow workers to choose the optimal communication method depending on their environment and preferences, further improving worker understanding and work efficiency. For example, in noisy environments, visual indicators and displays can be used, while in quieter environments, voice guidance can be used primarily.

[0072] The ensemble learning system achieves high judgment accuracy beyond the limits of a single AI model. This contributes to minimizing false positives, particularly in the picking of defective products and quality inspection processes. The system combines the results of multiple AI models to make a final decision, resulting in a more reliable work process.

[0073] The user interface and control panel allow operators to easily monitor the overall performance of the system and make immediate adjustments as needed, allowing for faster system maintenance and problem resolution, minimizing production line downtime.

[0074] As described above, the embodiment of the present invention combines advanced AI technology with a user-friendly interface to achieve efficient and highly accurate picking operations while simultaneously reducing the workload of operators, which is expected to improve productivity, reduce costs, and even improve the working environment in the manufacturing and logistics industries. [Example]

[0075] Below, we will show an example of the application of a multimodal AI picking system integrated into a factory conveyor belt system as an embodiment of the present invention.

[0076] Example 1: Automated defect detection and classification

[0077] System configuration: Image acquisition unit: A high-resolution camera that captures images of the products as they move along the conveyor belt. Image recognition and analysis module: An AI algorithm that identifies product features from captured images and detects the presence or absence of defects. Text generation AI engine: Generates text instructions on the type of defect detected and how to address it. Multimodal interface: Provides instructions to workers via voice, text messages, and visual feedback.

[0078] Operation process: Products moving along the conveyor belt are continuously photographed by an image capture unit. The captured images are sent to an image recognition and analysis module in real time, and AI detects product defects. If a defect is detected, a text generation AI engine generates countermeasures according to the type of defect and provides instructions to the worker through a multimodal interface. Workers will follow the instructions provided to properly dispose of defective items (e.g., repair, replace, dispose, etc.).

[0079] effect: The product quality control process is automated, enabling the rapid detection and appropriate disposal of defective products. This reduces the number of defective products that are mistakenly processed due to human error, improving overall work accuracy. This reduces the burden on workers and significantly improves the efficiency of the production line. This example demonstrates that the multimodal AI picking system of the present invention significantly improves the efficiency and accuracy of product quality control and defective product processing. Furthermore, the system is easily adaptable to different product lines and work environments, and is expected to be applied in a wide range of industrial fields.

[0080] Example 2: Automation of picking tasks for a wide variety of products in a warehouse

[0081] System configuration: Image acquisition unit: A high-resolution camera and barcode scanner for identifying products stored on shelves in the warehouse. Image recognition and analysis module: An AI algorithm that identifies the type and location of a product from product images taken with a camera and barcode information. Text generation AI engine: Generates picking instructions for the required items based on the picking list. Multimodal interface: Provides picking instructions to workers through voice guidance, text messages, and visual feedback.

[0082] Operation process: Once the picking list is entered into the system, an image capture unit identifies the location of the target items within the warehouse. The captured image and barcode information are sent to an image recognition and analysis module, where AI accurately identifies and locates the product. A text generation AI engine generates picking instructions for the identified items and provides instructions to workers through a multimodal interface. The worker follows the instructions provided to pick the items and prepare them for order fulfillment.

[0083] effect: The process of picking goods within the warehouse will be automated, greatly improving work efficiency. Increased accuracy in identifying and locating products, reducing picking errors and increasing customer satisfaction. The multimodal interface reduces the burden on workers and speeds up picking operations.

[0084] This example demonstrates that the multimodal AI picking system of the present invention improves the efficiency and accuracy of automated picking of a wide variety of products in a warehouse. Furthermore, this system enables rapid product picking and reduces order processing time, thereby contributing to improved customer service.

[0085] Example 3: Real-time order customization and picking

[0086] System configuration: Image Acquisition Unit: A high-resolution camera for capturing images of products and components for customized orders. Image recognition and analysis module: An AI algorithm that identifies the characteristics of products and parts from captured images and evaluates whether they meet customization requirements. Text generation AI engine: Generates specific picking and assembly instructions based on customized orders. Multimodal interface: Provides instructions to workers through voice guidance, text messages, and visual feedback.

[0087] Operation process: Once customized order information is entered into the system, an image capture unit takes images of the required products or parts. The captured images are sent to an image recognition and analysis module, where AI identifies the characteristics of the ordered products and parts and assesses their suitability. A text generation AI engine generates picking and assembly instructions based on identified items and parts, and provides instructions to workers through a multimodal interface. Workers follow the instructions provided to pick the items and perform customized assembly as needed.

[0088] effect: Customization for each order is done in real time, allowing for quick response to individual customer requests. Accurate product and part identification and compatibility assessment using AI will improve the quality of customized products. The multimodal interface enables workers to execute customized picking and assembly processes quickly and accurately. This example demonstrates that the multimodal AI picking system of the present invention improves the efficiency and accuracy of quickly processing customized orders based on individual customer needs, thereby increasing customer satisfaction while ensuring the production efficiency and quality of customized products. [Industrial Applicability]

[0089] The multimodal AI picking system of this invention has high automation capabilities and flexibility, making it applicable to a wide range of industrial fields. In particular, it is expected to be used in the following fields:

[0090] Manufacturing: This technology can be used to automate quality control and picking operations in factories that manufacture a variety of products, including automobiles, electronic devices, food, and pharmaceuticals. The invention contributes to detecting product defects, accurately sorting parts, and streamlining the manufacturing process.

[0091] Logistics and warehousing: This technology can be applied to automating and streamlining operations at logistics centers and warehouses, such as warehousing and shipping, inventory management, and order-based picking. This technology enables accurate and speedy product handling, reduces delivery delays, and improves customer satisfaction.

[0092] Retail: It can be used for order processing and product management in online shopping and in-store sales. In particular, it can improve service and reduce operating costs by quickly processing individual e-commerce orders and automating inventory management in stores.

[0093] Medical and healthcare: This technology can be used for precision picking tasks in the medical field, such as managing medical equipment and medicines, assembling treatment kits for each patient, etc. This technology can prevent mix-ups of medical supplies and ensure patient safety.

[0094] Custom product manufacturing: In the manufacture of customized products tailored to individual customer needs, it can be used to streamline the picking and assembly of parts according to orders, increasing production flexibility and enabling companies to respond quickly to diverse customer requests.

[0095] As described above, the present invention is expected to be adopted in many industrial fields due to its wide applicability and the economic benefits of improving business processes. Furthermore, the introduction of this system will promote the automation of work and the optimization of human resources, contributing to the improvement of the competitiveness and sustainability of the entire industry.

[0096] Agriculture: This technology can be used to automate quality-based picking and sorting in post-harvest processing and sorting of agricultural produce. The technology can identify the size, color, and shape of produce such as fruits and vegetables and sort them according to quality standards.

[0097] Food processing industry: This technology can also be applied to automating food processing, such as picking ingredients on a food packaging line or adding ingredients to a specific product. In the food industry, where hygiene control and accurate ingredient addition are required, this technology improves the accuracy and efficiency of work.

[0098] Entertainment and Events: Automate the picking and organization of exhibits and materials at events and exhibitions, enabling quick setup and teardown, and improving event management efficiency.

[0099] The multimodal AI picking system of the present invention not only automates tasks, but also improves work quality, reduces worker burden, and ultimately contributes to increased customer satisfaction. As such, this invention is expected to be put to practical use in a wide range of industrial fields as an important technology that promotes industrial digital transformation and efficiency.

[0100] Due to the industrial applicability described above, the present invention can provide effective solutions to modern challenges faced by many companies and organizations, and contribute to the creation of new value and the promotion of industrial development.

Claims

1. A system that automates the picking of specified items using text generation AI with image recognition and generation functions, comprising: an image recognition and analysis module that acquires images of the items and identifies and evaluates the condition of the items based on the images; a text generation AI engine that generates work instructions based on the results of the identification and condition evaluation; and a multimodal interface that communicates the work instructions to workers.

2. 2. The picking system according to claim 1, wherein the image recognition and analysis module uses a layer learning algorithm to identify features of the article from the image.

3. 3. The picking system of claim 1 or claim 2, wherein the multimodal interface provides voice guidance, text messages, and / or visual feedback.

4. The picking system according to claim 1, further comprising an ensemble learning system that integrates prediction results of a plurality of AI models.

5. 10. The picking system of any preceding claim, wherein the system is applied to manufacturing, logistics / warehousing, retail, medical / healthcare, or custom product manufacturing.

6. The picking system according to claim 1 , wherein the image acquisition unit includes a 3D scanner or an infrared camera.

7. 2. The picking system according to claim 1, wherein the image recognition and analysis module measures dimensions or estimates weight of the items in addition to identifying the items.

8. 2. The picking system according to claim 1, wherein the text generation AI engine takes into account the skill level and work history of the worker when generating work instructions.

9. 10. The picking system of any preceding claim, wherein the multi-modal interface receives feedback from an operator and adaptively adjusts system operation based on the feedback.

10. 10. A picking system according to any preceding claim, wherein the system directs further processing of picked items. Such additional processing may include, but is not limited to, packaging, labeling, quality inspection, or customized fabrication.