Intelligent tool cabinet system based on multi-mode large language model and control method

The intelligent tool cabinet system, which uses a multimodal large language model and prompt word engineering, solves the problems of high cost and low recognition accuracy, and realizes low-cost and efficient unmanned management. It is suitable for scenarios such as smart bookcases and smart retail.

CN120823664AActive Publication Date: 2025-10-21JILIN UNIVERSITY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511320782.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-21
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing smart tool cabinet systems have high initial investment costs and limited recognition accuracy. Traditional methods also have problems such as fragile RFID tags, electromagnetic interference, and recognition failures caused by similar visual features.

Method used

An intelligent tool cabinet system based on a multimodal large language model is adopted. Combining the multimodal large language model and prompt word engineering, it recognizes books through image analysis, reduces dependence on RFID tags and massive image data, and integrates the edge computing platform with the multimodal recognition module to achieve automated management.

Benefits of technology

It reduces initial costs, improves recognition accuracy and system stability, enables unmanned management, reduces manpower expenditure, and supports cross-scenario migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823664A_ABST
    Figure CN120823664A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and intelligent systems, and particularly relates to an intelligent tool cabinet system based on a multi-modal large language model and a control method. Comprising the steps that a core processing module obtains a state image of the intelligent tool cabinet and inputs the state image to a multi-mode recognition module; the cue word engineering module inputs cue words to the multi-modal recognition module, wherein the cue words comprise role information, output format information and spatial distribution guide words; the multi-modal identification module is integrated with a multi-modal big language model, and the multi-modal big language model realizes identification of tool names and borrowing and returning states based on state images and prompt words of the intelligent tool cabinet; and the data processing module receives an identification result of the multi-mode identification module, and completes extraction and storage of tool access and / or return records. According to the method, the problems of high initial input cost and limited accuracy in the prior art are solved on the basis of the multi-modal large language model and the cue word engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and intelligent systems, and in particular relates to an intelligent tool cabinet system and a control method based on a multimodal large language model. Background Art

[0002] Smart tool cabinets have a variety of uses, such as smart bookcases and smart snack cabinets. For example, with the rapid development of artificial intelligence technology, people are increasingly demanding smarter, more personalized reading spaces. Traditional bookcases simply store books. Consequently, traditional libraries rely heavily on a human workforce for book management and loan services, exposing drawbacks such as high labor costs and inefficient management. The emergence of smart bookcases has disrupted this traditional model, driving book lending towards intelligent, unmanned operations and becoming a key breakthrough in upgrading library services.

[0003] Existing technological approaches to smart bookcases can be categorized primarily into two categories: RFID (Radio Frequency Identification) and computer vision algorithms. RFID-based smart bookcases have numerous limitations. First, each book requires an additional RFID tag, resulting in high initial investment costs. Second, RFID tags are susceptible to recognition failures or missed reads due to physical damage or the proximity of adjacent book tags. Furthermore, communication between tags and readers is highly susceptible to environmental electromagnetic interference, severely impacting system stability and reliability. Smart bookcases based on currently mainstream computer vision algorithms require the acquisition of image datasets and model training for massive volumes of books, which is extremely costly. Furthermore, books in smart bookcases are densely packed, with similar spine shapes and a lack of distinctive patterns or visual features, further complicating the application of computer vision methods in book recognition. Summary of the Invention

[0004] In view of this, the present invention aims to provide an intelligent tool cabinet system and control method based on a multimodal large language model to solve the problems of high initial investment cost and limited accuracy in the existing technology. The intelligent tool cabinet designed based on the multimodal large language model of the present invention effectively avoids the above problems and provides a new direction for the development of intelligent tool cabinets.

[0005] To achieve the above object, the technical solution created by the present invention is implemented as follows: A smart tool cabinet system based on a multimodal large language model comprises: a core processing module, a prompt word engineering module, a multimodal recognition module, and a data processing module connected to the core processing module. The core processing module acquires a status image of the smart tool cabinet and inputs the status image to the multimodal recognition module; the prompt word engineering module inputs prompt words to the multimodal recognition module, wherein the prompt words include role information, output format information, and spatial distribution guide words; the multimodal recognition module integrates the multimodal large language model, and the multimodal large language model recognizes tool names and borrowing and returning status based on the status image and prompt words of the smart tool cabinet; and the data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool access and / or return records.

[0006] Furthermore, the core processing module is based on the edge computing platform and is connected to the identity authentication module, image acquisition module, lock control module, Hall sensor module, status sensing module and user interaction module.

[0007] Furthermore, the identity authentication module completes user identity authentication through face recognition, RFID card recognition or fingerprint recognition.

[0008] Furthermore, the prompt word engineering module inputs prompt words into the multimodal large language model through a preset prompt strategy. The preset prompt strategy includes: Based on the role-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as role information. The multimodal large language model sets the tool administrator role based on the role information to match the tool management scenario; Based on the template-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as format information, and the multimodal large language model outputs it in a fixed format based on the format information; Based on the task decomposition strategy, the prompt word engineering module inputs spatial distribution guide words into the multimodal large language model. The spatial distribution guide words guide the multimodal large language model to compare and analyze the changes in the spatial distribution of tools before and after the cabinet door of the smart tool cabinet is opened and closed.

[0009] Furthermore, the multimodal large language model is GPT-4o.

[0010] Furthermore, the state image of the smart tool cabinet obtained by the core processing module includes a state image of the smart tool cabinet before the user operation and a state image of the smart tool cabinet after the user operation.

[0011] A control method for an intelligent tool cabinet based on a multimodal large language model is implemented using an intelligent tool cabinet system based on a multimodal large language model, and specifically includes the following steps: S1: The core processing module obtains the status image of the intelligent tool cabinet and inputs the status image into the multimodal recognition module; S2: The prompt word engineering module inputs prompt words to the multimodal recognition module. The prompt words include role information, output format information, and spatial distribution guide words; S3: The multimodal recognition module integrates a multimodal large language model. The multimodal large language model recognizes tool names and borrowing and returning status based on the status image and prompt words of the intelligent tool cabinet. S4: The data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool use and / or return records.

[0012] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) The present invention creates an intelligent tool cabinet system and control method based on a multimodal large language model. The multimodal large language model is used to analyze the image differences of bookcases before and after borrowing and returning. No RFID tags or massive image training data are required. Only the model's image and text understanding capabilities are used to identify books, thus saving the cost of purchasing and maintaining RFID tags. At the same time, there is no need to collect image data sets for thousands of books, reducing the cost of data annotation and model training. In addition, the present invention can also avoid the problem of missed readings caused by physical damage, electromagnetic interference of RFID tags and similar spine features of visual algorithms.

[0013] (2) The present invention creates the intelligent tool cabinet system and control method based on the multimodal large language model, which improves the recognition efficiency and standardization based on the multi-strategy interaction mechanism optimized by the prompt word engineering. The present invention improves the output format consistency from 80% before optimization to 100% through template design, solving the problem that the traditional method needs manual correction due to format confusion and takes a long time. The present invention improves data processing efficiency by directly reducing manual intervention.

[0014] (3) The present invention creates an intelligent tool cabinet system and control method based on a multimodal large language model, integrating modules such as facial recognition, magnetic locks, and Hall sensors on the NVIDIA Jetson Orin NX platform, and automatically completes the "identity authentication, image acquisition, model reasoning, and record storage" process without manual supervision. The present invention can replace the repetitive work of 2-3 administrators in traditional libraries, reducing manpower expenditures, and supports the borrowing and returning of multiple books in a single operation. The Hall sensor and magnetic lock are used to control the cabinet door in conjunction to avoid recognition errors caused by the user not closing the door. In other words, the edge computing-driven full-process automated hardware system realizes unmanned management and efficient borrowing and returning.

[0015] (4) The present invention creates a smart tool cabinet system and control method based on a multimodal large language model. This system utilizes hardware interfaces (USB, GPIO) in conjunction with a software prompt word module to support flexible adjustments. By replacing prompt words (e.g., from "book identification" to "product identification") and adapting domain data, it can be migrated to scenarios such as smart retail. The present invention employs a modular architecture that is portable across scenarios, addressing the challenges of technology reuse and scenario adaptation. The present invention eliminates the need for redeveloping hardware or training models, significantly reducing the time and cost of deployment compared to traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings: Figure 1 A schematic diagram of the structure of an intelligent tool cabinet system based on a multimodal large language model according to an embodiment of the present invention; Figure 2 A schematic diagram of the workflow of the intelligent tool cabinet based on the multimodal large language model according to an embodiment of the present invention; Figure 3 The present invention is a flowchart of a method for controlling an intelligent tool cabinet based on a multimodal large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation of the present invention.

[0018] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0019] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second" and the like are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second" and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0020] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0021] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0022] like Figure 1 As shown, the present invention proposes an intelligent tool cabinet system based on a multimodal large language model, including: a core processing module, a prompt word engineering module, a multimodal recognition module and a data processing module, wherein the core processing module obtains a status image of the intelligent tool cabinet and inputs the status image to the multimodal recognition module; the prompt word engineering module inputs prompt words to the multimodal recognition module, and the prompt words include role information, output format information and spatial distribution guide words; the multimodal recognition module integrates a multimodal large language model, and the multimodal large language model realizes the recognition of tool names and borrowing and returning status based on the status image and prompt words of the intelligent tool cabinet; the data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool access and / or return records.

[0023] It should be noted that the smart tool cabinet proposed in this invention does not require 24 / 7 human supervision. Leveraging a self-service system, it can efficiently complete functions such as borrowing and returning items. It also leverages the powerful understanding and recognition capabilities of a multimodal large language model and optimizes prompt word engineering to accurately identify books, significantly reducing library labor costs and freeing staff from repetitive, low-value-added work. Furthermore, through targeted adjustments to modules such as prompt words, it can be flexibly applied to scenarios such as smart retail and smart bookcases.

[0024] Specifically, this invention introduces a multimodal large language model to construct an intelligent tool cabinet system that requires no additional physical tags, is immune to environmental interference, and does not require massive image training. This system achieves the following technological breakthroughs: 1) Reduced initial cost: No need to attach RFID tags to books or pre-collect large amounts of image data, thus avoiding hardware tagging and data training costs; 2) Improved recognition stability: Book recognition is achieved through text information interaction, avoiding issues that affect recognition, such as physical damage to RFID tags, electromagnetic interference, and similar visual features, thereby improving recognition accuracy and system reliability; 3) Simplified technical approach: Book management is accomplished through natural language interaction and semantic understanding, facilitating subsequent information extraction. This overcomes the technical bottleneck of traditional visual algorithms that rely on complex image feature extraction, reducing the difficulty of system deployment and maintenance.

[0025] In some embodiments, the core processing module, with the edge computing platform as the core, is connected to an identity authentication module, an image acquisition module, a lock control module, a Hall sensor module, a status sensing module and a user interaction module.

[0026] It should be noted that the multimodal interactive interface of the present invention supports data access such as image acquisition (USB camera), user touch (HDMI touch screen), and identity authentication (face recognition); the core processing module is based on the NVIDIA Jetson OrinNX edge computing platform, integrating a multimodal large language model and prompt word engineering optimization components, and can realize tool status recognition and borrowing and returning record generation through a standardized API interface; the hardware expansion capability supports external magnetic locks, Hall sensors and other peripheral modules through the GPIO interface, so as to meet the hardware configuration requirements of different scenarios through modular design.

[0027] Furthermore, the hardware system architecture of the present invention is based on the core hardware platform NVIDIA Jetson Orin NX, and combines the identity authentication module, image acquisition module, lock control module, status sensing module and user interaction module to realize corresponding functions, wherein the identity authentication module can capture the user's facial image through the camera and compare it with the database to realize identity authentication; the image acquisition module has a built-in high-definition camera for capturing the status of tools in the tool cabinet; the lock control module receives system instructions to control the cabinet door switch, and uses an electromagnetic lock to realize contactless unlocking; the Hall sensor module triggers the image acquisition and lock control logic through distance sensing; the user interaction module uses the touch screen for operation guidance (such as clicking the "borrow and return" button) and result display.

[0028] This invention builds a fully automated system, encompassing user authentication, door control, image acquisition, model inference, and data storage, to achieve a "human-free + real-time response" intelligent borrowing and returning experience. To enhance understanding of the entire process, the following detailed borrowing and returning process is illustrated, using a smart bookcase as an example: A user initiates an operation by clicking the "Borrow and Return" button via a user interaction module (e.g., a touchscreen). The user interaction module transmits this signal to the system core, triggering subsequent processes and simultaneously prompting the user to authenticate. Upon receiving this signal, the authentication module (e.g., employing facial recognition) captures the user's facial image using a camera and compares it to a database. If authentication fails, the user interaction module reports a result (e.g., "Identity not matched"), terminating the process and waiting for the user to retry. If authentication succeeds, the image acquisition module automatically captures the initial state of the door before it is opened, designated as Image1. The authentication module then sends an unlock command to the lock control module, while the user interaction module simultaneously prompts, "Door unlocked, please proceed." After the user completes the operation and closes the cabinet door, they confirm the end of the operation by clicking the "Process Completed" button in the user interaction module. The Hall sensor module detects the change in the cabinet door's state from "open" to "closed" and immediately sends a trigger signal to the image acquisition module (e.g., a high-definition camera). In response, the image acquisition module automatically captures the final state of the cabinet door after closing, recording it as Image2, and temporarily stores the state images (including Image1 and Image2) in the system. The image acquisition module uploads the state images to a multimodal large language model guided by prompt words to analyze the differences in the operation (e.g., adding or removing items). The analysis results (including the names of the borrowed and returned books) are displayed in the user interaction module for user confirmation. After the user confirms, the data (including timestamp, user name, and operation details) is stored in a database (e.g., a MySQL database can be used). The user interaction module displays "Operation Completed," and the system resets to its initial state.

[0029] In some embodiments, the identity authentication module completes user identity authentication through face recognition, RFID card recognition, or fingerprint recognition.

[0030] It should be noted that if fingerprint recognition is adopted, a fingerprint sensor needs to be integrated to confirm the user's identity through fingerprint comparison.

[0031] In some embodiments, the prompt word engineering module inputs prompt words to the multimodal large language model using a preset prompt strategy. The preset prompt strategy includes: Based on the role-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as role information. The multimodal large language model sets the tool administrator role based on the role information to match the tool management scenario; Based on the template-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as format information, and the multimodal large language model outputs it in a fixed format based on the format information; Based on the task decomposition strategy, the prompt word engineering module inputs spatial distribution guide words into the multimodal large language model. The spatial distribution guide words guide the multimodal large language model to compare and analyze the changes in the spatial distribution of tools before and after the cabinet door of the smart tool cabinet is opened and closed.

[0032] It should be noted that the present invention adopts a multimodal large language model and prompt word engineering in terms of tool identification. GPT-4o is an advanced multimodal large language model developed by OpenAI, which has the ability to process text and image information at the same time. The present invention uses its powerful visual recognition function to analyze the input image content, and at the same time understands and processes related text instructions and information through text interaction functions. Prompt word engineering refers to the careful design and optimization of prompt words input into the multimodal large language model to guide the model to better understand the task requirements and generate expected outputs. After receiving the image and prompt words, the GPT-4o model uses its multimodal processing capabilities to conduct an in-depth analysis of the image, combines the prompt words to understand the task requirements, and accurately identify the book title in the image. The prompt word engineering introduced in the present invention specifically adopts three prompt word acquisition methods: role-based, template-based design, and task decomposition-based.

[0033] Furthermore, using a smart tool cabinet as a smart bookcase as an example, a role-based prompt strategy assigns a specific role to the multimodal large language model, making it more targeted when handling tasks. For example, "You are a librarian." A template-based prompt strategy uses a fixed template to guide the multimodal large language model to output in a specific format. For example, "The borrowed books are: <Recognition Result 1>, <Recognition Result 2>; the returned books are: <Recognition Result 1>, <Recognition Result 2>." A task-decomposition-based prompt strategy guides the model to improve its image analysis. For example, "Observe the distribution of books in the bookcase and compare the changes before and after the door is opened." By comparing basic requirements with prompts based on role setting, template design, and task decomposition strategies, the effectiveness of different strategies in guiding the multimodal large language model to handle the "Smart Tool Cabinet (simulated smart bookcase) borrowed and returned book recognition" task was evaluated. The impact of each strategy on output accuracy, structure, and scenario adaptability was clarified. Experiments found that the multimodal large language model and prompt engineering can accurately obtain book borrowing and returning information. At the same time, this information is output according to the required template, which facilitates the subsequent code to extract the information.

[0034] In some embodiments, the multimodal large language model is GPT-4o.

[0035] In some embodiments, the state image of the smart tool cabinet obtained by the core processing module includes a state image of the smart tool cabinet before the user operation and a state image of the smart tool cabinet after the user operation.

[0036] The present invention also provides a control method for an intelligent tool cabinet based on a multimodal large language model, which is implemented using an intelligent tool cabinet system based on a multimodal large language model and specifically includes the following steps: S1: The core processing module obtains the status image of the intelligent tool cabinet and inputs the status image into the multimodal recognition module; S2: The prompt word engineering module inputs prompt words to the multimodal recognition module. The prompt words include role information, output format information, and spatial distribution guide words; S3: The multimodal recognition module integrates a multimodal large language model. The multimodal large language model recognizes tool names and borrowing and returning status based on the status image and prompt words of the intelligent tool cabinet. S4: The data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool use and / or return records.

[0037] It should be noted that the present invention adopts a label-free book recognition technology based on a multimodal large language model. By integrating the multimodal large language model and combining the image data before and after borrowing and returning, the model's image and text understanding ability is used to achieve accurate recognition of the tool name and status. The present invention also proposes a multi-strategy interaction mechanism for optimizing the prompt word engineering. By innovatively introducing the prompt word engineering, through strategies such as template design and task decomposition, the multimodal large language model is guided to quickly understand the scene information and generate borrowing and returning records that meet the needs, thereby improving recognition accuracy and reasoning efficiency and reducing dependence on manually labeled data. In addition, the hardware interface and software model used in the present invention support modular adjustment. By replacing the prompt word template and adapting the domain data (such as converting from book images to product packaging images), it can be quickly migrated to scenarios such as smart retail, logistics warehousing, etc., realizing the functional reuse of "smart bookcase-vending machine-inventory management system" and reducing cross-domain deployment costs.

[0038] Finally, it's important to note that the multimodal large language model keeps pace with the times and is not limited to the GPT-4o model. The edge computing platform uses domestically produced edge computing equipment, further reducing costs. Furthermore, for book recognition, a book knowledge base (containing structured data such as title, author, and spine features) is constructed. A knowledge base reference is included in the prompt (e.g., "Based on knowledge base ID: BN001, identify the book title corresponding to this spine") to guide the multimodal large language model to analyze images using prior knowledge.

[0039] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. This is not limited herein.

[0040] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. An intelligent tool cabinet system based on a multimodal large language model, characterized by: include: A core processing module, and a prompt word engineering module, a multimodal recognition module and a data processing module connected to the core processing module, wherein the core processing module obtains a status image of the intelligent tool cabinet and inputs the status image to the multimodal recognition module; the prompt word engineering module inputs prompt words to the multimodal recognition module, the prompt words including role information, output format information and spatial distribution guide words; the multimodal recognition module integrates a multimodal large language model, and the multimodal large language model realizes the recognition of tool names and borrowing and returning status based on the status image and prompt words of the intelligent tool cabinet; the data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool access and / or return records.

2. The intelligent tool cabinet system based on a multimodal large language model according to claim 1, characterized in that: The core processing module is based on the edge computing platform and is connected to the identity authentication module, image acquisition module, lock control module, Hall sensor module, status sensing module and user interaction module.

3. The intelligent tool cabinet system based on a multimodal large language model according to claim 2, characterized in that: The identity verification module completes user identity verification through face recognition, RFID card recognition or fingerprint recognition.

4. The intelligent tool cabinet system based on a multimodal large language model according to claim 1, characterized in that: The prompt word engineering module inputs prompt words into the multimodal large language model using preset prompt strategies. The preset prompt strategies include: Based on the role-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as role information. The multimodal large language model sets the tool administrator role based on the role information to match the tool management scenario; Based on the template-based prompt strategy, the prompt word engineering module inputs the prompt word into the multimodal large language model as format information, and the multimodal large language model outputs it in a fixed format based on the format information; Based on the task decomposition strategy, the prompt word engineering module inputs spatial distribution guide words into the multimodal large language model. The spatial distribution guide words guide the multimodal large language model to compare and analyze the changes in the spatial distribution of tools before and after the cabinet door of the smart tool cabinet is opened and closed.

5. The intelligent tool cabinet system based on a multimodal large language model according to claim 1 is characterized in that: The multimodal large language model is GPT-4o.

6. The intelligent tool cabinet system based on a multimodal large language model according to claim 1, characterized in that: The core processing module obtains a state image of the smart tool cabinet, including a state image of the smart tool cabinet before the user operation and a state image of the smart tool cabinet after the user operation.

7. A method for controlling an intelligent tool cabinet based on a multimodal large language model, implemented using the intelligent tool cabinet system based on a multimodal large language model according to any one of claims 1 to 6, characterized in that: The specific steps include: S1: The core processing module obtains the status image of the intelligent tool cabinet and inputs the status image into the multimodal recognition module; S2: The prompt word engineering module inputs prompt words to the multimodal recognition module. The prompt words include role information, output format information, and spatial distribution guide words; S3: The multimodal recognition module integrates a multimodal large language model. The multimodal large language model recognizes tool names and borrowing and returning status based on the status image and prompt words of the intelligent tool cabinet. S4: The data processing module receives the recognition results of the multimodal recognition module and completes the extraction and storage of tool use and / or return records.

Citation Information

Patent Citations

  • Intelligent management system and intelligent management cabinet thereof

    CN104778574A

  • Intelligent safety tool cabinet based on facial recognition technology and memory function and control method

    CN114120541A

  • Intelligent cabinet interaction method and device based on large language model and electronic equipment

    CN118468880A

  • Tool cabinet self-service borrowing and returning method and intelligent tool storage system

    CN119445720A

  • Fire-fighting facility detection system and method based on Internet of Things and artificial intelligence

    CN119578992A