A back kitchen supervision intelligent agent based on visual analysis

Through a layered design consisting of an edge perception layer, a cloud-edge collaborative transmission layer, and a cloud-based intelligent agent core layer, the existing kitchen supervision system addresses issues such as high latency, high cost, poor scalability, and lack of closed-loop systems. This enables real-time, multi-regional, and fully automated kitchen supervision, improving supervision efficiency and risk control capabilities.

CN122493388APending Publication Date: 2026-07-31SHENZHEN TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TECH UNIV
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing kitchen monitoring systems suffer from high latency, high cost, poor scalability, lack of functional closed loops, and inability to achieve full-coverage real-time management, making it difficult to meet the needs of large-scale, routine, and real-time monitoring.

Method used

It adopts a layered and decoupled design consisting of an edge perception layer, a cloud-edge collaborative transmission layer, and a cloud-based intelligent agent core layer. The edge end identifies violations in real time, while the cloud performs intelligent decision-making and full-process automated management. Combined with a lightweight model and an efficient communication protocol, it achieves second-level early warning and full-process closed-loop management.

Benefits of technology

It enables real-time alerts at the edge and second-level early warnings for intelligent decision-making in the cloud, supports centralized operation and maintenance in multiple regions, forms a fully automated closed loop, improves supervision efficiency and risk control capabilities, and reduces deployment costs and operation and maintenance difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493388A_ABST
    Figure CN122493388A_ABST
Patent Text Reader

Abstract

This invention discloses a visual analysis-based intelligent agent for kitchen supervision, belonging to the field of intelligent supervision technology. It includes an edge perception layer, a cloud-edge collaborative transmission layer, a cloud-based intelligent agent core layer, and an interactive application layer. The edge perception layer enables local real-time inference of kitchen violations, immediate alarms, and structured log generation. The cloud-edge collaborative transmission layer enables low-latency bidirectional transmission, remote equipment control, and breakpoint resume. The cloud-based intelligent agent core layer completes violation assessment, automated report generation, and rectification tracking. The interactive application layer provides visual monitoring, mobile alarm push notifications, and natural language interaction. This invention achieves real-time edge perception and cloud-based intelligent decision-making collaboration, constructing a fully automated regulatory closed loop encompassing video acquisition, violation identification, alarm intervention, regulatory matching, analysis and judgment, report generation, and rectification tracking. This significantly improves kitchen supervision efficiency and risk control capabilities, and is suitable for intelligent supervision of multi-regional, large-scale canteen kitchens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent supervision technology, and more specifically to an intelligent control system for kitchen supervision based on visual analysis. Background Technology

[0002] Currently, the hygiene, safety, and operational standardization of restaurant kitchens are directly related to consumer health, restaurant operations, and industry regulatory compliance, making them a core aspect of food safety supervision. Current kitchen supervision generally employs a traditional model of manual inspections, periodic spot checks, and manual review of monitoring data. This model suffers from drawbacks such as broad hardware coverage, limited intelligent application, low management efficiency, and weak risk control, and can no longer meet the demands for large-scale, routine, and real-time supervision.

[0003] Existing kitchen monitoring systems mostly rely on post-event traceability and manual monitoring, leading to issues such as inspector fatigue causing missed detections, delayed responses to violations, and untimely risk discovery. Furthermore, the characteristics of restaurant kitchens, including high staff turnover, inconsistent operational standards, dense risk points, and dispersed management across multiple canteens and areas, make it difficult for traditional manual methods to achieve comprehensive, seamless, and real-time control, as well as monitoring of factors such as rodent infestations and garbage. Simultaneously, the existing monitoring model lacks data support; alarm information is fragmented, handling processes are untraceable, and statistical analysis and report generation heavily depend on manual entry, failing to form a closed-loop management capability encompassing risk assessment, compliance analysis, and rectification tracking. Consequently, food safety hazards and violations continue to occur frequently.

[0004] In terms of the application of intelligent technologies, existing AI video monitoring solutions for kitchens have three major shortcomings: Architectural flaws: Either centralized cloud processing is adopted, requiring all video streams to be uploaded to the cloud, resulting in high end-to-end latency, large bandwidth consumption, high computing power costs, and poor high-concurrency stability, making it impossible to achieve second-level early warning and real-time intervention; or a single edge terminal is deployed in isolation, which can only identify locally and lacks the ability to achieve global cloud management, unified operation and maintenance, and multi-node data aggregation and analysis, making it difficult to adapt to the centralized supervision needs of multiple regions and multiple canteens.

[0005] Defects in implementation and adaptation: Functions are pre-built and fixed, and adding new violation types and supervision functions requires a lot of code modification, resulting in poor scalability; poor compatibility with existing monitoring equipment and logistics management systems for "transparent kitchens", often requiring the replacement of dedicated hardware, resulting in high deployment costs and long cycles; after scaling up, remote management of edge devices, model iteration, and configuration updates are difficult, resulting in high operation and maintenance costs.

[0006] Functional closed-loop defects: They generally only have a single function of "video recognition + basic alarm", and cannot connect the entire chain of video collection, violation identification, real-time alarm, regulation matching, analysis and judgment, report generation and rectification tracking. The process is fragmented and data is not shared. Compliance analysis, report production and rectification tracking are still completed manually, and a standardized, traceable and automated closed-loop supervision system has not been formed.

[0007] Therefore, how to provide an integrated technical solution that combines cloud-edge collaboration, intelligent decision-making, full-process automation, and compatibility with existing technologies to comprehensively improve the efficiency of kitchen supervision and risk control is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] In view of this, the present invention provides a kitchen supervision intelligent agent based on visual analysis to solve the problems existing in the background art.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: A kitchen supervision intelligent agent based on visual analysis includes: an edge perception layer, a cloud-edge collaborative transmission layer, a cloud-based intelligent agent core layer, and an interactive application layer; The edge perception layer is deployed at the kitchen monitoring points and is used to access existing monitoring video streams, complete image preprocessing, perform lightweight real-time reasoning of violations, provide local instant alarms, generate structured alarm logs, compress evidence data, and cache it locally. The edge-cloud collaborative transmission layer connects the edge perception layer and the cloud-based intelligent agent core layer. It is used to achieve low-latency bidirectional data transmission, distributed message peak reduction, remote status monitoring and unified management of edge devices, breakpoint resume and data synchronization based on a lightweight communication protocol. The cloud-based intelligent agent core layer serves as the decision-making brain of the intelligent agent, used for alarm data access, task complexity classification, dual-model dynamic routing scheduling, RAG knowledge matching based on the kitchen supervision knowledge base, violation judgment and analysis, automated report generation, rectification tracking, and full-process data memory storage. The interactive application layer is designed for regulatory personnel and is used for visualized global monitoring, real-time alarm push notifications, natural language command interaction, report display, and export in multiple formats.

[0010] Optionally, the edge-aware layer specifically includes: Kitchen surveillance camera group: directly connects to the existing "transparent kitchen" surveillance equipment to collect video streams; Video stream preprocessing module: performs preprocessing on the acquired video stream; Lightweight violation identification inference module: Deploys a scenario-optimized YOLOv8n lightweight object detection model to perform real-time inference on preprocessed video streams at the edge to identify violations; Alarm Structure Extraction Module: Extracts structured information from identified violations and automatically generates standardized alarm logs; Local caching and data compression module: compresses evidence-gathering videos; at the same time, it sets up a local caching mechanism to temporarily store alarm data when the network is interrupted, and automatically synchronizes it to the cloud through breakpoint resume after the network is restored; Edge-end local alarm push module: After identifying violations, it directly pushes local audio and visual alarms to the kitchen management personnel to achieve real-time intervention in violations.

[0011] Optionally, the preprocessing includes Gaussian denoising, HSV brightness enhancement, and image sharpening. Optionally, the model training method for the lightweight violation detection inference module includes: Construct a dedicated dataset for kitchen violations: Collect kitchen video streams covering different time periods, regions, and environments, extract video frames of typical violations and complete manual annotations to construct an annotated dataset; expand the training set through data augmentation algorithms; Basic Model Selection and Pre-training: The YOLOv8n lightweight object detection model was selected as the basic framework and pre-trained based on a dedicated dataset of kitchen violations to initially achieve the ability to identify kitchen violations; Model scenario-based optimization and lightweight transformation: The model is structured and pruned, and INT8 quantization is performed to reduce the number of model parameters and computational cost; the accuracy loss is compensated by teacher-student knowledge distillation; and the CBAM attention mechanism is introduced to enhance the extraction of violation features.

[0012] Optionally, the end-to-cloud collaborative transmission layer specifically includes: MQTT Lightweight Communication Protocol Module: Based on the MQTT 3.1.1 protocol, a bidirectional communication architecture is built, a unique Topic is assigned to each edge device, and a message transmission mechanism with QoS=1 is set. Distributed message queue module: It adopts Kafka distributed message queue to smooth out the peaks and fill the valleys of massive alarm logs uploaded from the edge, and supports stable reception of concurrent alarm data from thousands of cameras; Device status monitoring and remote management module: Real-time monitoring of the operating status, computing power usage, model version, and online status of all edge devices, and supports remote command issuance from the cloud; Breakpoint resume and data synchronization module: Enables bidirectional data synchronization between the edge and the cloud, automatically verifies data differences after network recovery, and completes breakpoint resume; at the same time, it supports batch synchronization and update of cloud monitoring rules and recognition models.

[0013] Optionally, the core layer of the cloud-based intelligent agent specifically includes: Intelligent Core Layer: The scheduling center of intelligent agents, responsible for the intelligent distribution, scheduling and execution of all supervisory tasks; Knowledge Service Layer: The knowledge support system for intelligent agents, providing accurate legal and business knowledge for analysis and judgment; External infrastructure layer: Provides underlying support for intelligent agents, including model services, data storage, and integration with third-party systems.

[0014] Optionally, the intelligent core layer specifically includes: MCP Scheduling Center: Built on the Model Context Protocol, it serves as the unified entry point for all tasks, responsible for receiving alarm logs from the edge and user-issued instructions, and completing task access, authentication, and distribution. Task Complexity Classifier (TCC): A predefined task feature lexicon for kitchen supervision is used to perform text normalization and feature matching on received tasks, automatically determine task complexity, and provide a basis for model scheduling. Scheduler: Based on the classification results of the task complexity classifier TCC, it performs dual-model dynamic routing, distributes tasks to the corresponding large language models, and simultaneously calls the RAG engine to obtain matching knowledge and issues tool call instructions to the executor. Executor: According to the scheduler's instructions, it matches and calls the corresponding tool function from the tool registry to complete the specific operation, and is also responsible for the return and verification of the execution results; Memory module: Based on SQLite database and sliding window mechanism, it realizes persistent storage and traceability of session history, alarm data and rectification records.

[0015] Optionally, the knowledge service layer specifically includes: Kitchen Supervision Knowledge Base: This base integrates national / local food safety regulations, kitchen management regulations, catering service operation standards, and equipment operation and maintenance manuals, serving as a fundamental resource library for knowledge retrieval. ChromaDB Vector Database: Stores knowledge base content that has been intelligently segmented and vectorized, and supports efficient vector similarity retrieval; RAG Engine: Adopts a lazy loading design, the embedded model and database connection are not loaded when the agent starts, and resources are initialized only when a retrieval request is triggered; it is responsible for vector retrieval, result rearrangement and context assembly; Tool Registry: Implemented using Python metaprogramming, it automatically scans and registers all monitoring tools, supporting automatic tool discovery, hot-swapping, and intelligent invocation. Optionally, the external facility layer specifically includes: Local lightweight large language model: The Qwen3-1.7B lightweight model deployed on the local server is responsible for handling simple alarm queries and command execution tasks; High-performance large language model in the cloud: It connects to the large cloud model API and is responsible for handling complex tasks; SQLite, a structured database, stores structured data. Kitchen Operations Platform API Interface: Provides standardized RESTful APIs that can seamlessly integrate with existing logistics management systems and food safety supervision platforms.

[0016] Optionally, the interactive application layer specifically includes: Web-based visual management platform: Built on Vue.js, it includes a real-time monitoring panel, alarm pop-ups, violation data statistics, rectification tracking management, and supports global visual control of multiple canteens and multiple areas; Mobile alarm push module: Supports real-time push of violation alarm information to managers via WeChat and SMS; Natural language interaction window: Allows administrators to issue commands via natural language, and the intelligent agent to automatically parse and complete the corresponding operations; Report Export and Display Module: Enables the visual display of compliant reports and supports export in multiple formats.

[0017] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a kitchen supervision intelligent agent based on visual analysis, which has the following beneficial effects: 1. Real-time edge sensing + cloud-based intelligent decision-making completely solves the latency problem. Edge devices complete identification and alarm locally with short latency, truly achieving "second-level early warning and on-site prevention" of violations. Supports centralized operation and maintenance, remote upgrades, and unified rule distribution for multiple canteens, multiple regions, and multiple devices, solving the shortcomings of pure edge devices that are "each fighting its own battle and unable to be managed in a unified manner".

[0018] 2. Full-process automation, eliminating manual supervision: Upgrading from "identification and alarm" to "full lifecycle supervision agent," it achieves a complete closed loop of video collection → violation identification → real-time alarm → regulatory matching → analysis and judgment → report generation → rectification tracking. Compliance reports are generated automatically, replacing manual table creation. Daily / weekly / monthly reports are automatically output, charts are automatically generated, and multiple formats can be exported, improving supervision efficiency. Natural language interaction, zero-threshold use: logistics personnel do not need professional training; they can query, analyze, and export reports using voice / text, truly making it "usable for everyone."

[0019] 3. Deeply Verticalized RAG for More Accurate and Resource-Efficient Regulatory Retrieval: Intelligent hierarchical segmentation ensures fragmented regulatory knowledge and more accurate retrieval. Intelligent segmentation by legal clauses and chapters avoids semantic fragmentation caused by traditional fixed-length segmentation, resulting in higher retrieval accuracy. Lazy-loading the RAG engine reduces resource consumption; it only loads when not in use and automatically releases resources when idle, solving the problem of traditional RAGs "filling up memory upon startup," enabling lightweight deployment. Every alert is automatically matched with legal basis, ensuring that violation analysis is "legally sound and verifiable," eliminating the illusion of large models and meeting regulatory inspection requirements. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the intelligent agent structure provided by the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] To address the core shortcomings of existing AI-powered kitchen monitoring technologies, this invention constructs an integrated technical solution combining real-time visual perception at the edge and a cloud-based intelligent agent for closed-loop decision-making throughout the entire process. Employing a layered decoupling and deep cloud-edge fusion design philosophy, this embodiment discloses a visual analysis-based intelligent agent for kitchen monitoring. Figure 1 As shown, the intelligent agent is divided into four layers: edge perception layer, cloud-edge collaborative transmission layer, cloud-cloud intelligent agent core layer, and interactive application layer. Each layer is loosely coupled and can be independently developed, iterated and expanded, while realizing seamless bidirectional flow of data and instructions.

[0024] Edge perception layer: undertakes video parsing and violation detection tasks with high real-time requirements, and solves the problems of high latency and high bandwidth consumption; Cloud-edge collaborative transmission layer: Enables low-latency, high-reliability data interaction between the edge and the cloud, while also achieving remote unified management and control of edge devices; The core layer of the cloud-based intelligent agent: the "decision-making brain" of the intelligent agent, which realizes the automated processing of complex tasks such as alarm analysis, regulation matching, compliance assessment, and report generation. It is the core innovative module that distinguishes this invention from existing technologies. Interactive application layer: Provides managers with a low-threshold, visual operation interface to enable the issuance of regulatory instructions and feedback of results.

[0025] (1) Edge sensing layer

[0026] The RK3588 edge computing devices deployed at various monitoring points in the kitchen are the "visual nerve endings" of the intelligent agent. All computing tasks with high real-time requirements are completed locally, without the need to upload the entire video stream to the cloud, thus fundamentally solving the problems of high latency and high bandwidth consumption in traditional centralized cloud solutions.

[0027] Kitchen monitoring camera group: Compatible with mainstream camera protocols such as RTSP / ONVIF, it can be directly connected to the existing monitoring equipment deployed in the "transparent kitchen" without the need to replace dedicated hardware, which greatly reduces deployment and transformation costs.

[0028] Video stream preprocessing module: Performs real-time frame extraction, Gaussian denoising, HSV brightness enhancement, image sharpening and other customized processing on the acquired video stream to solve the problem of poor image quality caused by kitchen fumes and uneven lighting, and provide high-quality input for subsequent inference.

[0029] Lightweight violation identification inference module: Deploys a scenario-optimized YOLOv8n lightweight target detection model to complete real-time inference of video frames locally at the edge, accurately identifying typical violations such as not wearing a mask / chef's hat, not wearing work clothes, rodent infestation, making or receiving phone calls, smoking, and littering. End-to-end inference latency is ≤300ms, and the recognition accuracy is ≥98%.

[0030] The alarm structure extraction module extracts structured information from identified violations and automatically generates standardized alarm logs containing the time of violation, region, type of violation, duration, and evidence materials, providing standardized data input for cloud analysis.

[0031] Local caching and data compression module: H.264 encoding is used to compress evidence videos to reduce transmission bandwidth consumption; at the same time, a local caching mechanism is set up to temporarily store alarm data when the network is interrupted, and automatically synchronize it to the cloud through breakpoint resume after the network is restored to avoid data loss.

[0032] Edge-end local alarm push module: After identifying violations, it directly pushes local audible and visual alarms to the kitchen management personnel, enabling immediate intervention in violations without waiting for cloud instructions, further improving the speed of regulatory response.

[0033] Lightweight visual violation recognition implementation process in the kitchen: Construct a dedicated dataset for kitchen violations: Collect kitchen video streams covering different time periods, regions, and environments, extract video frames of typical violations and complete manual annotation, and construct an annotated dataset with a scale of ≥100,000 frames; expand the training set through data augmentation algorithms such as rotation, blur, and lighting adjustment to improve the robustness of the model.

[0034] Basic model selection and pre-training: The YOLOv8n lightweight object detection model was selected as the basic framework and pre-trained based on a dedicated dataset to initially achieve the ability to identify violations in the kitchen.

[0035] Model scenario-based optimization and lightweight transformation: The model is structured and pruned and INT8 quantized to reduce the number of model parameters and computational cost; the accuracy loss is compensated by teacher-student knowledge distillation; and the CBAM attention mechanism is introduced to enhance the extraction of violation features. Ultimately, the number of model parameters is reduced by more than 70%, and the recognition accuracy is ≥98%.

[0036] Customized image preprocessing algorithm development: Develop preprocessing algorithms such as Gaussian denoising, HSV brightness enhancement, and image sharpening to solve image quality problems caused by kitchen fumes and uneven lighting, and improve recognition accuracy in complex environments.

[0037] Model deployment and inference optimization: The optimized model is converted to the RKNN format adapted to RK3588 to complete the NPU inference environment adaptation; the inference pipeline is optimized for multi-threading to achieve parallel execution of video frame acquisition, preprocessing and inference, ensuring that the inference latency of a single video stream is ≤300ms.

[0038] Real-time video stream inference and violation identification: Edge devices access the RTSP video stream from the camera, extract video frames in real time, and input them into the inference model after preprocessing to complete the real-time identification and location of violations.

[0039] Structured alarm extraction and local push: Structured information of violations is extracted to generate standardized alarm logs; at the same time, early warnings are pushed through the local audio and visual alarm module to achieve real-time intervention in violations.

[0040] (2) End-to-cloud collaborative transmission layer

[0041] The edge-cloud collaborative transmission layer is a "data highway" connecting the edge and the cloud. It is responsible for achieving low-latency, high-reliability two-way data interaction, while also completing the remote unified management and control of edge devices, solving the problem that traditional single edge solutions cannot achieve large-scale operation and maintenance and global management.

[0042] The MQTT lightweight communication protocol module is based on the MQTT 3.1.1 protocol to build a bidirectional communication architecture. It assigns a unique Topic to each edge device and sets a message transmission mechanism with QoS=1. Compared with the HTTP protocol, it reduces bandwidth consumption by more than 70% and the transmission latency is ≤1s, making it suitable for complex network environments in the kitchen.

[0043] Distributed message queue module: Using Kafka distributed message queue, it smooths out the peaks and fills the valleys of massive alarm logs uploaded from the edge, and can support the stable reception of alarm data from thousands of cameras, avoiding congestion and data loss in the cloud system under high concurrency scenarios.

[0044] Device status monitoring and remote management module: Real-time monitoring of the operating status, computing power usage, model version, and online status of all edge devices. Supports remote cloud-based command issuance of model updates, configuration modifications, restarts, and debugging instructions, enabling centralized operation and maintenance of distributed edge devices.

[0045] Breakpoint resume and data synchronization module: Enables bidirectional data synchronization between the edge and the cloud, automatically verifies data differences after network recovery, and completes breakpoint resume; at the same time, it supports batch synchronization and update of cloud monitoring rules and recognition models.

[0046] The cloud-edge collaborative low-latency data interaction technology process is as follows: Cloud-edge communication architecture construction: Build a two-way communication architecture of "edge-end proactive push + cloud-end on-demand retrieval", build the communication core based on the MQTT protocol, assign a unique Topic to each edge device, and set a message transmission mechanism with QoS=1 to ensure reliable message delivery.

[0047] Edge-end data preprocessing and compression: The edge end performs structured processing on alarm logs and uses H.264 encoding to compress evidence materials, increasing the compression ratio to over 100:1 and significantly reducing transmission bandwidth consumption.

[0048] Real-time data push based on MQTT: After generating alarm logs at the edge, the data is immediately pushed to the cloud via the MQTT protocol, with an end-to-end transmission delay of ≤1s, realizing real-time cloud upload of alarm data.

[0049] Cloud-based distributed message queue peak shaving: The cloud uses Kafka message queues to smooth out peaks and fill valleys in massive concurrent alarm data, buffering the impact of high-concurrency traffic and ensuring the stability of intelligent agents in scenarios with thousands of concurrent cameras.

[0050] Local caching and breakpoint resumption for network anomalies: When the network is interrupted, the edge device temporarily stores the data in the local cache; after the network is restored, the data differences are automatically verified and the breakpoint resumption is completed to ensure that alarm data is not lost or duplicated.

[0051] Remote status monitoring and management of edge devices: Real-time collection of information such as the operating status, computing power usage, and online status of edge devices in the cloud, and visualization display; automatic push of alarms when device failure occurs, realizing centralized operation and maintenance of distributed devices.

[0052] Cloud-based command delivery and batch model updates: The cloud sends commands such as configuration modification and model update to edge devices via the MQTT protocol. After the edge devices execute the commands, they send back the results, enabling remote unified management and batch iteration of edge devices.

[0053] (3) Cloud-based intelligent agent core layer

[0054] The core layer of the cloud-based intelligent agent is the "decision brain" of this invention. It adopts a hierarchical intelligent agent architecture to achieve fully automated processing from alarm reception, regulatory matching, analysis and judgment to report generation and instruction execution, completely solving the core defects of existing technologies such as fragmented functions and lack of closed-loop management.

[0055] The core intelligent layer serves as the central scheduling hub for intelligent agents, responsible for the intelligent distribution, scheduling, and execution of all oversight tasks. Specifically, it includes: MCP Scheduling Center: Built on Model Context Protocol (MCP), it serves as the unified entry point for all tasks, responsible for receiving alarm logs from the edge and user-issued instructions, and completing task access, authentication, and distribution.

[0056] Task Complexity Classifier (TCC): It uses a predefined feature lexicon for kitchen supervision tasks to perform text normalization and feature matching on received tasks, automatically determining task complexity and providing a basis for model scheduling.

[0057] Scheduler: Based on the TCC classification results, execute dual-model dynamic routing to distribute tasks to the corresponding large language models, while calling the RAG engine to obtain matching knowledge and issuing tool call instructions to the executor.

[0058] Executor: Based on the scheduler's instructions, it matches and calls the corresponding tool functions from the tool registry to complete specific operations such as data analysis and chart generation, and is also responsible for the return and verification of execution results.

[0059] Memory module: Based on SQLite database and sliding window mechanism, it realizes persistent storage and traceability of session history, alarm data and rectification records, and provides context support for large model inference.

[0060] Knowledge Service Layer: The knowledge support system for intelligent agents, providing accurate legal and business knowledge for analysis and judgment. Specifically, it includes: Kitchen Supervision Knowledge Base: Integrating exclusive content such as national / local food safety regulations, kitchen management charters, catering service operation standards, and equipment operation and maintenance manuals, it serves as the basic resource library for knowledge retrieval.

[0061] ChromaDB Vector Database B33: Stores knowledge base content that has been intelligently segmented and vectorized, and supports efficient vector similarity retrieval.

[0062] RAG Engine: Adopting a lazy-loading design, the embedded model and database connection are not loaded when the agent starts. Resources are only initialized when a retrieval request is triggered. Memory consumption is reduced by more than 70% when the agent is idle. It is responsible for vector retrieval, result rearrangement, and context assembly, providing accurate knowledge support for large model inference.

[0063] Tool Registry: Implemented based on Python metaprogramming, it automatically scans and registers all monitoring tools, supports automatic tool discovery, hot-swapping, and intelligent invocation, and allows adding new functions without modifying the main program code.

[0064] External infrastructure layer: Provides underlying support for intelligent agents, including model services, data storage, and integration with third-party systems. Specifically, it includes: Local lightweight large language model: The Qwen3-1.7B lightweight model deployed on the local server is responsible for handling simple alarm queries and command execution tasks.

[0065] High-performance cloud-based large language model: It connects to cloud-based large model APIs such as Minimax and is responsible for handling complex tasks such as violation trend analysis, compliance assessment, and customized report generation, giving full play to the complex reasoning capabilities of high-end models.

[0066] SQLite, a structured database, stores structured data such as alarm logs, agent configurations, rectification records, and report data, and supports efficient conditional queries and data statistics.

[0067] Kitchen Operations Platform API Interface: Provides standardized RESTful APIs that can seamlessly connect with existing logistics management systems and food safety supervision platforms to achieve data sharing and business collaboration.

[0068] The implementation process of the RAG-enhanced regulatory agent is as follows: Kitchen supervision knowledge base material collection and preprocessing: Collect unstructured documents such as food safety regulations, university kitchen management regulations, and operating procedures, complete format cleaning, deduplication, compliance verification, remove invalid content, and form standardized knowledge base materials.

[0069] Intelligent hierarchical segmentation based on priority strategy: The document is intelligently segmented using a priority strategy. Markdown structured segmentation is performed first, followed by segmentation according to legal clauses, then by manual chapters, and finally by catch-all recursive character segmentation. At the same time, parent heading information is injected into the metadata of child slices to avoid semantic fragmentation and ensure the semantic integrity of the segmented blocks.

[0070] Document vectorization and vector database construction: Using the bge-small-zh-v1.5 Chinese embedding model, the segmented document blocks are converted into vector data, which are then bound to the original text and metadata and stored in the ChromaDB vector database to build a dedicated vector knowledge base for kitchen supervision.

[0071] Lazy-loaded RAG engine development: Develop a lazy-loaded RAG engine where the embedded model and database connection are not loaded when the agent starts up. Resources are only initialized through a double-checked locking mechanism when a retrieval request is triggered for the first time. Resources are automatically released after the agent has been idle for more than a threshold, reducing idle memory consumption by more than 70%.

[0072] Automatic alarm data retrieval and knowledge matching: After receiving alarm logs, the agent automatically extracts keywords such as violation type and scenario to generate a search query. The RAG engine performs vector retrieval, matches relevant laws and regulations and management requirements from the database, and returns Top-K related document blocks.

[0073] Search results and context assembly: The retrieved legal knowledge, alarm data, historical records, and other content are assembled into a complete prompt according to a preset template, providing accurate support for large model inference and avoiding the generation of false information.

[0074] Large-scale model compliance analysis and judgment output: The assembled prompt is input into the scheduled large language model, and combined with legal knowledge and alarm data, the compliance analysis of violations, the matching of rectification basis, the generation of rectification suggestions are completed, and standardized judgment results are output.

[0075] Dual-model dynamic routing scheduling implementation process: Construct a dedicated task feature word library for kitchen supervision: a predefined scenario feature word library, with feature words for complex tasks including "analysis, report, judgment, trend, customized report", etc., and feature words for simple tasks including "query, status, basic Q&A, instruction execution", etc., to provide a basis for task classification.

[0076] Task access and text normalization: After receiving the task, the MCP scheduling center performs text normalization processing, including removing special characters, word segmentation, stop word filtering, and extracting the core keywords of the task.

[0077] TCC Task Complexity Classification and Judgment: The task complexity classifier matches core keywords with the feature word library, calculates the task complexity score, and determines the task as complex if it exceeds the threshold, and as simple if it does not exceed the threshold.

[0078] Task routing and distribution: Simple tasks are routed to local lightweight models for processing, ensuring data privacy and response speed; complex tasks are routed to high-performance models in the cloud for processing, leveraging the complex inference capabilities of high-end models.

[0079] Differentiated prompt word template loading: Prompt word templates are designed for local and cloud models respectively. Local templates focus on instruction execution and accurate question answering, simplifying reasoning steps; cloud templates focus on in-depth analysis and logical reasoning, guiding the model to complete complex judgment tasks.

[0080] Model execution inference and task processing: The corresponding model receives the assembled prompt, performs inference, and outputs the processing results.

[0081] Execution result verification and feedback: The executor verifies the format and compliance of the model output results. After the verification is passed, the results are fed back to the interactive application layer or the next execution stage.

[0082] (4) Interactive application layer

[0083] The interactive application layer is the interaction entry point between the intelligent agent and the user. It is designed with a low-threshold and visual operation interface, taking into account the characteristics of logistics management personnel who are not professional and technical personnel.

[0084] Web-based visual management platform: Built on Vue.js, it includes core functions such as real-time monitoring dashboard, alarm pop-ups, violation data statistics, and rectification tracking management, and supports global visual management and control of multiple canteens and multiple areas.

[0085] Mobile alarm push module: Supports real-time push of violation alarm information to managers via WeChat and SMS, enabling supervision response anytime, anywhere.

[0086] Natural language interaction window: It allows managers to issue commands in natural language, and the intelligent agent can automatically parse and complete the corresponding operations, which greatly reduces the threshold for intelligent agent operation.

[0087] Report Export and Display Module: Enables the visual display of compliance reports, supports exporting in multiple formats such as Excel, PDF, and images, and meets the reporting needs of daily management and superior inspections.

[0088] The entire workflow of this invention is executed automatically without human intervention, forming a complete regulatory closed loop, as detailed below: The kitchen's existing surveillance cameras collect real-time video streams, which are then connected to edge computing devices. At the edge, video preprocessing and lightweight model inference are performed to identify violations in real time, generate structured alarm logs, and push local alarms for immediate intervention. The edge device pushes alarm data and evidence materials to the cloud in real time via the MQTT protocol. The cloud uses a message queue to perform data peak smoothing and inputs the alarm data into the intelligent agent MCP scheduling center. The agent's TCC classifier determines the task complexity, schedules the corresponding model, and retrieves matching regulatory knowledge through the RAG engine to support alarm analysis. The executor calls the corresponding tools to complete multi-dimensional analysis of alarm data, judgment of violation trends, and automatically generate rectification suggestions and standardized compliance reports. Analysis results and alarm information are pushed to administrators via web platforms and mobile devices. Administrators can issue instructions such as rectification tracking and historical queries through natural language interaction windows, and the intelligent agent will automatically complete the corresponding operations. All alarm data, rectification records, and report data are persistently stored to achieve full-process traceability and form a closed-loop kitchen supervision process of "real-time monitoring - timely early warning - post-event analysis - closed-loop rectification".

[0089] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0090] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual analysis-based intelligent agent for kitchen supervision, characterized in that, include: Edge perception layer, cloud-edge collaborative transmission layer, cloud-cloud intelligent agent core layer, and interactive application layer; The edge perception layer is deployed at the kitchen monitoring points and is used to access existing monitoring video streams, complete image preprocessing, perform lightweight real-time reasoning of violations, provide local instant alarms, generate structured alarm logs, compress evidence data, and cache it locally. The edge-cloud collaborative transmission layer connects the edge perception layer and the cloud-based intelligent agent core layer. It is used to achieve low-latency bidirectional data transmission, distributed message peak reduction, remote status monitoring and unified management of edge devices, breakpoint resume and data synchronization based on a lightweight communication protocol. The cloud-based intelligent agent core layer serves as the decision-making brain of the intelligent agent, used for alarm data access, task complexity classification, dual-model dynamic routing scheduling, RAG knowledge matching based on the kitchen supervision knowledge base, violation judgment and analysis, automated report generation, rectification tracking, and full-process data memory storage. The interactive application layer is designed for regulatory personnel and is used for visualized global monitoring, real-time alarm push notifications, natural language command interaction, report display, and export in multiple formats.

2. The intelligent kitchen monitoring agent based on visual analysis according to claim 1, characterized in that, The edge-aware layer specifically includes: Kitchen surveillance camera group: Directly connects to the existing "transparent kitchen" surveillance equipment to collect video streams; Video stream preprocessing module: performs preprocessing on the acquired video stream; Lightweight violation identification inference module: Deploys a scenario-optimized YOLOv8n lightweight object detection model to perform real-time inference on preprocessed video streams at the edge to identify violations; Alarm Structure Extraction Module: Extracts structured information from identified violations and automatically generates standardized alarm logs; Local caching and data compression module: compresses evidence-gathering videos; at the same time, it sets up a local caching mechanism to temporarily store alarm data when the network is interrupted, and automatically synchronizes it to the cloud through breakpoint resume after the network is restored; Edge-end local alarm push module: After identifying violations, it directly pushes local audio and visual alarms to the kitchen management personnel to achieve real-time intervention in violations.

3. The intelligent kitchen monitoring agent based on visual analysis according to claim 2, characterized in that, The preprocessing includes Gaussian denoising, HSV brightness enhancement, and image sharpening.

4. The intelligent kitchen monitoring agent based on visual analysis according to claim 2, characterized in that, The model training method for the lightweight violation detection inference module includes: Construct a dedicated dataset for kitchen violations: Collect kitchen video streams covering different time periods, regions, and environments, extract video frames of typical violations and complete manual annotations to construct an annotated dataset; expand the training set through data augmentation algorithms; Basic Model Selection and Pre-training: The YOLOv8n lightweight object detection model was selected as the basic framework and pre-trained based on a dedicated dataset of kitchen violations to initially achieve the ability to identify kitchen violations; Model scenario-based optimization and lightweight transformation: The model is structured and pruned, and INT8 quantization is performed to reduce the number of model parameters and computational cost; the accuracy loss is compensated by teacher-student knowledge distillation; and the CBAM attention mechanism is introduced to enhance the extraction of violation features.

5. The intelligent kitchen monitoring agent based on visual analysis according to claim 1, characterized in that, The endpoint-cloud collaborative transmission layer specifically includes: MQTT Lightweight Communication Protocol Module: Based on the MQTT 3.1.1 protocol, a bidirectional communication architecture is built, a unique Topic is assigned to each edge device, and a message transmission mechanism with QoS=1 is set. Distributed message queue module: It adopts Kafka distributed message queue to smooth out the peaks and fill the valleys of massive alarm logs uploaded from the edge, and supports stable reception of concurrent alarm data from thousands of cameras; Device status monitoring and remote management module: Real-time monitoring of the operating status, computing power usage, model version, and online status of all edge devices, and supports remote command issuance from the cloud; Breakpoint resume and data synchronization module: Enables bidirectional data synchronization between the edge and the cloud, automatically verifies data differences after network recovery, and completes breakpoint resume; at the same time, it supports batch synchronous updates of cloud monitoring rules and recognition models.

6. The intelligent kitchen monitoring agent based on visual analysis according to claim 1, characterized in that, The core layer of the cloud-based intelligent agent specifically includes: Intelligent Core Layer: The scheduling center of intelligent agents, responsible for the intelligent distribution, scheduling and execution of all supervisory tasks; Knowledge Service Layer: The knowledge support system for intelligent agents, providing accurate legal and business knowledge for analysis and judgment; External infrastructure layer: Provides underlying support for intelligent agents, including model services, data storage, and integration with third-party systems.

7. A visual analysis-based intelligent agent for kitchen supervision according to claim 6, characterized in that, The intelligent core layer specifically includes: MCP Scheduling Center: Built on the Model Context Protocol, it serves as the unified entry point for all tasks, responsible for receiving alarm logs from the edge and user-issued instructions, and completing task access, authentication, and distribution. Task Complexity Classifier (TCC): A predefined feature lexicon for kitchen supervision tasks is used to perform text normalization and feature matching on received tasks, automatically determining task complexity and providing a basis for model scheduling. Scheduler: Based on the classification results of the task complexity classifier TCC, it performs dual-model dynamic routing, distributes tasks to the corresponding large language models, and simultaneously calls the RAG engine to obtain matching knowledge and issues tool call instructions to the executor. Executor: According to the scheduler's instructions, it matches and calls the corresponding tool function from the tool registry to complete the specific operation, and is also responsible for the return and verification of the execution results; Memory module: Based on SQLite database and sliding window mechanism, it realizes persistent storage and traceability of session history, alarm data and rectification records.

8. A visual analysis-based intelligent agent for kitchen supervision according to claim 6, characterized in that, The knowledge service layer specifically includes: Kitchen Supervision Knowledge Base: This base integrates national / local food safety regulations, kitchen management regulations, catering service operation standards, and equipment operation and maintenance manuals, serving as a fundamental resource library for knowledge retrieval. ChromaDB Vector Database: Stores knowledge base content that has been intelligently segmented and vectorized, and supports efficient vector similarity retrieval; RAG Engine: Adopts a lazy loading design, the embedded model and database connection are not loaded when the agent starts, and resources are initialized only when a retrieval request is triggered; it is responsible for vector retrieval, result rearrangement and context assembly; Tool Registry: Implemented based on Python metaprogramming, it automatically scans and registers all monitoring tools, and supports automatic tool discovery, hot-swapping, and intelligent invocation.

9. A visual analysis-based intelligent agent for kitchen supervision according to claim 6, characterized in that, The external facility layer specifically includes: Local lightweight large language model: The Qwen3-1.7B lightweight model deployed on the local server is responsible for handling simple alarm queries and command execution tasks; High-performance large language model in the cloud: It connects to the large cloud model API and is responsible for handling complex tasks; SQLite, a structured database, stores structured data. Kitchen Operations Platform API Interface: Provides standardized RESTful APIs that can seamlessly integrate with existing logistics management systems and food safety supervision platforms.

10. A visual analysis-based intelligent agent for kitchen supervision according to claim 1, characterized in that, The interactive application layer specifically includes: Web-based visual management platform: Built on Vue.js, it includes a real-time monitoring panel, alarm pop-ups, violation data statistics, rectification tracking management, and supports global visual control of multiple canteens and multiple areas; Mobile alarm push module: Supports real-time push of violation alarm information to managers via WeChat and SMS; Natural language interaction window: Allows administrators to issue commands via natural language, and the intelligent agent to automatically parse and complete the corresponding operations; Report Export and Display Module: Enables the visual display of compliant reports and supports export in multiple formats.