Intelligent and automated monitoring system and method for environments where food is prepared.
An intelligent system with interpretive AI and local triggers in security cameras provides real-time compliance monitoring in food handling environments, enhancing accuracy and efficiency through adaptive learning and feedback.
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Filing Date
- 2024-12-20
- Publication Date
- 2026-07-07
AI Technical Summary
Existing methods for monitoring food handling and preparation environments, such as supermarkets and bakeries, are costly and inefficient, requiring human oversight and limited automation.
An intelligent system using interpretive AI and security cameras with local triggers for real-time image analysis, providing probabilistic answers to specific questions about compliance with safety protocols, and incorporating a feedback loop for continuous improvement.
Enhances monitoring accuracy and efficiency by adapting to specific environments, allowing for rapid corrective actions and continuous learning to improve compliance monitoring.
Smart Images

Figure 00000000_0000_ABST
Description
1 / 12 Intelligent and automated monitoring system and method for environments where food is prepared. Field of application of the invention
[001] The present invention relates to a system and method for intelligent and automated monitoring of commercial environments, more specifically for environments where food is handled and prepared, such as supermarkets, bakeries and similar establishments, using advanced image analysis and artificial intelligence to answer specific operational questions, activating analyses at critical moments through local triggers and facilitating real-time interaction for proactive and informed management.
[002] Some acronyms and technical terms used here are:
[003] API - Application Programming Interface: a set of functions implemented in a program, made available for other programs and applications to use directly in a simplified way.
[004] Backend - Structure that enables the operation of a system, while the frontend is responsible for the visual part: presentation, design, languages, colors, etc.
[005] Chatbot - A computer program that simulates and processes written or spoken human conversations, allowing people to interact with digital devices as if they were communicating with a real person.
[006] EASiBox - Local machine equipped with GPU. Current models run on Nvidia GPUs for ease of use and integration, which may evolve depending on new hardware offerings dedicated to AI.
[007] PPE - Personal Protective Equipment: any and all equipment, device or product for individual use that must be used by the worker, in order to protect him or her against risks that could threaten his or her safety and health. Petition 870240108864, dated 12 / 20 / 2024, page 5 / 26 2 / 12
[008] GPT - Generative Pre-trained Transformer, a type of LLM (Large Language Model) or large-scale language model, composed of a neural network with many parameters (typically billions or possibly more).
[009] GPU - Graphics Processing Unit, or also VPU (visual processing unit), is a type of microprocessor specialized in processing graphics in personal computers, workstations or video games.
[010] Interpretive AI - Artificial Intelligence capable of producing details or reasons that led to the results presented, allowing it to be understood by humans and making its operation more transparent.
[011] LLaVA - Large Language and Vision Assistant: leverages vision and language models to provide detailed descriptions of images or image sequences, contextualizing objects, scenes, and actions, based on deep visual analysis and AI-generated semantic interpretation.
[012] OpenPose - Human pose estimation system, capable of detecting and tracking the human body in real time, accurately estimating the body's pose in 3D space.
[013] OpenAI - Company / laboratory that uses LLM to conduct AI research, with the intention of promoting and developing friendly AI. Description of the State of the Art
[014] The inspection of activities in food handling and preparation areas is traditionally carried out through on-site inspections or, in some cases, by installing video cameras, with the captured images being monitored by human operators in centralized facilities. Such methods involve high costs and limited efficiency.
[015] Patent document CN111461017 (A), entitled High-precision recognition method for city-scale catering kitchen work garment, refers to a method of identification. Petition 870240108864, dated 12 / 20 / 2024, page 6 / 26 3 / 12 for kitchen work attire (caps, gloves, aprons, etc.) in restaurants for use by municipal hygiene departments. To this end, a method is provided comprising the following steps:
[016] create a dedicated database;
[017] train the OpenPose human keypoint extraction network and extract human keypoints in the image;
[018] process the images of the circumscribed rectangular frame of the human body or the circumscribed rectangular frame of each body part (e.g., head, hands, etc.) according to the key points obtained in the previous step;
[019] forward the image of the area of the circumscribed rectangular frame of the human body part obtained in the previous step for comparison with a dedicated database, for recognition of the garment corresponding to the human body part;
[020] to provide feedback to the supervised establishment with the result of the comparison, that is, whether a missing piece of clothing was detected among the staff of that establishment.
[021] Although computer vision offers automation, it requires specific configurations for each context. Objectives of the Invention
[022] In view of the foregoing, the main objective of the invention is to introduce an intelligent system that employs local triggers for automatic context-based analysis, improving the accuracy and relevance of monitoring.
[023] Another objective is to optimize real-time monitoring through automatic image analysis, providing quick and accurate answers to specific questions.
[024] Yet another objective is to improve decision-making and operational management, with AI providing direct analysis to the responsible agents, who can then take action. Petition 870240108864, dated 12 / 20 / 2024, page 7 / 26 4 / 12 immediate corrective actions or refine the system with feedback.
[025] Yet another objective is to establish a cycle of continuous improvement, where team feedback adjusts and improves the accuracy of the AI, aligning the system with the specificities and needs of the retail environment. Brief Description of the Invention
[026] The invention aims to achieve the above objectives through the use of interpretive AI, which adapts flexibly to various monitoring scenarios, without the need for specific programming for each case.
[027] This is an innovative solution that combines the power of Interpretive Artificial Intelligence with security camera monitoring in food handling and preparation areas.
[028] The system is not limited to simple binary detections; instead, it employs OpenAI's GPT to analyze images and provide probabilistic answers to specific questions formulated in text, such as, for example, the probability of an employee not wearing safety gloves.
[029] The method comprises, firstly, capturing the image by at least one camera and processing the image by an embedded computer vision software, in order to detect the presence of human beings in the image. If such a presence is detected, the image data is analyzed by an AI system that interprets the context of the image in order to verify non-conformities.
[030] The captured images are processed by the system according to the following steps: i) the image captured by the camera identifies the presence of humans; ii) further processing, if humans are present, to identify the body part of interest (hands, head, etc.); Petition 870240108864, dated 12 / 20 / 2024, page 8 / 26 5 / 12 iii) capturing images of the body part being examined and forwarding them for comparison with images stored in a database; iv) if a non-conformity is determined, the probability of exceeding a predetermined threshold is calculated; (v) Generating a text message to alert the department or employee.
[031] If an evaluation error has occurred on the part of the system (for example, indicating that gloves are not being used when they are being used), the images captured by the camera are fed back into the database, after an evaluation stage by experts.
[032] The process begins with the system using local triggers, such as sensors or preliminary image analysis, to capture snapshots (images) at relevant moments. These triggers can be generated automatically periodically or according to criteria defined in advance by an operator.
[033] These images are then sent to the GPT model, along with predefined questions that seek to assess specific aspects of compliance and safety, such as, for example, “What is the probability (1 in 100) that the person in this image is not wearing gloves?” Such aspects are defined beforehand by an operator and vary according to the nature of the establishment.
[034] GPT analyzes the image and returns a probability, which the proposed system compares with a predefined threshold to determine whether an alert should be sent.
[035] If the probability exceeds the threshold, indicating a potential non-conformity, the system automatically notifies the responsible personnel at the handling location, allowing for rapid corrective action via voice or text message generated by the GPT.
[036] In addition, the proposed system allows supervised personnel to provide feedback on the alerts received. Petition 870240108864, dated 12 / 20 / 2024, page 9 / 26 6 / 12
[037] Thus, if an employee receives a non-compliance alert but disputes its validity (for example, if the gloves are skin-colored and are not correctly detected by the system), he or she can inform the system, which uses this information to refine its future analyses.
[038] This feedback loop enriches the model's learning and adjusts the system to the peculiarities of the specific retail environment, continuously improving its accuracy and relevance. Description of the Figures
[039] The invention will be better understood through the description of exemplary embodiments and the respective figures, in which:
[040] Figure 1 is a general diagram of the system, illustrating the information flows.
[041] Figure 2 is a flowchart of the invention method.
[042] Figure 3 illustrates the system's data layers.
[043] Figures 4, 5 and 6 illustrate the screens to be filled in, defining the trigger, the content of the question and the reset of the check interval.
[044] Figures 7, 8 and 9 show various images captured by a camera, in an example of the application of the invention. Detailed Description
[045] The proposed system is an innovative system that integrates security cameras (101a ... 101n) with interpretive artificial intelligence to monitor retail environments, such as supermarkets, in real time, capturing images that are analyzed by AI to answer predefined questions, such as, for example, Is the operator wearing gloves?
[046] The questions are linked to specific cameras and are triggered by new snapshots. When the AI identifies a significant probability of non-compliance, such as Petition 870240108864, dated 12 / 20 / 2024, page 10 / 26 7 / 12 If an operator is not wearing gloves, this information is immediately communicated to the responsible supervisor via a mobile application.
[047] The system is reinforced by a knowledge base (104) dedicated to each question, which improves with interactions, and a robust backend (102,103) that coordinates the relationship between cameras, questions and decision logic.
[048] Its configuration comprises the following aspects:
[049] Registration: the system allows the user to register the customer, the store, the sector, the cameras, the supervisors and the specific questions that they want the AI to analyze, creating an organizational structure within the software.
[050] Question Association: each question is associated with the relevant stores, cameras and sectors, establishing the basis for focused and contextual analyses.
[051] Backend: set of programs that manages the execution of the method.
[052] Knowledge Base: a specific knowledge base for each combination of establishment and questions, storing essential information and analysis histories, facilitating the evolution of the accuracy of AI responses.
[053] Access by Supervisors and Operators of the establishment: they can access the system to view analyses, receive alerts and interact with the system.
[054] Image monitoring: the system continuously monitors the arrival of new snapshots in the storage of the GPU-Accelerated Visual Description AI Central Processing Unit (103), preparing for analysis.
[055] Chatbot: an integrated chatbot (106) provides a direct communication channel between operators and the system, allowing the sending of feedback and the receiving of non-compliance alerts.
[056] Call to CPU: when the necessary snapshots are available, the system queries the GPU Accelerated Visual Description AI (103), sending the images and the query Petition 870240108864, dated 12 / 20 / 2024, page 11 / 26 8 / 12 corresponding to supervisors and / or previously registered employees.
[057] Response Analysis and Decision Making: based on the probability provided by GPU Accelerated Visual Description AI (103), the system decides whether to issue an alert and notifies the relevant supervisor if necessary;
[058] History of Analyses: the history of all analyses performed is maintained for future reference, allowing for review and improvement of the decision-making process.
[059] Knowledge Base Update and Snapshot Management: after each action or feedback, the knowledge base (104) is updated and the snapshots used are marked as obsolete to optimize storage and system efficiency.
[060] Referring now to Fig. 1, the proposed system comprises cameras (101) installed in the monitored environment, with the images captured by them, at predefined times and intervals, being sent to an EASiBox concentrator (102), which is equipped with a GPU (Graphical Processing Unit). In this device the images are subjected to a preliminary analysis, and those that meet certain criteria are selected, such as, for example, the presence of a human being.
[061] The selected images are sent to the GPU (103), which houses an Accelerated Visual Description AI, that is, an artificial intelligence model designed to analyze and describe images in real time, using the local processing power of GPUs to perform intensive computer vision calculations.
[062] This type of technology, an LLaVA (Large Language and Vision Assistant), leverages vision and language models to provide detailed descriptions of images or sequences of images, contextualizing objects, scenes, and actions based on deep visual analysis and AI-generated semantic interpretation.
[063] In the aforementioned analysis, the AI model subjects the images to analysis based on specific questions. These questions are previously formulated by the operator via the terminal (105), with the set of these questions being stored in the question knowledge base (104). Petition 870240108864, dated 12 / 20 / 2024, page 12 / 26 9 / 12
[064] The method illustrated by the flowchart in Fig. 2 comprises the following steps:
[065] 201 - The question is programmed to be executed by random trigger, at predefined time intervals dt1;
[066] 202 - Sensors and AI embedded in security cameras allow capturing images at the best possible view for interpretation;
[067] 203 - The camera images (the number of images is defined in the question) are forwarded along with the question and the return of the trigger information (e.g., human presence detected, white clothing detected, etc.) to the API (Application Programming Interface);
[068] 204 - The API provides a response comprising two pieces of information: the first is a numerical value indicating the probability of a non-conformity and the second is an explanatory text of this determination such as, for example: “Our analysis system indicated that you appear to be without gloves in the production area. If this is true, please put on gloves immediately (...). If this analysis is incorrect and you are wearing gloves, please reply only with 'NO' to inform us. Thank you!
[069] 205 - The numerical value of the probability is compared with the pre-established threshold;
[070] 206 - If the probability exceeds the threshold, a message is sent to the employee and / or supervisor, (106) in Fig. 1, alerting them to the non-conformity;
[071] 207 - The employee and / or local manager who receives the notification verifies whether the non-conformity actually occurs. If they agree with the observation, they will proceed with corrective action; otherwise, they will report that they disagree, consequently alleging that the AI made a mistake;
[072] 208 - Additionally, the system reduces the verification interval (capture of a new image) in order to verify if the non-conformity has been corrected within a given time; Petition 870240108864, dated 12 / 20 / 2024, page 13 / 26 10 / 12
[073] 209 - If the request is unjustified (for example, the employee is wearing skin-colored gloves), the supervisor returns a “NO”; the situation will be entered into the knowledge base after filtering by a specialized team. This is done to ensure that the model is not corrupted.
[074] Fig. 3 shows the system's data layers, based on the company's reference framework.
[075] Layer 1 is the customer, comprising one or more stores, where the system of the invention is installed and runs.
[076] Layer 2 is the hardware itself, comprising several cameras that can monitor PPE in hazardous areas or food handling areas. Each camera can be configured to respond to various questions triggered at the appropriate time (camera and / or sensor analysis) to capture a snapshot and analyze it.
[077] Layer 3 concerns the receipt and interpretation of the response, which allows intervention by a client employee (at the call center or in the store, which could be a supervisor or the operator himself.
[078] Fig. 4 illustrates the non-conformity configuration screen that is filled in by the user. The purpose is to define the data that allows the creation of an exception rule that occurs N times during a predefined time period and a certain day of the week, regardless of the number of occurrences in the area under observation.
[079] Thus, if an employee turns on a machine 600 times, but the operator has set N=10, then an analysis will be chosen at random once every 60 occurrences. In this way, the Large Language and Vision Assistant software or its equivalent can be used to process other exceptions.
[080] Fig. 5 shows the screen that defines the trigger firing rule, such as, for example, “daily random trigger”
[081] Note that several rules can be triggered at the same time (using the same trigger). Petition 870240108864, dated 12 / 20 / 2024, page 14 / 26 11 / 12
[082] Fig. 6 is a screen that allows the operator to choose the interval between successive checks.
[083] Figs. 7, 8 and 9 illustrate the operation of the method, comprising the following steps:
[084] Step a) User Definition
[085] Considering a camera with a view into the back office of a supermarket bakery, the user poses the following question: Confirm whether the employees at the workstation are wearing gloves.
[086] The rule is defined to occur every working day between 10:00 and 11:00 with a threshold of 60%.
[087] Stage b) Interactive Analysis Sequence between Generative Intelligence and Embedded Computer Vision Models Example
[088] On Monday at 10:12, the rule is activated in the EASiBox. The EASiBox triggers the camera API to capture a snapshot, along with the corresponding metadata. The initial metadata indicates that there is no human presence in the image (Fig. 7).
[089] The EASiBox waits 30 seconds and repeats the process.
[090] On the second attempt, the camera metadata indicates the presence of a human being (Fig. 8)
[091] EASiBox then asks the following question: I need an analysis of this image to answer whether the employees at the workstation are wearing gloves or not. In addition to the contextual answer, I need a number from 0 to 100 that assesses the probability of the answer being negative. As supplementary information, the camera's computer vision model reports: "human being in the photo". Petition 870240108864, dated 12 / 20 / 2024, page 15 / 26 12 / 12
[092] The answer is as follows: “Regarding the specific question about the use of gloves by employees at their workstations, the image is not clear enough to confirm with certainty the presence or absence of gloves.
[093] Based on this response, EASiBox requests a new image after, for example, 10 seconds (Fig. 9), and rephrases the question with the same request for details and probability: “I need an analysis of this image to answer whether the employees at the workstation are wearing gloves or not. In addition to the contextual answer, I need a number from 0 to 100 that assesses the probability of the answer being negative.
[094] As additional information, the camera's computer vision model reports: “clothing color: white, mask: yes, cap: yes.”
[095] In this case, the answer is: “Based on this more detailed image, the probability that the employees are not wearing gloves is 90% ...
[096] The EASiBox then compares the probability information with the predefined threshold and, upon finding that 90% > 60%, triggers execution and immediately sends the following message to the employee: Good morning! Our analysis system indicated that you appear to be without gloves in the production area. If this is true, please put on gloves immediately, as this is an essential part of our hygiene practice. If this analysis is incorrect and you are wearing gloves, please reply with 'NO' to inform us. Thank you!
[097] Due to the employee's NO response, the system forwards the images for analysis by human operators and possible incorporation into the database. Petition 870240108864, dated 12 / 20 / 2024, page 16 / 26
Claims
1 / 2 CLAIMS 1. INTELLIGENT AND AUTOMATED MONITORING SYSTEM FOR ENVIRONMENTS WHERE THERE ARE FOOD PREPARATION AREAS characterized by comprising: - at least one camera (101) equipped with an embedded computer vision system; - a concentrator equipped with a GPU (102); - a central processing unit (103); - a knowledge base of questions (104); - an operator interface (105); - communication channels (106) with the personnel in charge of the tasks, as well as with the supervisors.
2. INTELLIGENT AND AUTOMATED MONITORING METHOD FOR ENVIRONMENTS WHERE FOOD PREPARATION IS PRESENTED characterized by comprising the steps of: - capturing images (201) using cameras equipped with an embedded computer vision system; - selecting (202) images containing people; - analyzing (203) images using interpretive artificial intelligence to detect non-conformities in the people included in the images; - generating responses (204) indicating probabilities related to the presence of said non-conformities; - comparing (205) said probabilities with predefined threshold values - submitting the results of said comparisons to GPT processing (206) generating statements in human language; Petition 870240108864, dated 12 / 20 / 2024, page 17 / 26 2 / 2 - transmitting these statements (206) to said people; - verify (207) the correction of said non-conformities by means of new image capture.
3. METHOD according to claim 2 characterized in that said nonconformities comprise the absence of clothing elements on the persons depicted in said images.
4. METHOD according to claim 3 characterized in that said clothing items comprise one or more of the group comprising a cap, gloves, mask and apron.
5. METHOD according to claim 2 characterized in that said image capture is carried out periodically, according to predefined criteria.
6. METHOD according to claim 2 characterized in that said image capture is done from predefined triggers. Petition 870240108864, dated 12 / 20 / 2024, pp. 18 / 26