Intelligent monitoring system and method for multi-mode behavioral specifications of catering employees based on deep learning

Through the deep learning-based intelligent monitoring system, the problem of multi-modal behavior recognition in complex environments of restaurant store camera monitoring systems has been solved, efficient detection and management of abnormal employee behavior has been achieved, and the level of intelligence in the catering industry has been improved.

CN120673340APending Publication Date: 2025-09-19COLORFUL GUIZHOU IMPRESSION NETWORK MEDIA CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510799732.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing camera surveillance systems in catering stores are unable to efficiently identify and manage employees' multi-modal abnormal behaviors under complex network environments, hardware facilities and scenarios, resulting in low inspection efficiency and high management costs.

Method used

An intelligent monitoring system based on deep learning is adopted, including an active video inspection module, image preprocessing, employee target element detection, multi-modal abnormal behavior analysis and abnormal information alarm management. It processes high-concurrency data through an asynchronous mechanism and combines it with the lightweight target detection model YOLOv5 to achieve accurate identification of employee behavior in different areas and graded alarms.

Benefits of technology

It improves the ability to detect abnormal behavior of employees in catering stores, improves operational efficiency and management level, reduces the workload of manual inspections, reduces management costs, and has strong technical replicability and promotion capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673340A_ABST
    Figure CN120673340A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent monitoring systems, and particularly relates to a catering employee multi-mode behavior specification intelligent monitoring system based on deep learning, which comprises an active video inspection module, an image preprocessing module, an employee target element detection module, a multi-mode abnormal behavior analysis module and an abnormal information alarm management module. According to the invention, the video inspection management platform, the data processing platform, the target detection model and the modeled behavior analysis algorithm are fused, and the capability of detecting the abnormal behaviors of the store employees under the conditions of poor network environment, poor hardware facilities, poor monitoring environment and complex scenes is greatly improved. And intelligent identification of nonstandard abnormal behaviors of catering store employees in various scenes and different modes is realized with low computing power cost. And the operation efficiency and the management level of the store are further improved, a solid foundation is laid for intelligent transformation of the catering industry, and the method has relatively high technical replicability and generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent monitoring systems, and specifically to a deep learning-based intelligent monitoring system and method for multi-modal behavior norms of catering employees. Background Art

[0002] (1) The development of artificial intelligence technology has promoted the intelligent transformation and upgrading of various industries. With the widespread application of deep learning algorithms in the field of video intelligent analysis, the manual video inspection and offline inspection methods for catering store management standards have gradually shifted to online intelligent methods. Taking the abnormal behavior and status recognition scenario of catering store employees as an example, the inspection personnel generally use offline inspections or video playback to check whether there are reception staff at the reception position, strangers intruding in the kitchen, employees smoking, employees playing with mobile phones, employees not wearing epaulettes, employees not wearing masks or hats, and abnormal behavior in work clothes. The manual offline inspection or video inspection method increases costs, is inefficient, and the inspection management effect is not obvious. The camera monitoring data is not fully utilized. At present, through deep learning computer vision technology, the management of catering employees' hygienic dress, on-the-job behavior, and strangers in the kitchen can be captured in real time and analyzed and automatically warned. The irregular behavior of employees or stores can be recorded and relevant pictures and videos can be saved for manual re-inspection, providing basic data for subsequent employee behavior standard management, improving monitoring efficiency and reducing the workload of manual inspections. By intelligently identifying and warning employee behavior, improper behavior can be corrected in a timely manner, improving the overall quality of catering services and reducing store management costs.

[0003] (2) The cameras installed at catering stores are mainly ordinary cameras without snapshot functions. These cameras are installed at different angles and positions, and the image frame rate configurations are different. The resolutions include 720*450, 1920*1080, 2560*1440 and 3660*2880. The lighting conditions in the stores are complex, and the cameras are affected by exposure, oil smoke, and back-kitchen processing equipment, making it difficult to guarantee the quality of the images obtained by the cameras. The effects of the same camera monitoring area vary significantly. The high-angle camera is far away from the activity area and covers a wide monitoring area, resulting in the employee target appearing smaller in the distant image. The low-positioned camera captures the employee in a larger area in the image. The behavioral standards for employees in the front hall, back kitchen and reception areas are different. The images obtained by ordinary cameras need to be fully recognized. These factors have posed challenges to the inspection speed, abnormal behavior detection and inference speed, recall rate and accuracy, concurrent processing capabilities, and multi-modal behavior recognition of the AI ​​system. Summary of the Invention

[0004] The purpose of the present invention is to provide a deep learning-based intelligent monitoring system and method for multimodal behavioral norms of catering staff, which solves the problems raised in the background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: A deep learning-based intelligent monitoring system for multimodal behavioral norms of catering employees, including five modules: active video inspection module, image preprocessing, employee target element detection, multimodal abnormal behavior analysis, and abnormal information alarm management.

[0006] Preferably, the active video inspection module is responsible for exchanging data between multiple cameras in multiple stores managed by the enterprise, including: 1) Define the store organizational structure, connect store cameras to the video networking platform, and obtain store camera data streams through the platform; 2) Encode store cameras and define the store and device locations corresponding to each camera, including the front office, back kitchen, reception, and stranger capture locations. Software development is then used to map camera information to the platform's parameter configuration management module for easy modification and maintenance. 3) Through software development, camera device locations and corresponding scheduled inspection tasks are mapped to the platform parameter configuration management module. Relevant personnel can use the platform parameter management module to set the device location, device task, and inspection time for each camera. For example, a camera located at the reception desk corresponds to the receptionist presence recognition task. This camera uses an electronic fence to analyze the receptionist's work area within a specified timeframe, with scheduled inspections set at 5-minute intervals. A camera located in the kitchen desk corresponds to the kitchen abnormal behavior analysis task, with scheduled inspections set at 30-minute intervals, and so on. 4) Since the number of camera channels connected to the platform exceeds 500, the backend pulls device video images in a concurrent manner to request AI model services. Therefore, there are practical problems such as long AI inference time and even timeouts caused by large data volume and high concurrency. Therefore, an asynchronous mechanism is used to implement message docking between the backend and the AI ​​model to ensure service timeout and consistent data transmission.

[0007] Preferably, the video stream data processing module performs validity verification and post-processing on the video stream acquired by the store camera; 1) The image data stream and timestamp at a specific moment are obtained through interception. To ensure the validity and integrity of the intercepted data stream, the image integrity is verified through a highly real-time file header and tail mark check + metadata consistency check algorithm; 2) To address resolution inconsistency, the image resolutions captured through video streams range from low to high: 720*450, 1920*1080, 2560*1440, and 3660*2880. Images with a 720*450 resolution are blurry, making it difficult to detect and analyze human features, necessitating camera replacement. Images with a 3660*2880 resolution are typically captured by wide-angle cameras deployed at higher locations, resulting in large image size and distortion, a small human area in the image, and unclear details. The other two resolutions can both effectively display store scenes. To balance the store's existing network conditions, AI recognition speed, and accuracy, and combined with the more complex constraints of analyzing abnormal human behavior in the front office, targeted processing is performed on camera images from different device locations. Specifically, the image resolution of cameras located in the kitchen is standardized to 1920*1080, while that of cameras located in the front office is standardized to 2560*1440. Images from front office cameras with resolutions above the standard are compressed, while those below the standard remain unchanged.

[0008] Preferably, human element detection adopts the lightweight target detection model yolov5 and innovatively designs the detection categories. Taking into account the actual conditions that store employees do not wear uniform clothing, the elements of the same category are split into multiple subcategories, which can effectively complete the recognition of abnormal behavior and action and improve the recognition accuracy.

[0009] Optimally, multi-modal abnormal behavior analysis converts the service standard requirements of catering employees into achievable AI problems. Abnormal behaviors include: smoking, playing with mobile phones, not wearing masks, not wearing hats, not wearing work clothes, not wearing epaulettes, receptionists not being on duty, and strangers breaking into the kitchen. Different areas of the store have different behavioral standard requirements for employees. The personnel factor detection results based on the camera are analyzed and judged according to different modes of logic to realize the recognition of abnormal behaviors in different modes.

[0010] Preferably, abnormal information alarm management uses human element detection and abnormal behavior judgment algorithms to retain data on abnormal behaviors in the system background. The retained fields include store, time, location, equipment code, image, and abnormal behavior, and provides retrieval and export functions using store, time range, and abnormal behavior as key fields. Enterprise managers have different early warning requirements and management specifications for different abnormal behaviors. According to actual management needs, abnormal behaviors are graded and counted, and according to different early warning levels, managers at all levels are notified through in-site messages and text messages.

[0011] A method for an intelligent monitoring system of multimodal behavior norms of catering employees based on deep learning is as follows: Abnormal behaviors include employees smoking, using mobile phones, not wearing hats or masks, strangers entering the kitchen entrance, front desk staff not wearing work uniforms, and being absent from the reception area. To detect and inspect such abnormal behaviors, the video inspection strategy module is responsible for accessing multiple on-site cameras, implementing a scheduled inspection task mechanism, and acquiring camera image data. After the inspection strategy is issued, the acquired store camera image data is processed in batches to achieve resolution and image size control; Then, after image preprocessing, the image is uploaded to the human feature detection model to obtain target category and location information. The detection capability is trained by the detection category data designed by the scheme. Next, we will analyze the human element recognition results in different scenarios using the multi-modal abnormal behavior analysis algorithm based on store business rules proposed in this solution. The results will be asynchronously returned to the requester and managed and displayed through the back-end management system. Finally, the alarms of different abnormal behaviors are graded and the number of times is counted, and the abnormal behavior information of different levels is sent to managers of different levels via text messages.

[0012] Compared with the prior art, the present invention has the following beneficial effects: This invention integrates a video inspection management platform, a data processing platform, a target detection model, and a patterned behavior analysis algorithm to significantly improve the ability to detect abnormal behaviors of store employees under conditions of poor network environments, hardware facilities, and monitoring environments, as well as complex scenarios. It achieves intelligent identification of irregular and abnormal behaviors of restaurant store employees in various scenarios and modes at a low computing cost. This further improves store operational efficiency and management level, lays a solid foundation for the intelligent transformation of the restaurant industry, and has strong technical replicability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a framework diagram of the intelligent monitoring system for multi-modal behavior norms of catering employees based on deep learning of the present invention; Figure 2 This is the prior knowledge graph for distinguishing between store clerks and customers based on epaulettes in the present invention; Figure 3 This is a diagram of the abnormal behavior analysis and management platform of the present invention. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0015] See also Figure 1-3 , a deep learning-based intelligent monitoring system for multimodal behavioral norms of catering employees, including five modules: active video inspection module, image preprocessing, employee target element detection, multimodal abnormal behavior analysis, and abnormal information alarm management.

[0016] Preferably, the active video inspection module is responsible for exchanging data between multiple cameras in multiple stores managed by the enterprise, including: 1) Define the store organizational structure, connect store cameras to the video networking platform, and obtain store camera data streams through the platform; 2) Encode store cameras and define the store and device locations corresponding to each camera, including the front office, back kitchen, reception, and stranger capture locations. Software development is then used to map camera information to the platform's parameter configuration management module for easy modification and maintenance. 3) Through software development, camera device locations and corresponding scheduled inspection tasks are mapped to the platform parameter configuration management module. Relevant personnel can use the platform parameter management module to set the device location, device task, and inspection time for each camera. For example, a camera located at the reception desk corresponds to the receptionist presence recognition task. This camera uses an electronic fence to analyze the receptionist's work area within a specified timeframe, with scheduled inspections set at 5-minute intervals. A camera located in the kitchen desk corresponds to the kitchen abnormal behavior analysis task, with scheduled inspections set at 30-minute intervals, and so on. 4) Since the number of camera channels connected to the platform exceeds 500, the backend pulls device video images in a concurrent manner to request AI model services. Therefore, there are practical problems such as long AI inference time and even timeouts caused by large data volume and high concurrency. Therefore, an asynchronous mechanism is used to implement message docking between the backend and the AI ​​model to ensure service timeout and consistent data transmission.

[0017] Preferably, the video stream data processing module performs validity verification and post-processing on the video stream acquired by the store camera; 1) The image data stream and timestamp at a specific moment are obtained through interception. To ensure the validity and integrity of the intercepted data stream, the image integrity is verified through a highly real-time file header and tail mark check + metadata consistency check algorithm; 2) To address resolution inconsistency, the image resolutions captured through video streams range from low to high: 720*450, 1920*1080, 2560*1440, and 3660*2880. Images with a 720*450 resolution are blurry, making it difficult to detect and analyze human features, necessitating camera replacement. Images with a 3660*2880 resolution are typically captured by wide-angle cameras deployed at higher locations, resulting in large image size and distortion, a small human area in the image, and unclear details. The other two resolutions can both effectively display store scenes. To balance the store's existing network conditions, AI recognition speed, and accuracy, and combined with the more complex constraints of analyzing abnormal human behavior in the front office, targeted processing is performed on camera images from different device locations. Specifically, the image resolution of cameras located in the kitchen is standardized to 1920*1080, while that of cameras located in the front office is standardized to 2560*1440. Images from front office cameras with resolutions above the standard are compressed, while those below the standard remain unchanged.

[0018] Preferably, human element detection adopts the lightweight target detection model yolov5 and innovatively designs the detection categories. Taking into account the actual conditions that store employees do not wear uniform clothing, the elements of the same category are split into multiple subcategories, which can effectively complete the recognition of abnormal behavior and action and improve the recognition accuracy.

[0019] (1) Specifically, the English and Chinese labels for the human element detection categories are as follows.

[0020]

[0021] Optimally, multi-modal abnormal behavior analysis converts the service standard requirements of catering employees into achievable AI problems. Abnormal behaviors include: smoking, playing with mobile phones, not wearing masks, not wearing hats, not wearing work clothes, not wearing epaulettes, receptionists not being on duty, and strangers breaking into the kitchen. Different areas of the store have different behavioral standard requirements for employees. The personnel factor detection results based on the camera are analyzed and judged according to different modes of logic to realize the recognition of abnormal behaviors in different modes.

[0022] (2) For example, in the front office scene, the human element detection results of non-employees are first excluded. Employees and customers can be distinguished by epaulettes, and the focus is on detecting whether employees smoke, play with mobile phones, wear epaulettes, or wear work clothes; the back kitchen fully detects whether employees smoke, play with mobile phones, wear work clothes, hats, and masks; the reception area mainly detects whether there are reception staff wearing epaulettes and work clothes within the specified time area; the back kitchen entrance mainly detects the faces of people entering and leaving the back kitchen to determine whether they are store employees. Specifically: 1) Pseudo-algorithm for identifying abnormal behavior of front office employees:

[0023] 2) Pseudo-algorithm for identifying abnormal behavior in the kitchen:

[0024] 3) Guest reception presence detection pseudo-algorithm: Detects employees in the reception area to determine whether there are any employees on duty within the specified time zone.

[0025]

[0026] 4) A stranger in the kitchen breaks into the pseudo-algorithm:

[0027] Preferably, abnormal information alarm management uses human element detection and abnormal behavior judgment algorithms to retain data on abnormal behaviors in the system background. The retained fields include store, time, location, equipment code, image, and abnormal behavior, and provides retrieval and export functions using store, time range, and abnormal behavior as key fields. Enterprise managers have different early warning requirements and management specifications for different abnormal behaviors. According to actual management needs, abnormal behaviors are graded and counted, and according to different early warning levels, managers at all levels are notified through in-site messages and text messages.

[0028] A method for an intelligent monitoring system of multimodal behavior norms of catering employees based on deep learning is as follows: Abnormal behaviors include employees smoking, using mobile phones, not wearing hats or masks, strangers entering the kitchen entrance, front desk staff not wearing work uniforms, and being absent from the reception area. To detect and inspect such abnormal behaviors, the video inspection strategy module is responsible for accessing multiple on-site cameras, implementing a scheduled inspection task mechanism, and acquiring camera image data. After the inspection strategy is issued, the acquired store camera image data is processed in batches to achieve resolution and image size control; Then, after image preprocessing, the image is uploaded to the human feature detection model to obtain target category and location information. The detection capability is trained by the detection category data designed by the scheme. Next, we will analyze the human element recognition results in different scenarios using the multi-modal abnormal behavior analysis algorithm based on store business rules proposed in this solution. The results will be asynchronously returned to the requester and managed and displayed through the back-end management system. Finally, the alarms of different abnormal behaviors are graded and the number of times is counted, and the abnormal behavior information of different levels is sent to managers of different levels via text messages.

[0029] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-modal behavior norm intelligent monitoring system for catering staff based on deep learning, characterized by: It includes five modules: active video inspection module, image preprocessing, employee target element detection, multi-mode abnormal behavior analysis, and abnormal information alarm management.

2. The deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: The active video inspection module is responsible for exchanging data between multiple cameras in multiple stores managed by the enterprise, including: 1) Define the store organizational structure, connect store cameras to the video networking platform, and obtain store camera data streams through the platform; 2) Encode store cameras and define the store and device locations corresponding to each camera, including the front office, back kitchen, reception, and stranger capture locations. Software development is then used to map camera information to the platform's parameter configuration management module for easy modification and maintenance. 3) Through software development, camera device locations and corresponding scheduled inspection tasks are mapped to the platform parameter configuration management module. Relevant personnel can use the platform parameter management module to set the device location, device task, and inspection time for each camera. For example, a camera located at the reception desk corresponds to the receptionist presence recognition task. This camera uses an electronic fence to analyze the receptionist's work area within a specified timeframe, with scheduled inspections set at 5-minute intervals. A camera located in the kitchen desk corresponds to the kitchen abnormal behavior analysis task, with scheduled inspections set at 30-minute intervals, and so on. 4) Since the number of camera channels connected to the platform exceeds 500, the backend pulls device video images in a concurrent manner to request AI model services. Therefore, there are practical problems such as long AI inference time and even timeouts caused by large data volume and high concurrency. Therefore, an asynchronous mechanism is used to implement message docking between the backend and the AI ​​model to ensure service timeout and consistent data transmission.

3. The deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: The video stream data processing module verifies the validity and post-processes the video streams obtained by store cameras; 1) The image data stream and timestamp at a specific moment are obtained through interception. To ensure the validity and integrity of the intercepted data stream, the image integrity is verified through a highly real-time file header and tail mark check + metadata consistency check algorithm; 2) To address resolution inconsistency, the image resolutions captured through video streams range from low to high: 720*450, 1920*1080, 2560*1440, and 3660*2880. Images with a 720*450 resolution are blurry, making it difficult to detect and analyze human features, necessitating camera replacement. Images with a 3660*2880 resolution are typically captured by wide-angle cameras deployed at higher locations, resulting in large image size and distortion, a small human area in the image, and unclear details. The other two resolutions can both effectively display store scenes. To balance the store's existing network conditions, AI recognition speed, and accuracy, and combined with the more complex constraints of analyzing abnormal human behavior in the front office, targeted processing is performed on camera images from different device locations. Specifically, the image resolution of cameras located in the kitchen is standardized to 1920*1080, while that of cameras located in the front office is standardized to 2560*1440. Images from front office cameras with resolutions above the standard are compressed, while those below the standard remain unchanged.

4. The deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: Human element detection uses the lightweight target detection model yolov5 and innovatively designs detection categories. Taking into account the actual conditions that store employees do not wear uniform clothing, elements of the same category are split into multiple subcategories, which can effectively complete the recognition of abnormal behavior and actions and improve recognition accuracy.

5. The deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: Multimodal abnormal behavior analysis transforms the service standard requirements of catering employees into achievable AI problems. Abnormal behaviors include: smoking, playing with mobile phones, not wearing masks, not wearing hats, not wearing work clothes, not wearing epaulettes, the receptionist is not on duty, and strangers break into the kitchen. Different areas of the store have different behavioral standards for employees. The personnel factor detection results based on the camera are analyzed and judged according to different modes of logic to realize the recognition of abnormal behaviors in different modes.

6. The deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: Abnormal information alarm management uses human element detection and abnormal behavior judgment algorithms to retain data on abnormal behaviors in the system background. The retained fields include store, time, location, equipment code, image, and abnormal behavior, and provides retrieval and export functions using store, time range, and abnormal behavior as key fields. Enterprise managers have different warning requirements and management specifications for different abnormal behaviors. According to actual management needs, abnormal behaviors are classified and counted, and according to different warning levels, managers at all levels are notified through in-site messages and text messages.

7. The method of a deep learning-based intelligent monitoring system for multi-modal behavior norms of catering staff according to claim 1 is characterized in that: The details are as follows: Abnormal behaviors include employees smoking, using mobile phones, not wearing hats or masks, strangers entering the kitchen entrance, front desk staff not wearing work uniforms, and being absent from the reception area. To detect and inspect such abnormal behaviors, the video inspection strategy module is responsible for accessing multiple on-site cameras, implementing a scheduled inspection task mechanism, and acquiring camera image data. After the inspection strategy is issued, the acquired store camera image data is processed in batches to achieve resolution and image size control; Then, after image preprocessing, the image is uploaded to the human feature detection model to obtain target category and location information. The detection capability is trained by the detection category data designed by the scheme. Next, we will analyze the human element recognition results in different scenarios using the multi-modal abnormal behavior analysis algorithm based on store business rules proposed in this solution. The results will be asynchronously returned to the requester and managed and displayed through the back-end management system. Finally, the alarms of different abnormal behaviors are graded and the number of times is counted, and the abnormal behavior information of different levels is sent to managers of different levels via text messages.

Citation Information

Cited By

  • Store operating state detection method and system, terminal and storage medium

    CN121438235A

  • Intelligent chain store patrol method and system based on multi-source video stream

    CN121640374A