Intelligent canteen processing method and system

By leveraging the two-way collaboration between audio pickup and monitoring equipment, and combining audio and video data, the needs of canteen users can be quickly identified, solving the problem of low service efficiency in traditional canteens and achieving efficient and accurate service response and intelligent management.

CN120852098APending Publication Date: 2025-10-28NANJING JIUYUE ZHILIAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760260.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-10-28

Smart Images

  • Figure CN120852098A_ABST
    Figure CN120852098A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent canteen processing method and system, and relates to the technical field of intelligent catering, and the method can collect the information of a target area from the two dimensions of audio and video through the cooperative work of a pickup device and a monitoring device, can accurately judge the demands of customers, achieves the efficient service response, and timely generates and distributes a service task. The efficiency and quality of canteen service are improved, the dining experience of customers is optimized, and meanwhile intelligent and fine management of canteen operation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart catering technology, and in particular to a smart canteen processing method and system. Background Technology

[0002] In today's era of rapid digital and intelligent development, canteens, as places where people gather frequently, directly impact users' dining experience and operational efficiency through their service efficiency and management level. As users increasingly demand more immediate and precise dining services, traditional service models relying on manual observation and passive response are no longer adequate to meet the diverse and complex service needs.

[0003] In traditional canteen service models, the collection and response to user needs primarily rely on two methods: manual registration and pager calls. Manual registration typically requires users to verbally state their needs to the service counter or mobile staff before or during their meal, such as requesting additional utensils or dishes. This method is not only inefficient, but during peak hours, staff often struggle to attend to all users, leading to missed requests or delayed responses. Furthermore, manual recording is prone to errors, affecting service accuracy. While pagers allow users to signal their needs by pressing a button, they only convey basic information about the service requested, failing to specify the exact requirements. Staff still need to physically go to the pager after receiving the signal, increasing response time. Additionally, issues like signal interference and pager malfunctions can prevent timely communication of requests.

[0004] Therefore, how to quickly and accurately identify user needs in complex environments has become an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a smart canteen processing method and system that can quickly and accurately identify user needs in complex environments.

[0006] A first aspect of the present invention provides a smart canteen processing method, comprising: Real-time acquisition of audio and video data corresponding to target areas that meet monitoring conditions; A two-way collaborative mechanism is constructed. When a trigger condition is detected in the audio data, the user status of the video data within the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal. When the user status and table status in the video data are detected as being in a triggered state, a service task is generated and sent to the execution terminal by combining the second keyword of the audio data within the second acquisition interval.

[0007] Optionally, in one possible implementation of the first aspect, real-time acquisition of audio and video data corresponding to the target area that meets the monitoring conditions includes: The monitoring conditions are met by determining that a user's dining area meets the monitoring criteria and is therefore identified as the target area. Control and monitoring equipment to collect video data of the target area in real time; Based on the user's location distribution in the video data, the directional sound pickup direction and pickup range of the sound pickup device are determined, and the sound pickup device is controlled to collect audio data in real time based on the directional sound pickup direction and pickup range.

[0008] Optionally, in one possible implementation of the first aspect, determining the directional pickup direction and pickup range of the audio pickup device based on the user's position distribution in the video data includes: The user's location distribution is determined based on the position coordinates of the person's outline in the video data; The coordinates of positions with a spacing less than a preset spacing are determined to be the same aggregation group. The device position coordinates of the pickup device are obtained, and the direction from the device position coordinates to the center coordinates of the aggregation group is determined as the directional pickup direction corresponding to the aggregation group. Based on the directional pickup direction, the angle is gradually expanded to both sides according to the angle increment until the expanded area completely covers all users in the aggregation combination. The range corresponding to the expanded area is determined as the pickup range corresponding to the directional pickup direction.

[0009] Optionally, in one possible implementation of the first aspect, it also includes: The average volume of the pickup device is obtained, and when the average volume is greater than or equal to a volume threshold, the pickup range is increased based on a preset angle.

[0010] Optionally, in one possible implementation of the first aspect, a two-way collaborative mechanism is constructed. When a trigger condition is detected in the audio data, the user status of the video data within the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal, including: Compare the first keyword of the audio data with the preset keyword. If the first keyword is the same as the preset keyword, it is determined that the audio data meets the triggering condition. Based on the current time, the video data within the first acquisition interval is obtained by tracing back a preset duration, and the user status in the video data is identified; When the user's state is a preset action, a service task is generated by combining the first keyword and sent to the execution terminal.

[0011] Optionally, in one possible implementation of the first aspect, when the user state and table state in the video data are detected as trigger states, a service task is generated and sent to the execution terminal by combining the second keyword of the audio data within the second acquisition interval, including: When the user state is a preset action, the second keyword and the preset keyword of the audio data in the second collection interval are compared. When the second keyword is the same as the preset keyword, a service task is generated and sent to the execution terminal. The detection frequency for identifying the table status is adjusted based on the amount of food changes on each plate and the frequency of the user's picking actions. When the table status is such that the number of empty plates is greater than or equal to the empty plate threshold, a service task is generated and sent to the execution terminal. The triggering status includes the user status being a preset action and the number of empty plates being greater than or equal to the empty plate threshold.

[0012] Optionally, in one possible implementation of the first aspect, the detection frequency for recognizing the table status is adjusted based on the amount of food change on each plate and the frequency of the user's picking-up actions, including: The initial coverage area of ​​each plate is obtained, the remaining coverage area of ​​the plate is recorded according to the initial recognition frequency, and the change in food quantity is obtained according to the ratio of the difference between the initial coverage area and the remaining coverage area to the initial coverage area. If the change in the food exceeds a first threshold, the detection frequency is adjusted by a first increasing frequency; if the change in the food is less than the first threshold, the detection frequency is adjusted by a first decreasing frequency. The frequency of the user's gripping actions on each plate is obtained. If the gripping action frequency is greater than a second threshold, the detection frequency is adjusted by a second increase frequency. If the gripping action frequency is less than the second threshold, the detection frequency is adjusted by a second decrease frequency.

[0013] Optionally, in one possible implementation of the first aspect, obtaining the frequency of the user's picking up actions on each plate includes: Identify the user's hand trajectory and obtain the duration of the hand trajectory's stay in the plate area and the number of times the direction changes; When the dwell time exceeds the clamping duration threshold and / or the number of orientation changes exceeds the number threshold, record one clamping count, and obtain the clamping action frequency based on the number of clampings per unit time.

[0014] Optionally, in one possible implementation of the first aspect, obtaining the number of directional changes of the hand trajectory in the plate area includes: Obtain the angle between the tangents of the hand trajectory at the previous and next moments. When the angle between the tangents is greater than the angle threshold, increase the number of direction changes by 1.

[0015] A second aspect of the present invention provides a smart canteen processing system, comprising: The acquisition module is used to acquire audio and video data corresponding to the target area that meets the monitoring conditions in real time; The first collaboration module is used to build a two-way collaboration mechanism. When a trigger condition is detected in the audio data, the user status of the video data in the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal. The second collaboration module is used to detect when the user status and table status in the video data are in a triggered state, and then, in combination with the second keyword of the audio data in the second acquisition interval, generate a service task and send it to the execution terminal.

[0016] The beneficial effects of this invention are as follows: 1. By constructing a two-way collaborative mechanism between audio pickup and monitoring devices, user needs are cross-validated and deeply analyzed from both audio and video dimensions. In audio-triggered scenarios, audio keywords are compared with preset keywords, and combined with user actions, expressions, and other state information in the video to avoid misjudgments caused by environmental noise; in video-triggered scenarios, by analyzing user actions and table status, combined with audio keywords for the corresponding time period, specific needs are clarified.

[0017] 2. Data acquisition and processing resources are dynamically adjusted based on changes in food quantity on plates, frequency of user picking up food, and ambient noise levels. By comparing changes in food quantity with preset thresholds and judging the frequency of picking up food with corresponding thresholds, the frequency of table status detection is adjusted in real time. The detection frequency is increased during peak dining hours to promptly capture changes in demand, and decreased during off-peak hours to conserve system resources. Simultaneously, the average volume of the audio pickup device is monitored and compared with a volume threshold to dynamically adjust the pickup range. In noisy environments, the pickup range is expanded to ensure the acquisition of effective audio information, while in quiet environments, the range is narrowed to improve audio acquisition accuracy. This dynamic adjustment mechanism achieves rational allocation of system resources, improves data processing efficiency, reduces resource consumption, and enhances the intelligence level of canteen operation and management.

[0018] 3. Based on accurate demand identification and intelligent resource allocation, service tasks can be quickly generated and assigned, enabling waiters to receive accurate service instructions in a timely manner, shortening response time and reducing user waiting time. Real-time monitoring of table status can promptly identify empty plates and generate tasks, maintaining a clean dining environment and improving user comfort. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present invention; Figure 2 This is a flowchart illustrating a smart canteen processing method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a smart canteen processing system provided in an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] See Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of the present invention. The present invention utilizes the collaborative work of a sound pickup device and a monitoring device to collect information from a target area from both audio and video dimensions. This allows for accurate assessment of customer needs, efficient service response, timely generation and allocation of service tasks, improved efficiency and quality of canteen services, optimized customer dining experience, and intelligent and refined management of canteen operations. The sound pickup device can be a microphone, the monitoring device can be a camera, and the target area refers to the area around the dining tables where users are dining.

[0022] See Figure 2 This is a flowchart illustrating a smart canteen processing method provided in an embodiment of the present invention. Figure 2 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps S1 to S3 are detailed as follows: S1 collects audio and video data in real time for the target area that meets the monitoring conditions.

[0023] Among them, monitoring conditions refer to determining whether the target area meets the data collection standards, which usually refers to the dining area where users are eating; audio data refers to the sound information in the canteen collected by the microphone, including user voices; video data refers to the images in the canteen captured by the camera, covering user actions, table status, and other content.

[0024] By collecting audio and video data, a comprehensive and timely data foundation can be provided for accurately capturing user service needs.

[0025] Specifically, based on the cafeteria's spatial layout and employee flow patterns during meals, high-definition cameras and high-sensitivity microphones can be installed in various dining areas. The camera distribution ensures coverage of all tables, clearly capturing every customer's movement and the table's status. Each table is equipped with a corresponding microphone to collect sound from the surrounding area. Once an employee enters and sits in a dining area, that area is identified as a target area meeting monitoring criteria. The cameras and microphones then begin collecting corresponding video and audio data in real time, enabling timely, comprehensive, and accurate acquisition of audio and video information related to user service needs within the cafeteria.

[0026] Based on the above embodiments, step S1 can be implemented in the following ways: The system determines that the dining area where users are present meets the monitoring conditions and identifies it as the target area; it controls the monitoring equipment to collect video data of the target area in real time; based on the location distribution of users in the video data, it determines the directional sound pickup direction and sound pickup range of the sound pickup device, and controls the sound pickup device to collect audio data in real time based on the directional sound pickup direction and sound pickup range.

[0027] Clearly defining the areas requiring focused monitoring and data collection, and avoiding ineffective data collection in areas without users, can save system resources while ensuring timely capture of user service needs and improving service response efficiency. Indiscriminate data collection from all areas of the cafeteria would generate a large amount of redundant information, increasing the system's processing burden. Precisely locating target areas allows system resources to be concentrated on areas with actual demand. Furthermore, accurately adjusting the pickup direction and range of the audio equipment based on the user's specific location within the target area can reduce interference from ambient noise, improve the targeting and accuracy of audio collection, and ensure that the collected audio data truly reflects the user's service needs.

[0028] In practical applications, a microphone can be composed of multiple individual microphones, such as combining multiple microphones with different directional properties, with each microphone responsible for a specific direction and range. For example, arranging several cardioid microphones at different angles can cover areas in multiple different directions.

[0029] Directional pickup direction refers to the direction in which the pickup device focuses on capturing sound based on the user's location. Pickup range refers to the area within which the pickup device captures audio. User location distribution refers to the distribution of users in various locations within the restaurant.

[0030] For example, by analyzing video data from table number 3, the system determines that the student is sitting on the left side of the table. Based on this location information, the system adjusts the combination microphones near the table, maximizing the sensitivity of one of the left-pointing microphones to determine its directional pickup direction. Using this direction as a reference, the system defines the corresponding pickup range. At this point, the pickup device only captures sound within this range, effectively filtering out noise from other directions, and can clearly capture the student's possible service requests, such as "Waiter, more soup."

[0031] In some embodiments, the directional pickup direction and pickup range of the audio pickup device can be determined based on the location distribution of users in the video data through the following steps: The location distribution of users is determined based on the position coordinates of the personnel outlines in the video data; the position coordinates of positions with a spacing less than a preset spacing are determined to be the same aggregation group, the device position coordinates of the sound pickup device are obtained, and the direction from the device position coordinates to the center coordinates of the aggregation group is determined as the directional sound pickup direction corresponding to the aggregation group; with the directional sound pickup direction as the reference, the angle is gradually expanded to both sides according to the angle increment until the expanded area completely covers all users in the aggregation group, and the range corresponding to the expanded area is determined as the sound pickup range corresponding to the directional sound pickup direction.

[0032] In practical applications, when determining the user's location coordinates, the target detection and localization model in deep learning algorithms can be used to identify human features and postures in the image and determine the pixel coordinates of the human outline in the image. Then, based on the pre-calibrated camera parameters and the actual coordinates of the fixed point in the canteen, a mapping relationship between the image coordinates and the actual physical coordinates can be established, converting the pixel coordinates into location coordinates in the actual space of the canteen, thereby accurately determining the user's location coordinates.

[0033] Understandably, grouping users in similar locations into the same cluster facilitates centralized voice acquisition, improving sound pickup efficiency. By determining the directional pickup direction, the pickup device can more accurately point at the target user group, reducing interference from other directions and improving the targeting and effectiveness of audio acquisition. This ensures clear capture of the service requests of users within the same cluster. By gradually expanding the angle based on the directional pickup direction, a pickup range that completely covers all users within the cluster is determined, ensuring no user's voice is missed and achieving comprehensive voice acquisition of the target user group. Simultaneously, dynamically adjusting the pickup range according to actual conditions adapts to different table layouts and user distribution, increasing flexibility.

[0034] The preset spacing is a pre-defined distance threshold used to determine whether users with similar location coordinates belong to the same cluster, for example, set to 1 meter. A cluster is a group of users whose location coordinates are less than the preset spacing. The device location coordinates are the coordinates of the microphone installed in the cafeteria on the restaurant floor plan, used to determine the relative positional relationship between the microphone and the user. The angle increment is a pre-defined value for each expansion to both sides, used to gradually adjust the pickup range, for example, set to 5°. The expanded area is a fan-shaped area formed by gradually expanding the angle to both sides based on the directional pickup direction; this area increases with the expansion of the angle.

[0035] For example, after determining the directional pickup direction, the angle increment is set to 5°. Starting from the directional pickup direction, the angle is gradually increased by 5° to the left and right. Using trigonometric functions, the fan-shaped area covered within this angle range, with the pickup device as the vertex and the maximum distance from the microphone to the edge of the table as the radius, is calculated and projected onto the actual floor plan of the cafeteria. It is determined whether this fan-shaped area completely covers the two users of aggregation combination 1. If not, the angle is further increased to both sides in 5° increments. After several adjustments, when the area expands to 30° from the left to the right of the directional pickup direction, the fan-shaped area completely covers the two users. At this point, this area from 30° to the right is determined as the pickup range corresponding to the directional pickup direction of aggregation combination 1.

[0036] In addition to the above embodiments, the following embodiments are also included: The average volume of the pickup device is obtained, and when the average volume is greater than or equal to a volume threshold, the pickup range is increased based on a preset angle.

[0037] The ambient noise level of a microphone significantly impacts its ability to capture user voice. In environments with high overall volume and significant ambient noise, the original pickup range may not be sufficient to capture enough user voice information, causing the user's intended speech signal to be drowned out by noise. By acquiring the average volume of the microphone and comparing it to a volume threshold, increasing the pickup range when the average volume is greater than or equal to the threshold allows for the collection of more sound data in noisy environments. This provides ample material for signal processing algorithms to extract and enhance the user's voice signal from the background noise, ensuring accurate capture of the user's speech even in noisy conditions and improving the reliability and effectiveness of audio acquisition.

[0038] Average volume refers to the average loudness of sound captured by the microphone over a period of time, usually measured in sound pressure level (dB). It reflects the overall intensity of the sound received by the microphone during this period. Volume threshold is a pre-set standard value used to determine whether the current environment is noisy. When the average volume of the microphone reaches or exceeds this threshold, the ambient noise is considered high, and appropriate measures need to be taken to adjust the pickup range. Preset angle is a pre-determined angle value; when it is necessary to increase the pickup range, the pickup angle range of the microphone is expanded in units of this angle.

[0039] By monitoring the average volume of the microphone in real time and dynamically adjusting the pickup range based on the volume, the microphone can better adapt to different environmental noise conditions. In noisy environments, timely expansion of the pickup range increases the amount of sound collected, providing more data support for extracting user voice signals from noise. This improves the ability to accurately capture user voice needs in complex environments, ensuring the timeliness and accuracy of the smart canteen system's response to user service requests.

[0040] S2, construct a two-way collaborative mechanism, when the trigger condition is detected for the audio data, obtain the user status of the video data within the first acquisition interval, and generate a service task by combining the first keyword of the audio data and send it to the execution terminal.

[0041] The cafeteria environment is complex and noisy, making it easy to misjudge service needs by relying solely on audio data. For example, there might be other unrelated conversations in the cafeteria containing keywords like "waiter," leading to numerous invalid tasks if triggered only by audio. By combining audio data with video data, image recognition algorithms can analyze the video to confirm whether customers are actually making corresponding actions, such as waving or standing up. Only when the demand keywords in the audio match the user's state shown in the video can service needs be determined more accurately.

[0042] By verifying and supplementing audio and video data, misjudgments that may result from a single data source can be avoided, user service needs can be more accurately determined, accurate service tasks can be generated, and service quality and efficiency can be improved.

[0043] The two-way collaborative mechanism is a working mode in which audio and video data cooperate and verify each other. Its core lies in closely linking audio and video data to analyze and judge user needs from different dimensions. The trigger condition is the appearance of keywords or phrases related to preset service needs in the audio data, reaching the standard set by the system to initiate subsequent analysis processes. The first acquisition interval is the video data acquisition period determined by tracing back a certain time from the audio data trigger time, used to obtain the user's status information at the time of triggering. User status refers to the user's actions and behaviors during the dining process, reflecting their service needs. The first keyword is the key text content related to service needs in the audio data, such as "waiter" or "get tissues." The service task is a specific service instruction generated based on the audio and video data analysis results, such as "[table number] customer needs tissues." The execution terminal is a device such as VR glasses worn by waiters, used to receive and display service task information.

[0044] For example, during lunchtime, a microphone at a table captures the audio of a customer saying, "Waiter, my rice is spilled." The system identifies "waiter" and "spilled rice" as the primary keywords, matches them with preset keywords, and determines that the trigger condition is met. The server immediately retrieves video data from the cameras in that area (the first acquisition interval) from 3 seconds before to 3 seconds after the audio trigger, based on the location information of the microphone at that table. By analyzing the video using an image recognition algorithm, it is found that the customer is gesturing with their hands, confirming that the user's state matches the audio request. The system integrates the request content in the audio with the customer's location information obtained from the image analysis, generates a service task "[Table 8] Customer spilled rice, needs to clean up and refill," and sends it to the VR glasses of the waiter closest to that table.

[0045] By leveraging the collaborative analysis of audio and video data, the generation of invalid tasks caused by misjudgment of a single audio signal is effectively reduced, ensuring that the generated service tasks accurately reflect the user's real needs, improving the relevance and effectiveness of the service, and enhancing the user's dining experience.

[0046] Based on the above embodiments, step S2 can be implemented in the following ways: Compare the first keyword and the preset keyword in the audio data. If the first keyword is the same as the preset keyword, it is determined that the audio data meets the triggering condition. Based on the current time, backtrack for a preset duration to obtain the video data within the first acquisition interval, and identify the user status in the video data. When the user status is a preset action, combine the first keyword to generate a service task and send it to the execution terminal.

[0047] By comparing the first keyword in the audio data with preset keywords, the system can quickly and accurately determine whether the user's voice information expresses a service need, thereby determining whether to trigger the subsequent task generation process. This avoids unnecessary processing of irrelevant audio data, improves system efficiency and targeting, and ensures that the system only responds to audio data related to service needs, promptly capturing user service requests. The preset keywords are a pre-defined set of keywords related to service needs, which the system uses to identify user service requirements. These keywords are set based on common service scenarios and user needs in the cafeteria.

[0048] After determining that the audio data meets the triggering conditions, video data within the first acquisition interval is obtained by tracing back a preset duration. This allows for the acquisition of user status information such as actions and expressions before and after the audio triggering time. Combining the user status in the video data can further verify whether the user truly has a service need, avoiding the generation of invalid tasks due to simple audio misjudgments (such as keywords accidentally appearing in environmental noise). By identifying user status, a more comprehensive basis is provided for more accurately determining the user's specific needs.

[0049] When the user's state in the video data is a preset action, such as waving or raising a hand, indicating a need for service, combining this with the first keyword in the audio data can accurately determine the user's specific service requirement and generate a corresponding service task, which is then sent to the execution terminal, such as the waiter's VR glasses. This ensures that the service task accurately reflects the user's needs, enabling waiters to provide timely and accurate services, thereby improving service quality and user satisfaction.

[0050] Preset actions are pre-defined actions that indicate a user's need for services, such as waving or raising their hand.

[0051] Through the above methods, a complete process from comprehensive analysis of audio and video data to service task generation and delivery is realized, accurately transforming users' service needs into specific service tasks and sending them to the execution terminal, enabling waiters to respond in a timely manner, improving the accuracy and efficiency of services, and enhancing users' dining experience and the service level of the canteen.

[0052] S3, when the user status and table status in the video data are detected as triggered, a service task is generated and sent to the execution terminal by combining the second keyword of the audio data in the second acquisition interval.

[0053] While video can capture visual information such as user actions and table status, it's insufficient to pinpoint specific user needs. For instance, when a camera detects a customer standing up and waving, it can only infer a possible service request, but it's unclear whether they need more food, payment, or something else. Furthermore, in some cases, a wave might be a call to a companion, making misjudgment highly likely when relying solely on video data. However, combining video data with audio data—analyzing the second keyword in the audio within the second capture interval after video triggering—can clarify the user's exact needs, improve service accuracy, and achieve multi-dimensional, comprehensive user need capture. This avoids overlooking user requests and enhances the comprehensiveness and proactiveness of cafeteria services.

[0054] The table status refers to the actual condition of the table, including the number of empty plates, the arrangement of tableware, and the amount of food remaining. The trigger status is when a user performs a preset action or when the number of empty plates is greater than or equal to a threshold. The second acquisition interval is the range of audio data collected within a certain period before and after the user's triggered state in the video data. The second keyword is a keyword related to the service requirement found in the audio data within the second acquisition interval.

[0055] By combining video and audio data and using precise audio collected by a combination microphone, we can promptly identify and meet the service needs expressed by users, improve the comprehensiveness of services and user satisfaction, and optimize the canteen service process to enhance operational efficiency.

[0056] Based on the above embodiments, step S3 can be implemented in the following ways: When the user's state is a preset action, the second keyword and the preset keyword of the audio data in the second acquisition interval are compared. When the second keyword is the same as the preset keyword, a service task is generated and sent to the execution terminal. The detection frequency when recognizing the table status is adjusted according to the change in the amount of food on each plate and the frequency of the user's picking action. When the table status is that the number of empty plates is greater than or equal to the empty plate threshold, a service task is generated and sent to the execution terminal. The triggering status includes the user's state being a preset action and the number of empty plates being greater than or equal to the empty plate threshold.

[0057] When the camera detects a user's pre-defined action, indicating a potential service need, the specific service requirement can be confirmed by comparing the second keyword in the audio data within the second acquisition interval with the pre-defined keyword. This avoids blindly generating service tasks based solely on user actions, improving the accuracy and relevance of task generation, ensuring that the generated service tasks effectively meet user needs, improving the quality and efficiency of cafeteria services, and reducing the generation of invalid tasks.

[0058] Different food changes and the frequency of user picking up food reflect the user's dining progress and needs. By adjusting the detection frequency when identifying the table status based on these factors, system resources can be allocated more rationally, improving the system's sensitivity and response speed to changes in table status. Increasing the detection frequency when food changes rapidly or picking up food frequently can promptly capture new user needs, such as adding more food; decreasing the detection frequency when food changes slowly or picking up food infrequently reduces unnecessary system computation and resource consumption, making the system operate more intelligently and efficiently.

[0059] When the number of empty plates is greater than or equal to the empty plate threshold, it indicates that the customer may have finished eating or is about to require new food or service. Generating a service task at this time can promptly meet the customer's potential needs, remove empty plates in a timely manner, improve the cleanliness of the dining environment, help to rationally plan plate clearing and cleaning work, improve work efficiency, and optimize the service process. For example, when the number of empty plates is greater than or equal to the empty plate threshold, a service task can be generated: "The customer has a large number of empty plates; it may be necessary to clear the plates or inquire whether they need more food." This task can then be sent to the VR glasses of nearby waiters, allowing them to understand the customer's needs in a timely manner.

[0060] By using preset actions and empty plate counts (greater than or equal to an empty plate threshold) as trigger states, the system expands its ability to capture user needs. By comprehensively judging user actions and the actual state of the dining table, the system can more fully discover user service needs, improve the initiative and comprehensiveness of canteen services, avoid overlooking user needs, and enhance user dining satisfaction.

[0061] Food change rate refers to the ratio of the difference between the remaining amount of food on the plate and the initial amount to the initial amount, used to measure the degree of food consumption. For example, if a plate initially contains 100 grams of food and 20 grams remain, the food change rate is 0.8. User picking-up frequency is the number of times a user picks up food from the plate per unit of time, reflecting the user's dining activity level and food demand. For example, if the user picks up food 5 times per minute, the picking-up frequency is 5 times per minute. Detection frequency is the time interval at which the system detects the table status. For example, the initial detection frequency is once every 2 minutes, and this frequency can be adjusted based on food change rate and picking-up frequency. Empty plate count is the number of empty plates on the table. The empty plate threshold is a pre-set standard for the number of empty plates. For example, setting the empty plate threshold to 3 means that when a table has 3 or more empty plates, the trigger condition is met.

[0062] In some embodiments, the detection frequency for identifying the table status can be adjusted based on the amount of food change on each plate and the frequency of the user's picking actions through the following steps: The system acquires the initial coverage area of ​​each plate, records the remaining coverage area of ​​the plates based on the initial recognition frequency, and calculates the food change based on the ratio of the difference between the initial and remaining coverage areas to the initial coverage area. If the food change is greater than a first threshold, the detection frequency is adjusted by a first increase frequency; if the food change is less than the first threshold, the detection frequency is adjusted by a first decrease frequency. The system also acquires the frequency of the user's gripping actions on each plate. If the gripping action frequency is greater than a second threshold, the detection frequency is adjusted by a second increase frequency; if the gripping action frequency is less than the second threshold, the detection frequency is adjusted by a second decrease frequency.

[0063] By dynamically adjusting the detection frequency of the food tray, system resources can be allocated more rationally. Different changes in food quantity and the frequency of user picking up food reflect varying dining progress and needs. When food quantity changes significantly or the picking up frequency is high, it indicates rapid changes in the tray's state; in this case, the detection frequency can be increased to accurately capture these changes and respond promptly to potential user needs, such as adding food. Conversely, when food quantity changes slightly or the picking up frequency is low, the tray's state is relatively stable; reducing the detection frequency minimizes unnecessary data processing, saves system resources, and improves system efficiency.

[0064] The initial coverage area of ​​each plate is the area covered by food when the user begins eating, obtained through image recognition technology, reflecting the initial amount of food on the plate. The initial recognition frequency is a pre-set interval for detecting and recognizing the plates at the beginning of the meal, such as once every 2 to 3 seconds, used to acquire initial data. The remaining coverage area is the area covered by remaining food on the plate after a certain time interval as the meal progresses, also recorded through image recognition technology. The first threshold is a pre-set standard value used to judge the magnitude of food change. When the food change is greater than this threshold, it is considered a significant decrease in food quantity; when it is less than this threshold, it is considered a slow change in food quantity. For example, the first threshold is set to 0.3. The first increase frequency is the magnitude by which the detection frequency is increased when the food change is greater than the first threshold. For example, if the original detection frequency is once every 3 seconds, it becomes once every 2 seconds after adjustment using the first increase frequency. The first decrease frequency is the magnitude by which the detection frequency is decreased when the food change is less than the first threshold. For example, the original detection frequency was once every 2 seconds, which is changed to once every 4 seconds after the first frequency reduction adjustment. The second threshold is a pre-set standard value used to judge the frequency of the user's gripping actions. When the gripping action frequency is greater than this threshold, the eating pace is considered fast; when it is less than this threshold, the eating pace is considered slow. For example, the second threshold is set to 5 times per minute. The second increase frequency is used to increase the detection frequency by a certain amount when the gripping action frequency is greater than the second threshold. The second decrease frequency is used to decrease the detection frequency by a certain amount when the gripping action frequency is less than the second threshold.

[0065] For example, at the start of the meal, the system uses image recognition technology from cameras installed in the cafeteria to obtain the initial coverage area of ​​each plate. Assuming an initial coverage area of ​​80 square centimeters for a plate, the initial recognition frequency is set to once every 2 seconds. At the 10th second after the start of the meal (after 5 recognitions), the remaining coverage area of ​​the plate is recorded as 50 square centimeters. The change in food quantity is calculated using a formula of 0.375, and a first threshold of 0.3 is set. Since 0.375 is greater than 0.3, meaning the change in food quantity exceeds the first threshold, the detection frequency is adjusted by increasing the detection frequency, for example, from once every 2 seconds to once every 1.5 seconds, to more promptly monitor changes in the amount of food on the plates.

[0066] Meanwhile, the system recognizes the user's hand movements through image recognition and counts the frequency of the user's picking up the plate. Within 1 minute, the system detects 7 picking up actions by the student. A second threshold is set at 5 times per minute. Since 7 is greater than 5, the picking up action frequency is greater than the second threshold. Based on the second increased frequency, for example, the detection frequency is further adjusted from once every 1.5 seconds to once every 1 second to quickly capture changes in the plate's state.

[0067] As the meal continues, after a period of time, the change in the amount of food on the plate is recalculated. If it becomes 0.1, which is less than the first threshold of 0.3, the detection frequency is adjusted back to once every 2 seconds according to the first reduction frequency. If the frequency of the picking action becomes once per minute, which is less than the second threshold of 5 times, the detection frequency is adjusted back to once every 3 seconds according to the second reduction frequency.

[0068] The detection frequency is dynamically adjusted based on changes in food quantity and the frequency of user picking up food. The detection frequency is increased when the plate status changes rapidly and decreased when changes are slow, avoiding resource waste and enabling the system to operate more efficiently, reducing unnecessary data processing. Timely adjustment of the detection frequency allows for more accurate monitoring of changes in the amount of food on the plate and the user's dining progress, helping to promptly identify user needs, such as adding food or clearing empty plates, thus improving the user's dining experience. By comprehensively considering changes in food quantity and the frequency of picking up food, the detection frequency is dynamically adjusted, providing richer and more accurate data support for more accurately determining whether a plate is empty, reducing the possibility of misjudgment and optimizing the cafeteria's service process.

[0069] In some embodiments, the frequency of a user's picking up and dropping off each plate can be obtained through the following steps: The system identifies the user's hand trajectory and obtains the duration of the hand trajectory in the plate area and the number of times the direction changes. When the duration of the hand trajectory is greater than the clamping duration threshold and / or the number of direction changes is greater than the number threshold, the system records one clamping count and obtains the clamping action frequency based on the number of clamping counts per unit time.

[0070] Accurately determining the frequency of a user's picking-up and picking-up actions is crucial for understanding their dining behavior and needs. By recognizing hand trajectories, the duration of time spent in the plate area, and the number of times the direction changes, it's possible to more precisely determine whether a user is actually picking up food. This avoids misinterpreting simply passing the plate as a picking-up action, thus improving the accuracy of the judgment.

[0071] The hand trajectory is the path formed by the user's hand movement in the video image, tracked and acquired using a target tracking algorithm. The plate area is the region within the video image determined by image recognition technology. The dwell time is the duration the hand trajectory remains within the plate area, calculated by recording the time the hand enters and leaves the area. The number of direction changes is the number of times the hand's direction changes significantly during its movement within the plate area. The gripping duration threshold is a pre-set time standard value used to determine if the hand dwell time is sufficient to constitute a gripping action. For example, setting the gripping duration threshold to 5 seconds means that a gripping action is satisfied when the hand dwells in the plate area for more than 5 seconds. The number of times threshold is a pre-set frequency standard value used to determine if the number of hand direction changes reaches the level required to constitute a gripping action. For example, setting the number of times threshold to 2 means that a gripping action is satisfied when the hand changes direction more than 2 times within the plate area. The number of gripping actions is recorded when the duration of hand stay in the plate area exceeds the gripping duration threshold and / or the number of directional changes exceeds the number threshold. The cumulative number of gripping actions is the gripping count.

[0072] For example, target tracking algorithms can be used to track a user's hand and obtain their hand trajectory. Suppose that within a certain time period, a user's hand is detected entering the plate area. The system begins recording the time the hand enters and leaves the plate area, calculating the dwell time as 1.5 seconds. Simultaneously, by analyzing changes in the hand trajectory, the number of directional changes is counted as 3. Setting a threshold of 1 second for the gripping duration and 2 for the number of gripping actions, since the dwell time exceeds the threshold and the number of directional changes exceeds the threshold, the system records this as one gripping action. Over the next minute, the system continuously monitors the user's hand movements, recording a total of 8 gripping actions using the same method. Therefore, the user's gripping action frequency during this minute is 8 times per minute.

[0073] By comprehensively considering the duration of hand movement on the plate area and the number of times the movement changes direction, and comparing this with a set threshold, it is possible to more accurately determine whether a user is actually picking up food. This reduces the possibility of misinterpreting a mere hand movement passing over the plate as a picking action, thus improving the accuracy of the judgment. The frequency of picking actions can intuitively reflect how often a user picks up food during the meal, providing quantitative data support for the cafeteria to understand users' dining habits.

[0074] In some embodiments, the number of times the hand trajectory changes direction in the plate area can be obtained through the following steps: Obtain the angle between the tangents of the hand trajectory at the previous and next moments. When the angle between the tangents is greater than the angle threshold, increase the number of direction changes by 1.

[0075] When analyzing hand trajectories to determine whether a user's hand action is for picking up food, changes in the direction of hand movement are a crucial factor. By acquiring the tangent angles of the hand trajectory at different moments and comparing them with an angle threshold, the changes in hand direction can be quantified more precisely. This is because when the hand is simply passing over the plate, its movement direction is usually relatively stable, and the tangent angle is small; however, during the process of picking up food, the hand needs to make various adjustments, leading to frequent changes in movement direction and often larger tangent angles. Therefore, this method can effectively distinguish between a simple passing motion and a picking action, thereby improving the accuracy of judging user hand movements.

[0076] The included angle threshold is a pre-defined standard value used to determine whether the included angle of the tangents is large enough to be considered a significant change in hand orientation. For example, if the included angle threshold is set to 30 degrees, a condition for determining a change in hand orientation is met when the included angle of the tangents is greater than 30 degrees.

[0077] For example, suppose at a certain time 1, the system acquires position point 1 on the user's hand trajectory at that time 1 and calculates tangent line 1 of the trajectory at that point; at the next time 2, it acquires position point 2 and calculates tangent line 2 of the trajectory at that point. Through geometric calculation, the angle between tangent line 1 and tangent line 2 is determined to be 40 degrees. Given that the system's set angle threshold is 30 degrees, since 40 degrees is greater than 30 degrees (i.e., the tangent angle is greater than the angle threshold), the system increases the direction change count by 1. As the user's hand continues to move, the system continuously repeats the above process, continuously monitoring and counting the number of direction changes to help determine the user's hand movements. For example, when the hand passes over a plate, if the calculated tangent angles are all small and do not exceed the angle threshold, the increase in the direction change count is small; however, when the hand is making a grasping motion, there will be more instances where the tangent angle is greater than the angle threshold, and the number of direction changes will increase accordingly.

[0078] By calculating the angle between the tangents and comparing it with a threshold, the change in the direction of hand movement can be accurately quantified, and the action of simply passing the plate can be more effectively distinguished from the action of picking up food.

[0079] See Figure 3 This is a schematic diagram of a smart canteen processing system provided in an embodiment of the present invention. The smart canteen processing system includes: The acquisition module is used to acquire audio and video data corresponding to the target area that meets the monitoring conditions in real time; The first collaboration module is used to build a two-way collaboration mechanism. When a trigger condition is detected in the audio data, the user status of the video data in the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal. The second collaboration module is used to detect when the user status and table status in the video data are in a triggered state, and then, in combination with the second keyword of the audio data in the second acquisition interval, generate a service task and send it to the execution terminal.

[0080] Figure 3 The apparatus of the illustrated embodiment can be used to perform corresponding actions. Figure 2 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A smart canteen processing method, characterized in that, include: Real-time acquisition of audio and video data corresponding to target areas that meet monitoring conditions; A two-way collaborative mechanism is constructed. When a trigger condition is detected in the audio data, the user status of the video data within the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal. When the user status and table status in the video data are detected as being in a triggered state, a service task is generated and sent to the execution terminal by combining the second keyword of the audio data within the second acquisition interval.

2. The method according to claim 1, characterized in that, Real-time acquisition of audio and video data corresponding to the target area that meets the monitoring conditions, including: The monitoring conditions are met by determining that a user's dining area meets the monitoring criteria and is therefore identified as the target area. Control and monitoring equipment to collect video data of the target area in real time; Based on the user's location distribution in the video data, the directional sound pickup direction and pickup range of the sound pickup device are determined, and the sound pickup device is controlled to collect audio data in real time based on the directional sound pickup direction and pickup range.

3. The method according to claim 2, characterized in that, Based on the user location distribution in the video data, the directional sound pickup direction and pickup range of the audio pickup device are determined, including: The user's location distribution is determined based on the position coordinates of the person's outline in the video data; The coordinates of positions with a spacing less than a preset spacing are determined to be the same aggregation group. The device position coordinates of the pickup device are obtained, and the direction from the device position coordinates to the center coordinates of the aggregation group is determined as the directional pickup direction corresponding to the aggregation group. Based on the directional pickup direction, the angle is gradually expanded to both sides according to the angle increment until the expanded area completely covers all users in the aggregation combination. The range corresponding to the expanded area is determined as the pickup range corresponding to the directional pickup direction.

4. The method according to claim 3, characterized in that, Also includes: The average volume of the pickup device is obtained, and when the average volume is greater than or equal to a volume threshold, the pickup range is increased based on a preset angle.

5. The method according to claim 1, characterized in that, A two-way collaborative mechanism is established. When a trigger condition is detected in the audio data, the user status of the video data within the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal, including: Compare the first keyword of the audio data with the preset keyword. If the first keyword is the same as the preset keyword, it is determined that the audio data meets the triggering condition. Based on the current time, the video data within the first acquisition interval is obtained by tracing back a preset duration, and the user status in the video data is identified; When the user's state is a preset action, a service task is generated by combining the first keyword and sent to the execution terminal.

6. The method according to claim 1, characterized in that, When the user status and table status in the video data are detected as triggered, a service task is generated and sent to the execution terminal by combining the second keyword of the audio data within the second acquisition interval, including: When the user state is a preset action, the second keyword and the preset keyword of the audio data in the second collection interval are compared. When the second keyword is the same as the preset keyword, a service task is generated and sent to the execution terminal. The detection frequency for identifying the table status is adjusted based on the amount of food changes on each plate and the frequency of the user's picking actions. When the table status is such that the number of empty plates is greater than or equal to the empty plate threshold, a service task is generated and sent to the execution terminal. The triggering status includes the user status being a preset action and the number of empty plates being greater than or equal to the empty plate threshold.

7. The method according to claim 6, characterized in that, The detection frequency for recognizing the table status is adjusted based on the amount of food changes on each plate and the frequency of the user's picking-up actions, including: The initial coverage area of ​​each plate is obtained, the remaining coverage area of ​​the plate is recorded according to the initial recognition frequency, and the change in food is obtained according to the ratio of the difference between the initial coverage area and the remaining coverage area to the initial coverage area. If the change in the food exceeds a first threshold, the detection frequency is adjusted by a first increasing frequency; if the change in the food is less than the first threshold, the detection frequency is adjusted by a first decreasing frequency. The frequency of the user's gripping actions on each plate is obtained. If the gripping action frequency is greater than a second threshold, the detection frequency is adjusted by increasing the frequency by a second factor. If the gripping action frequency is less than the second threshold, the detection frequency is adjusted by decreasing the frequency by a second factor.

8. The method according to claim 7, characterized in that, Get the frequency of the user's picking up actions on each plate, including: Identify the user's hand trajectory and obtain the duration of the hand trajectory's stay in the plate area and the number of times the direction changes; When the dwell time exceeds the clamping duration threshold and / or the number of orientation changes exceeds the number threshold, record one clamping count, and obtain the clamping action frequency based on the number of clampings per unit time.

9. The method according to claim 8, characterized in that, The number of times the hand trajectory changes direction in the plate area is obtained, including: Obtain the angle between the tangents of the hand trajectory at the previous and next moments. When the angle between the tangents is greater than the angle threshold, increase the number of direction changes by 1.

10. A smart canteen processing system, characterized in that, include: The acquisition module is used to acquire audio and video data corresponding to the target area that meets the monitoring conditions in real time; The first collaboration module is used to build a two-way collaboration mechanism. When a trigger condition is detected in the audio data, the user status of the video data in the first acquisition interval is obtained, and a service task is generated by combining the first keyword of the audio data and sent to the execution terminal. The second collaboration module is used to detect when the user status and table status in the video data are in a triggered state, and then, in conjunction with the second keyword of the audio data within the second acquisition interval, generate a service task and send it to the execution terminal.