A method and system for responding to dining needs based on voice interaction
By extracting user speech in the restaurant environment using microphone arrays and dynamic filtering masking technology, and combining it with a service conflict rule base for semantic completion and scheduling optimization, the problems of noise interference and uneven task allocation during peak dining hours are solved, thereby improving the accuracy of speech recognition and the response efficiency of restaurant services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNXIANG (SHANGHAI) CATERING MANAGEMENT CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
In restaurant environments during peak dining hours, existing technologies struggle to effectively suppress environmental noise interference, leading to a decrease in speech recognition rates. Furthermore, the service scheduling scheme lacks dynamic detection of the real-time status of waiters and service requests, resulting in uneven task allocation and response delays.
The system acquires the user's sound source coordinates and the frequency domain features of ambient noise through a microphone array, constructs a dynamic filtering mask for speech enhancement, performs semantic completion and scheduling optimization in conjunction with a pre-set service conflict rule base, generates a service path allocation weight table, and monitors the load coefficient in real time to trigger secondary compensation decisions.
It enables high-fidelity extraction of user voice in noisy environments, ensuring the accuracy and clarity of service instructions, and optimizes the task allocation of multiple waiters, avoiding uneven task distribution and response delays, thereby improving the overall operational efficiency of restaurant services.
Smart Images

Figure CN122088950A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction technology, and in particular to a method and system for responding to dining needs based on voice interaction. Background Technology
[0002] In existing technologies, during peak dining hours, the ambient noise spectrum is complex and dynamically changing. Current methods often only employ general noise reduction algorithms or filters with fixed thresholds, lacking dynamic analysis and targeted suppression of the frequency domain characteristics of ambient noise. This makes it difficult to stably extract clear user speech in strong noise backgrounds, resulting in a significant decrease in speech recognition rate and frequent false triggers and command misunderstandings. Furthermore, most systems only perform speech-to-text conversion, lacking a deep understanding and structuring ability for users' colloquial and fragmented expressions. Due to the lack of a knowledge rule base for the service scenario, the system cannot effectively complete omitted information, correct ambiguities, or identify logical conflicts between different requests, resulting in incomplete, semantically ambiguous, or self-contradictory service instructions that cannot directly drive accurate service responses.
[0003] Furthermore, existing service scheduling schemes are mostly based on simple polling or static priority rules, assigning identified instructions to idle waiters. This method lacks real-time detection and dynamic balancing of multi-dimensional conflicts between the real-time status of waiters and the inherent requirements of service requests. This easily leads to uneven task allocation, with some waiters overloaded and experiencing delayed responses, while others are underutilized. Simultaneously, the system lacks continuous tracking of waiter load status after task execution, as well as closed-loop monitoring of the synergy between allocation strategies and actual execution load, making it impossible to perform dynamic secondary compensation and scheduling optimization when system efficiency becomes unbalanced. Therefore, there is an urgent need to develop a dining demand response method and system that deeply integrates environmental awareness, accurate semantic parsing, and dynamic intelligent scheduling to solve the problems of inaccurate semantic understanding and service element extraction, as well as the suboptimal dynamic task allocation among multiple waiters, thereby improving the accuracy, timeliness, and overall operational efficiency of restaurant service responses. Summary of the Invention
[0004] This invention provides a method and system for responding to dining needs based on voice interaction, the main purpose of which is to address the problems raised in the background section above.
[0005] To achieve the above objectives, the present invention provides a method for responding to dining needs based on voice interaction, comprising: S1: Obtain the coordinates of user sound sources and the frequency domain characteristics of ambient noise to meet dining needs; S2: Construct a dynamic filtering mask for the dining needs based on the frequency domain characteristics of the environmental noise, and extract the semantic fragments of the dining needs by combining the user's sound source coordinates; S3: Based on a preset service conflict rule base, the semantic fragment of the demand is structurally completed to generate a set of service elements for the dining demand; S4: Perform conflict detection on the set of service elements and the status of the waiter, and generate a service path allocation weight table for the dining demand; S5: Based on the service path weight allocation table, drive the interactive terminal of the dining demand to output service guidance instructions, and at the same time update the dynamic load coefficient of the waiter status. S6: Monitor the synergy between the dynamic load coefficient and the service path allocation weight table in real time, and trigger secondary demand response compensation decisions.
[0006] Preferably, the step of obtaining the coordinates of the user's sound source and the frequency domain characteristics of ambient noise to determine dining needs includes: The mixed audio signal of the restaurant was acquired by a microphone array deployed in the restaurant; The mixed audio signal is preprocessed and framed, and a fast Fourier transform is performed to obtain the frequency domain representation of the mixed audio signal; Based on the frequency domain representation, the environmental noise frequency domain features of the mixed audio signal are extracted; Perform generalized cross-correlation time delay estimation on the mixed audio signal to calculate the microphone time difference for the dining requirement; Based on the microphone time difference and the geometry of the microphone array, the coordinates of the user's sound source for the dining request are calculated.
[0007] Preferably, the step of constructing a dynamic filtering mask for the dining needs based on the frequency domain features of the environmental noise, and extracting the semantic fragments of the dining needs by combining the user's sound source coordinates, includes: Based on the frequency domain characteristics of the environmental noise, the environmental noise energy distribution of the dining demand is calculated, and an adaptive filtering threshold is determined based on a preset energy threshold. Based on the adaptive filtering threshold, a dynamic filtering mask for the dining requirements is constructed. Based on the user's sound source coordinates, the speech signal of the microphone array is spatially filtered to generate the user's enhanced speech signal for the dining needs; The user voice enhancement signal is filtered based on the dynamic filtering mask to obtain a clean user voice signal for the dining requirement. Endpoint detection is performed on the clean user voice signal to segment out the effective voice segment of the dining requirement, and speech recognition is performed on the voice segment to generate the semantic fragment of the dining requirement.
[0008] Preferably, the preset service conflict rule base includes colloquial expression grammar rules, service item conflict rules, and semantic slot templates.
[0009] Preferably, the step of performing structural completion on the semantic fragment of the demand based on a preset service conflict rule base to generate a set of service elements for the dining demand includes: Parse the initial user intent of the semantic fragment of the requirement; Based on the colloquial expression grammar rules, the initial user intent is grammatically corrected and completed to generate a structured expression of the dining needs; Based on the service item conflict rules, logical conflict items in the structured intent representation are identified and resolved to generate an advanced structured intent representation of the dining needs. Based on the advanced structured intent description, fill in the set of service elements for the dining needs.
[0010] Preferably, the step of performing conflict detection on the set of service elements and the waiter status to generate a service path allocation weight table for the dining demand includes: Match the task requirements in the set of service elements with the status of the waiter; Calculate the expected path length, skill matching degree, and load impact for each server; The expected path length, the skill matching degree, and the load impact degree are integrated into a multi-dimensional conflict index for the dining demand. Based on the preset conflict resolution strategy, the multi-dimensional conflict indicators are comprehensively evaluated to generate the service path allocation weight for each service worker. The allocation weights of each waiter are integrated to generate a service path allocation weight table for the dining needs.
[0011] Preferably, the step of driving the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and simultaneously updating the dynamic load coefficient of the waiter status, includes: Parse the service path allocation weight table to obtain the waiter identifier, service path sequence, and expected service time for the dining demand; Based on the waiter's identifier and the service path sequence, a service guidance instruction for the dining request is generated, and the interactive terminal for the dining request is driven to output the service guidance instruction. Based on the expected path length, skill matching degree, and expected service time, calculate the additional load value for each waiter; Update the dynamic load coefficient of the waiter status based on the newly added load value.
[0012] Preferably, the real-time monitoring of the synergy between the dynamic load coefficient and the service path allocation weight table, triggering a secondary demand response compensation decision, includes: Based on the dynamic load coefficient and the service path allocation weight table, the system coordination deviation of the dining demand is calculated; When the system coordination deviation exceeds a preset first threshold, a secondary demand response compensation decision is triggered. The secondary demand response compensation decision includes initiating path reallocation for the dining demand and requesting auxiliary scheduling for the dining demand.
[0013] Preferably, the formula for calculating the system's cooperative deviation is: ; in, The system's cooperative deviation is... Index for waiters, This represents the total number of service staff currently in service. waiter Real-time dynamic load factor, waiter Cumulative path weights Let be the standard complexity constant of the stated dining requirements. It is an exponential function. As an exponential adjustment factor, waiter The degree to which the skill attributes match the requirements of the current task. This refers to real-time environmental interference factors.
[0014] A voice-interactive dining demand response system, the system comprising: The environmental perception module is used to obtain the coordinates of the user's sound source and the frequency domain characteristics of environmental noise when dining. The speech processing module is used to construct a dynamic filtering mask for the dining needs based on the frequency domain characteristics of the ambient noise, and extract the semantic segments of the dining needs by combining the user's sound source coordinates. The semantic parsing module is used to complete the structure of the semantic fragments of the demand based on a preset service conflict rule base, and generate a set of service elements for the dining demand. The scheduling optimization module is used to perform conflict detection on the set of service elements and the status of the waiters, and generate a service path allocation weight table for the dining demand. The instruction execution module is used to drive the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and at the same time update the dynamic load coefficient of the waiter status. The collaborative monitoring module is used to monitor the synergy between the dynamic load coefficient and the service path allocation weight table in real time, and trigger secondary demand response compensation decisions.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention dynamically analyzes the frequency domain characteristics of environmental noise and constructs an adaptive filter mask, combining this with sound source spatial localization for joint noise reduction. The system effectively suppresses background noise interference and extracts user speech with high fidelity from noisy sound fields, significantly improving the first-pass pass rate and accuracy of speech recognition. Furthermore, through a pre-set service knowledge base containing grammatical and conflict rules, the system can intelligently complete, correct, and resolve logical conflicts in recognized colloquial and fragmented semantics, transforming ambiguous user expressions into standard service instructions that are structurally complete, clearly defined, and logically consistent. This fundamentally solves problems such as misinterpretation of service requests, missing elements, or contradictory instructions caused by environmental interference and arbitrary expression, ensuring an accurate foundation for subsequent service scheduling.
[0016] This invention achieves dynamic optimization of task allocation for multiple service staff and continuous improvement of system operating efficiency. When allocating tasks, the system does not simply rely on idle status, but comprehensively considers multi-dimensional conflict indicators such as service elements, real-time location of service staff, skill matching degree, and dynamic load coefficient to generate an optimal service path allocation weight table, achieving the best match between resources and demand. Simultaneously, through a closed-loop feedback mechanism, the system updates the service staff load status in real time after instruction execution and continuously monitors the degree of deviation between the actual load and the allocation strategy. Once a system imbalance is detected, a secondary compensation decision is immediately triggered, thereby dynamically adjusting the scheduling strategy. This effectively avoids the uneven task allocation and response delay problems caused by traditional polling or static scheduling, achieving dynamic balancing of service load and optimization of overall operational efficiency. Attached Figure Description
[0017] Figure 1 A flowchart illustrating a voice-interactive-based method for responding to dining needs, as provided in an embodiment of the present invention. Figure 2 A functional block diagram of a voice-interactive dining demand response system provided in an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a voice-interactive method for responding to dining needs. The executing entity of this voice-interactive method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the voice-interactive method for responding to dining needs can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a voice-interactive-based method for responding to dining needs according to an embodiment of the present invention. In this embodiment, the voice-interactive-based method for responding to dining needs includes: In this embodiment of the invention, the coordinates of the user's sound source and the frequency domain characteristics of environmental noise are obtained; The acquisition of user sound source coordinates and ambient noise frequency domain characteristics for dining needs includes: The mixed audio signal of the restaurant was acquired by a microphone array deployed in the restaurant; The mixed audio signal is preprocessed and framed, and a fast Fourier transform is performed to obtain the frequency domain representation of the mixed audio signal; Based on the frequency domain representation, the environmental noise frequency domain features of the mixed audio signal are extracted; Perform generalized cross-correlation time delay estimation on the mixed audio signal to calculate the microphone time difference for the dining requirement; Based on the microphone time difference and the geometry of the microphone array, the coordinates of the user's sound source for the dining request are calculated.
[0021] Specifically, dining needs refer to service requests made by customers via voice in a restaurant environment, such as ordering food, adding dishes, paying the bill, and calling for service.
[0022] Specifically, the user's voice source coordinates refer to the spatial coordinates of the location of the customer making the voice request within the physical space of the restaurant.
[0023] Specifically, the frequency domain characteristics of environmental noise refer to the frequency characteristics of interference sounds that are not user speech in the background environment of a restaurant.
[0024] Specifically, mixed audio signals refer to the raw audio data collected by a microphone array deployed in the restaurant.
[0025] Specifically, a microphone array refers to a collection of multiple microphones arranged in a specific geometric structure within a restaurant space.
[0026] Specifically, frequency domain representation refers to the signal representation with frequency as the independent variable obtained by mathematically transforming the mixed audio signal in the time domain.
[0027] Specifically, microphone time difference refers to the time difference between the arrival of a sound signal from the same sound source at different microphone units in a microphone array. This difference arises from the different propagation distances of sound waves and is a core parameter for calculating the direction and distance of the sound source.
[0028] Specifically, geometric structure refers to the relative spatial relationships of the microphone units in a microphone array, such as spacing and arrangement shape.
[0029] Furthermore, microphone arrays distributed across different locations within the restaurant begin to operate synchronously. Each microphone unit continuously captures air vibrations at its location, converting the physical sound signals into continuous electrical signals. These electrical signals from multiple spatial points are simultaneously captured by a data acquisition card or audio interface and input into the system in real time, forming a multi-channel mixed audio signal that includes various components such as the target user's voice, conversations with other customers, background music, and equipment operating sounds.
[0030] Furthermore, the continuous multi-channel mixed audio signal is first pre-processed and framed. The system cuts the continuous signal stream into short, stationary segments according to a preset frame length and frame shift. A window function is typically applied to each frame to reduce spectral leakage.
[0031] Then, a Fast Fourier Transform (FFT) is performed on each frame of the signal. This transform converts each frame of the time-domain signal independently to the frequency domain, calculating the complex spectrum of the signal at each frequency point within that time period. By organizing the spectra of all frames in chronological order, a complete frequency domain representation of the mixed audio signal is obtained, namely a time-spectrum diagram, which clearly shows the distribution of signal energy over time and frequency.
[0032] Furthermore, by comparing with a preset background noise model for quiet periods, or by using sound activity detection technology, frames that may contain speech are initially distinguished from frames containing pure noise. In frames determined to be pure noise or speech intervals, the system statistically analyzes their spectral characteristics: calculating the average energy of each frequency band and identifying the frequency bands with higher energy as the main noise bands; analyzing the stability of the spectrum to distinguish between steady-state noise with slow energy changes over time and non-steady-state noise with sudden energy changes. Integrating these analyzed noise energy distribution patterns and spectral characteristic patterns extracts the specific environmental noise frequency domain characteristics for the current moment.
[0033] Furthermore, to locate the sound source, the system selects one or all pairs of microphone channels from the microphone array. For the mixed audio signals acquired by each pair of channels, the system performs generalized cross-correlation time delay estimation. The core of this operation is to calculate the cross-correlation function of the two signals at different time offsets and enhance its robustness to noise and reverberation through pre-filtering, thereby obtaining a generalized cross-correlation function with a sharp peak. The peak position of this function is searched, and the time offset corresponding to this position is the microphone time difference between the sound and the two microphones. By processing multiple pairs of microphones, a set of time difference observations can be obtained.
[0034] Furthermore, the system uses the set of microphone time difference observations obtained in the previous step, along with the pre-determined geometric parameters of the microphone array, as known inputs. Based on the sound wave propagation model, a set of hyperbolic equations is constructed. By solving this set of equations, the most probable sound source location point that satisfies all observation time differences can be calculated in three-dimensional space or a two-dimensional plane. The coordinates of this location point are the user's sound source coordinates.
[0035] In summary, acquiring mixed audio signals enables the digital capture of the complete acoustic scene, including user needs, providing the raw data foundation for all subsequent processing. Preprocessing, framing, and performing Fast Fourier Transform to obtain the frequency domain representation transforms the time-domain signal into a domain more suitable for analyzing spectral characteristics. This transformation is an essential step and foundation for subsequent advanced processing such as noise feature analysis and frequency domain filtering.
[0036] In summary, extracting the frequency domain features of ambient noise directly serves the construction of a dynamic filter mask in subsequent steps, enabling the system to perform intelligent and adaptive noise reduction instead of using a fixed filter, thus maintaining good speech capture capabilities in changing environments. Generalized cross-correlation delay estimation to calculate the microphone time difference utilizes the physical properties of sound propagation to extract key observations for spatial positioning from multi-channel signals.
[0037] In summary, the calculation of user sound source coordinates based on microphone time difference and geometric structure provides beamforming direction for the voice processing module, enabling spatially selective sound pickup and further suppressing noise from non-target directions; it also provides the scheduling optimization module with location information of the demand initiation, which is one of the core inputs for optimizing the waiter delivery path and calculating the expected path length.
[0038] In this embodiment of the invention, a dynamic filtering mask for the dining needs is constructed based on the frequency domain characteristics of the environmental noise, and the semantic fragments of the dining needs are extracted by combining the user's sound source coordinates. The process of constructing a dynamic filtering mask for dining needs based on the frequency domain features of the environmental noise, and extracting semantic fragments of dining needs by combining the user's sound source coordinates, includes: Based on the frequency domain characteristics of the environmental noise, the environmental noise energy distribution of the dining demand is calculated, and an adaptive filtering threshold is determined based on a preset energy threshold. Based on the adaptive filtering threshold, a dynamic filtering mask for the dining requirements is constructed. Based on the user's sound source coordinates, the speech signal of the microphone array is spatially filtered to generate the user's enhanced speech signal for the dining needs; The user voice enhancement signal is filtered based on the dynamic filtering mask to obtain a clean user voice signal for the dining requirement. Endpoint detection is performed on the clean user voice signal to segment out the effective voice segment of the dining requirement, and speech recognition is performed on the voice segment to generate the semantic fragment of the dining requirement.
[0039] Specifically, environmental noise energy distribution refers to the arrangement of noise intensity in different frequency ranges, which is further quantified based on the frequency domain characteristic analysis of environmental noise.
[0040] Specifically, the adaptive filtering threshold refers to a numerical limit that is dynamically adjusted based on the real-time calculated distribution of ambient noise energy. The system compares the noise energy with a preset energy threshold to determine a filtering threshold that is most suitable for the current noise level.
[0041] Specifically, a dynamic filtering mask refers to a matrix or function that is generated in real time based on an adaptive filtering threshold and is used to weight signals in the frequency domain.
[0042] Specifically, spatial filtering refers to an operation that processes audio signals using the coordinates of the user's sound source and the geometric characteristics of the microphone array.
[0043] Specifically, the user voice enhancement signal refers to the single-channel output signal generated after spatial filtering of the original multi-channel signal acquired by the microphone array.
[0044] Specifically, a clean user speech signal refers to the signal obtained by applying a dynamic filtering mask to the user speech enhancement signal and then performing further frequency domain filtering.
[0045] Specifically, endpoint detection refers to the operation of analyzing continuous, clean user voice signals and automatically segmenting the start and end points that contain valid voice parts.
[0046] Specifically, a valid speech segment refers to an audio fragment containing the user's complete speech request that is precisely segmented from a clean user speech signal through endpoint detection.
[0047] Specifically, the semantic fragment of demand refers to the text form of the result generated after speech recognition of a valid speech segment.
[0048] Furthermore, firstly, the entire audible frequency range is divided into multiple sub-bands. For each sub-band, the system calculates the average or statistical value of the noise energy within the current analysis time period, thereby calculating the environmental noise energy distribution map across the entire frequency spectrum.
[0049] Next, the system compares this distribution map with a preset energy threshold curve. The preset energy threshold is typically a baseline derived from extensive experimental data, indicating the energy level of typical noise at different frequencies. Based on the degree to which real-time noise energy exceeds or falls below this baseline, the system dynamically determines a new adaptive filtering threshold line that fits the current noise environment. For example, if low-frequency noise is abnormally prominent, the threshold in that frequency band will be increased accordingly for stronger suppression.
[0050] Furthermore, the system constructs a matrix with the same dimension as the signal spectrum—that is, a dynamic filtering mask—based on the determined adaptive filtering threshold.
[0051] For each frequency point in the spectrum, its energy value is compared with the adaptive filtering threshold value at that frequency point. If the energy at that point is significantly lower than the threshold, a very small weight value close to 0 is assigned to the corresponding position in the mask, indicating that the frequency point should be significantly suppressed; if the energy is close to or higher than the threshold, a weight value close to 1 or adjusted according to the signal-to-noise ratio is assigned to the mask, indicating that the frequency point should be preserved or appropriately enhanced. This weight matrix is the dynamic filtering mask used for the next filtering step.
[0052] Furthermore, the precise coordinates of the user's sound source and the position coordinates of each unit in the microphone array are obtained. Based on this spatial geometric information, the system calculates the expected delay of sound from the target sound source direction to each microphone.
[0053] Subsequently, it processes the multi-channel raw speech signals acquired by the microphone array in real time, applying corresponding time delay compensation and weight adjustment to each channel signal. This ensures that the speech signals from the target direction are aligned and superimposed in phase across all channels, thus enhancing them. Meanwhile, interference signals from other directions cancel each other out or weaken each other because they cannot be perfectly aligned. After this processing, a user speech enhancement signal dominated by the target speech is generated.
[0054] Furthermore, the time-domain enhanced user speech signal is transformed into the frequency domain to obtain its spectrum. Then, the constructed dynamic filtering mask is multiplied onto this spectrum. This means that the amplitude of each frequency component in the spectrum is scaled according to the weight value at the corresponding position in the mask: the amplitude of frequencies identified as noise is significantly reduced, while frequencies identified as potentially containing speech are largely preserved. After completing the frequency domain weighting, the system inversely transforms the processed spectrum back into the time domain. After this series of filtering processes, the final time-domain signal is the clean user speech signal, in which the spectral components of background noise have been significantly and specifically filtered out.
[0055] Furthermore, endpoint detection is performed on the clean user speech signal. Typically, by calculating features such as short-time energy and zero-crossing rate, and setting dynamic thresholds, the start and end times of the speech are accurately determined, thus segmenting a valid speech segment containing only the user's voice. Next, the system calls the speech recognition engine to decode and analyze this valid speech segment. The recognition engine maps audio features to text units and generates the most probable text sequence by searching for the optimal path.
[0056] In summary, calculating the environmental noise energy distribution and determining the adaptive filtering threshold transforms the qualitative perception of noise into a quantitative and actionable decision criterion. Constructing a dynamic filtering mask is a specific noise reduction tool generated based on these decision criteria. This mask is the soul of frequency domain filtering; it directly determines which frequency components are preserved and which are suppressed, serving as the core carrier for achieving high-precision, adaptive noise reduction.
[0057] In summary, spatial filtering is used to generate enhanced user speech signals. By utilizing sound source location information, the voices of target customers are picked up first, and speech and spatially dispersed noise are initially separated from the source, providing a cleaner starting point for subsequent processing.
[0058] In summary, applying dynamic filtering masks to obtain clean user speech signals is a second round of refinement in the frequency dimension, building upon spatial enhancement. Endpoint detection segmenting of effective speech segments and speech recognition generating required semantic fragments are crucial for the transformation from clear audio to explicit textual instructions. Endpoint detection ensures the accuracy of the processed object and avoids interference from silent segments; speech recognition enables semantic coherence in human-computer interaction, converting user speech requests into standardized text that the system can subsequently parse and process, laying an accurate information input foundation for the entire service response chain.
[0059] In this embodiment of the invention, the semantic fragment of the demand is structurally completed based on a preset service conflict rule base to generate a set of service elements for the dining demand. The preset service conflict rule base includes colloquial expression grammar rules, service item conflict rules, and semantic slot templates.
[0060] The system performs structural completion on the semantic fragments of the demand based on a preset service conflict rule base, generating a set of service elements for the dining demand, including: Parse the initial user intent of the semantic fragment of the requirement; Based on the colloquial expression grammar rules, the initial user intent is grammatically corrected and completed to generate a structured expression of the dining needs; Based on the service item conflict rules, logical conflict items in the structured intent representation are identified and resolved to generate an advanced structured intent representation of the dining needs. Based on the advanced structured intent description, fill in the set of service elements for the dining needs.
[0061] Specifically, the pre-defined service conflict rule base refers to a knowledge database that is pre-established and stored before system deployment.
[0062] Specifically, the colloquial expression grammar rules are a subset of a pre-defined service conflict rule base. It defines the mapping between non-standard expressions commonly used by customers in everyday conversations and standard service request phrases.
[0063] Specifically, service conflict rules are another subset of the pre-defined service conflict rule base. They define the possible relationships between different service requests that are logically mutually exclusive or have strict requirements on the execution order.
[0064] Specifically, semantic slot templates are another subset of the pre-defined service conflict rule base. They define the standardized information field structure that different service types must include.
[0065] Specifically, initial user intent refers to the summary of user purpose parsed by the system after the system has made the most preliminary and direct understanding of the semantic fragments of the demand.
[0066] Specifically, structured intent representation refers to a more complete and standardized form of intent description generated by standardizing the initial user intent according to the rules of colloquial expression grammar and filling in missing information.
[0067] Specifically, advanced structured intent representation refers to the final version of the intent representation without logical contradictions, generated after applying service item conflict rules to the structured intent representation and resolving any identified contradictions.
[0068] Specifically, the service element set refers to a structured data set containing complete service instruction information generated after the advanced structured intent expression is filled into the corresponding semantic slot template.
[0069] Further, it receives semantic fragments of the user's needs from the previous module. First, it performs a shallow analysis of the text using keyword matching and basic natural language processing techniques. For example, it identifies verbs such as "point," "want," and "add," which may correspond to the intent to "place an order"; words such as "checkout" and "pay the bill," which correspond to the intent to "pay"; and words such as "slow" and "not yet," which may correspond to the intent to "urge" or "query." The result of this step is a rough, framework-like initial user intent, indicating the general direction of the need, but the details are vague.
[0070] Furthermore, the initial user intent is matched and mapped against colloquial expression grammar rules in a pre-defined service conflict rule base. This rule base acts like a grammar correction manual, guiding the system on how to complete missing components and correct ambiguous expressions.
[0071] Furthermore, for example, given the initial user intent to urge and the fragment "boiled beef," the rule specifies that the "urge" intent must be associated with a specific "dish name." The system will then fill in "boiled beef." For the fragment "this," the rule might indicate that the referent needs to be determined by combining the dialogue context or visual information.
[0072] By applying these rules, the system can reconstruct the original fragment in a standardized manner, generating a more complete structured expression of intent that conforms to the standard service language format, such as urging for the dish: boiled beef.
[0073] Furthermore, the system compares the structured intent representation with the service item conflict rules to check whether it has any logical contradiction with the current system state or other valid requests.
[0074] The identification process includes: checking whether the ordered dish exists on the menu, whether the same dish has been ordered before but a cancellation request has been made, and whether a new dish has been requested at the same time as the bill is being processed. Once a conflict is identified, the system initiates a resolution strategy.
[0075] The conflict resolution methods may include: overriding the old instruction with the latest one, rejecting the conflicting instruction and prompting the user, or marking the conflict request as requiring manual intervention. After completing the conflict checking and resolution, the system outputs a logically consistent advanced structured intent statement.
[0076] Furthermore, based on the service type determined by the advanced structured intent statement, a corresponding standard template is selected from the semantic slot template library. Then, each parameter explicitly defined in the advanced structured intent statement is filled into the corresponding slot of the template. For certain elements required by the template but not directly provided in the statement, the system will automatically complete them according to preset rules. Finally, a set of service elements that has all necessary slots filled and can be seamlessly read and processed by the machine is generated.
[0077] In summary, parsing the initial user intent and translating it from natural language into a preliminary, machine-operable intent label is the trigger for initiating subsequent deep understanding processes. Grammatical correction and completion based on colloquial expression rules translates incomplete and arbitrary spoken language into standard, unambiguous service terms, resolving comprehension barriers caused by differences in expression styles in human-computer interaction. This is a crucial step in ensuring the accuracy of intent.
[0078] In summary, the logic conflict resolution based on service item conflict rules can identify and handle potentially contradictory instructions in advance before service execution, effectively avoiding system chaos or execution errors caused by instruction conflicts and ensuring the rationality and orderliness of the service process.
[0079] In summary, by filling the service element set with advanced structured intent representations, the fully understood and validated user intent is encapsulated into a standardized service ticket with a unified format and complete information. This service element set serves as a bridge connecting the semantic understanding layer and the service scheduling and execution layer. It transforms the user's unstructured voice request into a structured task object containing all execution details, providing precise input for subsequent intelligent scheduling and resource allocation.
[0080] In this embodiment of the invention, conflict detection is performed on the set of service elements and the status of the waiter to generate a service path allocation weight table for the dining demand. The step of performing conflict detection on the set of service elements and the status of the waiters to generate a service path allocation weight table for the dining demand includes: Match the task requirements in the set of service elements with the status of the waiter; Calculate the expected path length, skill matching degree, and load impact for each server; The expected path length, the skill matching degree, and the load impact degree are integrated into a multi-dimensional conflict index for the dining demand. Based on the preset conflict resolution strategy, the multi-dimensional conflict indicators are comprehensively evaluated to generate the service path allocation weight for each service worker. The allocation weights of each waiter are integrated to generate a service path allocation weight table for the dining needs.
[0081] Specifically, the waiter status refers to the dynamic and quantifiable set of work status information for each waiter at the current point in time in the system.
[0082] Specifically, status matching refers to the process of comparing and associating the specific task requirements in the set of service elements with the various information in the status of the service personnel one by one.
[0083] Specifically, the expected path length refers to the estimated total distance that a specific waiter would need to walk from their real-time geographical location to the target table, possible pick-up points, and finally to provide the service if the task of fulfilling the dining request were assigned to that waiter.
[0084] Specifically, skill matching refers to the degree to which a waiter's skills and expertise match the abilities required for the task within the set of service elements.
[0085] Specifically, load impact refers to the estimated impact on a waiter's existing current task load if a new task is assigned to them.
[0086] Specifically, the multi-dimensional conflict index refers to integrating the evaluation results of three different aspects—expected path length, skill matching degree, and load impact degree—calculated for each waiter into a comprehensive evaluation vector or dataset.
[0087] Specifically, the pre-set conflict resolution strategy refers to a set of rules or algorithms pre-defined by the system to guide how to weigh and reconcile the contradictions among various factors in multi-dimensional conflict indicators.
[0088] Specifically, the service path allocation weight refers to a numerical result generated after applying a preset conflict resolution strategy to comprehensively evaluate the multi-dimensional conflict indicators of a certain waiter.
[0089] Specifically, the service path allocation weight table refers to the global view formed by integrating the allocation weights of each service worker.
[0090] Furthermore, the system first reads the set of service elements to clarify the core requirements of the task. Simultaneously, it obtains the real-time status of all online service staff. Next, it performs a status matching process: for example, it calculates the spatial relationship between the task target table A12 and the real-time geographical location of each service staff member; it verifies the task requirement of using a tablet for checkout against each service staff member's skill and expertise list; and it performs a pre-assessment by overlaying the overall complexity of the task with each service staff member's current workload.
[0091] Furthermore, for each waiter or those initially matched and screened, the system calculates three key indicators: Expected path length: Based on the restaurant's digital map or topology model, combined with the waiter's real-time geographical location and the target location and possible intermediate points of the task, the system calculates the optimal or typical walking path through the path planning algorithm and calculates its total length.
[0092] Skills Matching: The system compares the skill requirement tags in the service element set with the skill expertise tags in the waiter's profile. The matching score can be assigned different quantitative scores based on levels such as perfect match, partial match, and no match.
[0093] Load Impact: The system analyzes the waiter's current task load, then adds the estimated time for new tasks to assess whether the total load is close to or exceeds their individual capacity threshold. Load impact can be measured by the percentage increase in load or the estimated delay in response time for subsequent tasks.
[0094] Furthermore, the system creates a unified evaluation framework for each server, integrating the three values or levels calculated separately—expected path length, skill matching degree, and load impact degree—into a preset format.
[0095] Furthermore, for example, a triple can be formed such as: path length, skill matching degree, and load impact degree. This triple is a multi-dimensional conflict indicator for the waiter and this specific task, which comprehensively reflects the potential conflict in the three dimensions of efficiency, quality, and load.
[0096] Furthermore, the system invokes a pre-defined conflict resolution strategy. This strategy defines how to normalize and weighted aggregate multi-dimensional metrics such as path length, skill matching degree, and load impact.
[0097] Furthermore, for example, the strategy might stipulate: first, convert path length into time cost, skill matching degree into quality benefits, and load impact degree into risk coefficient; then, assign appropriate weights to costs, benefits, and risks based on the current operating model; finally, perform a comprehensive evaluation through an aggregation function to generate a single, comparable service path weight score. The higher the score, the smaller the overall conflict and the higher the suitability.
[0098] Furthermore, the system collects the identity identifiers of all evaluated waiters and their corresponding service path allocation weight scores. It then integrates these waiter-weight data pairs, typically sorting them from highest to lowest weight score, and organizes them into a well-structured table or list. This service path allocation weight table provides clear, data-driven task allocation suggestions for the interactive terminal.
[0099] In summary, the multi-dimensional conflict index standardizes and organizes evaluation results of different dimensions and properties, laying a data structure foundation for the next step of scientific and fair comprehensive comparison.
[0100] In summary, the comprehensive evaluation and weighting based on conflict resolution strategies transforms management strategies into mathematical rules, intelligently weighs multidimensional conflicts, and ultimately outputs a quantitative score representing comprehensive priority, achieving a key transformation from multidimensional assessment to a single decision-making basis.
[0101] In summary, the service path allocation weight table visually displays the priority ranking of all candidates, serving as a direct basis for the system to automate task dispatch or provide decision support for managers. It signifies the completion of the conflict detection and optimization process, seamlessly integrating the processed requirements into the execution phase.
[0102] In this embodiment of the invention, the interactive terminal for the dining demand is driven to output service guidance instructions according to the service path allocation weight table, and the dynamic load coefficient of the waiter status is updated at the same time. The step of driving the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and simultaneously updating the dynamic load coefficient of the waiter status, includes: Parse the service path allocation weight table to obtain the waiter identifier, service path sequence, and expected service time for the dining demand; Based on the waiter's identifier and the service path sequence, a service guidance instruction for the dining request is generated, and the interactive terminal for the dining request is driven to output the service guidance instruction. Based on the expected path length, skill matching degree, and expected service time, calculate the additional load value for each waiter; Update the dynamic load coefficient of the waiter status based on the newly added load value.
[0103] Specifically, the waiter identifier refers to the unique identification code that represents a particular waiter in the system, such as an employee number or a unique ID assigned by the system.
[0104] Specifically, the service path sequence refers to the orderly list of action steps planned for the selected waiter to fulfill the current dining needs.
[0105] Specifically, the expected service time refers to the total time estimated by the system for a waiter to complete all the tasks covered in the entire service path sequence.
[0106] Specifically, service guidance instructions refer to explicit operational instructions generated by the system to guide waiters in performing tasks. They typically include specific actions such as delivery, introduction, operation, target objects such as dishes, table numbers, and route directions, and are presented in text, voice, or graphic form.
[0107] Specifically, an interactive terminal refers to an electronic device carried by a waiter or permanently equipped in their work area for receiving system instructions.
[0108] Specifically, the additional workload value refers to the quantitative value that is expected to be calculated and added to a server's existing workload due to the assignment of current dining demand tasks to a server.
[0109] Specifically, the dynamic load coefficient is a comprehensive quantitative indicator that reflects the overall workload and stress level of waiters in real time.
[0110] Further, the system reads the generated service path allocation weight table. First, it parses the table structure, locates and retrieves the waiter identifier corresponding to the record with the highest weight, and determines the final task executor. Next, based on the waiter's location and task requirements, the system retrieves or plans a customized service path sequence for the task from the backend database. Simultaneously, based on data such as the path sequence length and the historical average time for the task type, the system obtains or calculates the expected service time to complete this sequence. These three pieces of information together constitute the core instruction content for task dispatch.
[0111] Furthermore, using waiter identification and service path sequence as input, specific and operable instructions are generated.
[0112] Furthermore, the instructions will be personalized. For example, "Wang Wu, corresponding to the waiter's label, please go to the No. 3 preparation station to get the 'boiled beef' and deliver it to table A12. Please enjoy your meal."
[0113] Then, the system locates the interactive terminal that the waiter is bound to or closest to based on the waiter's identifier, sends the generated instruction text, voice package or navigation map to the terminal, and drives the terminal to output through screen display, voice broadcast or vibration prompts to ensure that the instruction is effectively received by the target waiter.
[0114] Furthermore, for the waiters assigned tasks, the system initiates a load calculation process. This is based on three key parameters: 1) The expected path length for the waiter to perform this task; 2) The waiter's skill level in this task. The lower the skill level, the more extra preparation or psychological burden may be required. 3) Expected service time of the task.
[0115] The three parameters are weighted and combined using an internally pre-set load calculation model to calculate a value representing the additional workload brought by this task, i.e., the additional load value. For example, tasks with long paths, long durations, or skill mismatches will have a higher additional load value.
[0116] Furthermore, the system reads the waiter's current dynamic load coefficient from memory or the database. Then, it adds the newly calculated load value to the waiter's current dynamic load coefficient. The result is the updated total load level for the waiter after accepting the new task. The system immediately writes this new value back to the waiter's status record, completing the real-time update of the dynamic load coefficient. This updated coefficient will immediately affect the system's judgment when detecting conflicts with subsequent new requests.
[0117] In summary, generating and driving terminal output service guidance instructions transforms the digitized path sequence into personalized operation guidelines that service staff can directly understand and follow. Through interactive terminals, accurate and timely instruction issuance is achieved, ensuring that the optimized scheduling scheme is accurately implemented.
[0118] In summary, the calculation of new workload values not only considers time, but also takes into account the impact of physical strength and skill complexity, making the measurement of waiters' workload more precise and humane, and providing key data input for maintaining the overall healthy operation of the system.
[0119] In summary, the dynamic load factor ensures that the status of the service staff in the system can evolve dynamically and accurately with each task assignment, providing the latest and most accurate decision-making basis for responding to new demands. It is the foundation for maintaining the dynamic balance of the entire system and continuously optimizing scheduling accuracy.
[0120] In this embodiment of the invention, the synergy between the dynamic load coefficient and the service path allocation weight table is monitored in real time to trigger a secondary demand response compensation decision.
[0121] The real-time monitoring of the synergy between the dynamic load coefficient and the service path allocation weight table, triggering secondary demand response compensation decisions, includes: Based on the dynamic load coefficient and the service path allocation weight table, the system coordination deviation of the dining demand is calculated; When the system coordination deviation exceeds a preset first threshold, a secondary demand response compensation decision is triggered. The secondary demand response compensation decision includes initiating path reallocation for the dining demand and requesting auxiliary scheduling for the dining demand.
[0122] The formula for calculating the system's cooperative deviation is: ; in, The system's cooperative deviation is... Index for waiters, This represents the total number of service staff currently in service. waiter Real-time dynamic load factor, waiter Cumulative path weights Let be the standard complexity constant of the stated dining requirements. It is an exponential function. As an exponential adjustment factor, waiter The degree to which the skill attributes match the requirements of the current task. This refers to real-time environmental interference factors.
[0123] Specifically, real-time monitoring refers to the process by which a system continuously and automatically checks and analyzes specific operational indicators at extremely short intervals.
[0124] Specifically, synergy refers to the degree of cooperation and coordination between the actual workload reflected by the dynamic load coefficient of the waiter status and the task allocation plan reflected in the latest service path allocation weight table.
[0125] Specifically, the system coordination deviation refers to a specific value obtained after quantitatively evaluating the aforementioned coordination.
[0126] Specifically, the preset first threshold refers to a critical value pre-set by the system to determine whether the system's cooperative deviation has reached a dangerous or unacceptable level.
[0127] Specifically, secondary demand response compensation decision refers to the auxiliary decision-making process triggered by the system to correct or compensate for potential problems after the service path allocation weight table is generated and executed for the first time, when the system discovers potential problems through real-time monitoring.
[0128] Specifically, the auxiliary scheduling request refers to the automatic generation and sending of an alert and assistance request to the restaurant manager's terminal when the system is unable to effectively solve the problem through automatic path reallocation, prompting the manager to make manual intervention and scheduling.
[0129] Furthermore, the system periodically collects real-time data in two aspects: first, the latest dynamic load coefficient set of all online service providers; and second, the currently effective service path allocation weight table representing the future short-term task allocation plan.
[0130] The two sets of data are compared and analyzed: for example, analyzing whether the dynamic load coefficients of waiters with high weights in the service path allocation weight table are also high; or checking whether there are waiters with already high dynamic load coefficients who are still being assigned tasks in the latest weight table. The system uses specific internal logic and formulas to synthesize this difference between the plan and reality into a single, measurable value, that is, to calculate the current system coordination deviation.
[0131] Furthermore, the system compares the calculated system coordination deviation value with a preset first threshold stored in the configuration in real time. If the deviation value is detected to be continuously greater than the first threshold, the system determines that the current coordination state is unbalanced and there is an operational risk. Once this condition is met, the system immediately and automatically triggers the preset anomaly handling process, that is, it initiates the secondary demand response compensation decision mechanism. This process is automatic and has no delay, ensuring that problems can be detected and addressed in a timely manner.
[0132] Furthermore, the triggered secondary demand response compensation decision-making mechanism is not a single action, but a set of decisions including multiple alternatives. The system will intelligently select to execute one or more of these options based on the severity of the deviation and its specific causes.
[0133] The system attempts to automatically rerun the optimized conflict detection and scheduling algorithm in the background, focusing on currently high-load areas or personnel. The aim is to find better assignees for some assigned but unexecuted tasks and generate new local scheduling instructions to cover them.
[0134] When the system determines that automatic reallocation may not solve the problem, or when the situation requires more flexible handling, it will automatically generate an early warning message containing a specific problem description and drive the administrator's interactive terminal to output the auxiliary scheduling request, requesting the intervention of artificial intelligence.
[0135] In summary, the system coordination deviation, through continuous quantitative analysis of the difference between the plan and the actual situation, provides the system with an accurate metric for perceiving whether its operation is coordinated and healthy, and is the core of realizing intelligent monitoring.
[0136] In summary, the secondary decision-making mechanism sets a safety threshold, enabling the system to automatically switch from normal operation mode to emergency compensation mode before the deviation becomes too large and may affect the overall service efficiency or the service staff's condition. This demonstrates the system's proactive defense and self-adjustment capabilities.
[0137] In summary, providing compensatory measures such as path reallocation and auxiliary scheduling requests empowers the system to self-correct or seek external assistance when initial scheduling is imperfect or unexpected situations arise. This significantly enhances the robustness and practicality of the entire approach, ensuring that the service response chain does not break due to localized problems in complex real-world restaurant operations, thereby improving the overall reliability and service assurance level of the system.
[0138] Specifically, The system coordination deviation refers to a single comprehensive indicator calculated using this formula, which quantifies the degree of overall coordination imbalance between the current actual workload of the entire restaurant server team and the system task allocation plan.
[0139] Specifically, The waiter index is a loop variable used to iterate through and identify each waiter who is currently serving.
[0140] Specifically, The total number of servers currently in service refers to the total number of servers in the restaurant who are available to work at the time of calculation.
[0141] Specifically, waiter The real-time dynamic load factor refers to the load factor obtained by the system in real time at the calculation moment, which represents the load factor of the service staff. The quantitative value of the overall workload and stress level currently experienced.
[0142] Specifically, waiter The cumulative path weight refers to the quantitative value that represents the future task load that the system plans to allocate to service worker i, which is statistically derived based on the latest service path allocation weight table.
[0143] Specifically, The standard complexity constant for dining needs refers to a baseline complexity value predefined based on the task types in the set of service elements.
[0144] Specifically, An exponential function is a function that operates on an exponential basis with the natural constant e as its base.
[0145] Specifically, The exponential adjustment factor is a constant greater than 0 used to adjust the skill matching degree. Sensitivity to the influence within the exponential term. Adjusted by... The value of can control the degree to which the skill matching degree has a non-linear effect on the final adjustment coefficient.
[0146] Specifically, waiter The matching degree between a waiter's skill attributes and the requirements of the current task refers to the numerical value of the degree of conformity between a waiter i's personal skill expertise and the ability required by the currently evaluated dining needs task. The value range is typically [value range missing]. .
[0147] Specifically, The real-time environmental interference factor is a dynamic, positive-zero value used to quantify the complexity and interference level of the restaurant's overall operating environment at any given moment.
[0148] Furthermore, It directly captures the fundamental contradiction between the plan and reality for each server. This is the core deviation detection item, ensuring that the system can sensitively detect which server's assigned tasks are mismatched with their actual capacity.
[0149] Furthermore, divide by the standard complexity constant. Its function is to achieve deviation normalization. Intelligent weighting is introduced to modulate the original bias based on two key contextual factors.
[0150] In summary, this invention extracts user speech and recognizes it as text with high fidelity in a noisy restaurant environment by integrating sound source localization and environmental noise analysis; then, by using a pre-set catering rule knowledge base, it intelligently completes, corrects, and resolves logical conflicts in colloquial and fragmented requests, generating structured service instructions. This invention systematically solves three major problems in complex acoustic environments: unreliable voice interaction, inaccurate semantic understanding, and suboptimal dynamic scheduling of multiple waiters, thereby achieving a comprehensive improvement in the accuracy and timeliness of restaurant service response and overall operational efficiency.
[0151] like Figure 2 The diagram shown is a functional block diagram of a voice-interactive dining demand response system provided in an embodiment of the present invention.
[0152] The voice-interactive dining demand response system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the voice-interactive dining demand response system 100 may include an environment perception module 101, a voice processing module 102, a semantic parsing module 103, a scheduling optimization module 104, an instruction execution module 105, and a collaborative monitoring module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0153] In this embodiment, the functions of each module / unit are as follows: The environmental perception module 101 is used to obtain the coordinates of the user's sound source and the frequency domain characteristics of the ambient noise for those who require dining. The voice processing module 102 is used to construct a dynamic filtering mask for the dining needs based on the frequency domain features of the environmental noise, and extract the semantic segments of the dining needs by combining the user's sound source coordinates. The semantic parsing module 103 is used to perform structural completion on the semantic fragment of the demand based on a preset service conflict rule base, and generate a set of service elements for the dining demand. The scheduling optimization module 104 is used to perform conflict detection on the set of service elements and the status of the waiters, and generate a service path allocation weight table for the dining demand. The instruction execution module 105 is used to drive the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and at the same time update the dynamic load coefficient of the waiter status. The collaborative monitoring module 106 is used to monitor the synergy between the dynamic load coefficient and the service path allocation weight table in real time, and trigger secondary demand response compensation decisions.
[0154] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0155] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0157] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0158] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for responding to dining needs based on voice interaction, characterized in that, The method includes: S1: Obtain the coordinates of user sound sources and the frequency domain characteristics of ambient noise to meet dining needs; S2: Construct a dynamic filtering mask for the dining needs based on the frequency domain characteristics of the environmental noise, and extract the semantic fragments of the dining needs by combining the user's sound source coordinates; S3: Based on a preset service conflict rule base, the semantic fragment of the demand is structurally completed to generate a set of service elements for the dining demand; S4: Perform conflict detection on the set of service elements and the status of the waiter, and generate a service path allocation weight table for the dining demand; S5: Based on the service path weight allocation table, drive the interactive terminal of the dining demand to output service guidance instructions, and at the same time update the dynamic load coefficient of the waiter status. S6: Monitor the synergy between the dynamic load coefficient and the service path allocation weight table in real time, and trigger secondary demand response compensation decisions.
2. The method for responding to dining needs based on voice interaction as described in claim 1, characterized in that, The acquisition of user sound source coordinates and ambient noise frequency domain characteristics for dining needs includes: The mixed audio signal of the restaurant was acquired by a microphone array deployed in the restaurant; The mixed audio signal is preprocessed and framed, and a fast Fourier transform is performed to obtain the frequency domain representation of the mixed audio signal; Based on the frequency domain representation, the environmental noise frequency domain features of the mixed audio signal are extracted; Perform generalized cross-correlation time delay estimation on the mixed audio signal to calculate the microphone time difference for the dining requirement; Based on the microphone time difference and the geometry of the microphone array, the coordinates of the user's sound source for the dining request are calculated.
3. The method for responding to dining needs based on voice interaction as described in claim 2, characterized in that, The process of constructing a dynamic filtering mask for dining needs based on the frequency domain features of the environmental noise, and extracting semantic fragments of dining needs by combining the user's sound source coordinates, includes: Based on the frequency domain characteristics of the environmental noise, the environmental noise energy distribution of the dining demand is calculated, and an adaptive filtering threshold is determined based on a preset energy threshold. Based on the adaptive filtering threshold, a dynamic filtering mask for the dining requirements is constructed. Based on the user's sound source coordinates, the speech signal of the microphone array is spatially filtered to generate the user's enhanced speech signal for the dining needs; The user voice enhancement signal is filtered based on the dynamic filtering mask to obtain a clean user voice signal for the dining requirement. Endpoint detection is performed on the clean user voice signal to segment out the effective voice segment of the dining requirement, and speech recognition is performed on the voice segment to generate the semantic fragment of the dining requirement.
4. The method for responding to dining needs based on voice interaction as described in claim 1, characterized in that, The preset service conflict rule base includes colloquial expression grammar rules, service item conflict rules, and semantic slot templates.
5. A method for responding to dining needs based on voice interaction as described in claim 4, characterized in that, The system performs structural completion on the semantic fragments of the demand based on a preset service conflict rule base, generating a set of service elements for the dining demand, including: Parse the initial user intent of the semantic fragment of the requirement; Based on the colloquial expression grammar rules, the initial user intent is grammatically corrected and completed to generate a structured expression of the dining needs; Based on the service item conflict rules, logical conflict items in the structured intent representation are identified and resolved to generate an advanced structured intent representation of the dining needs. Based on the advanced structured intent description, fill in the set of service elements for the dining needs.
6. The method for responding to dining needs based on voice interaction as described in claim 1, characterized in that, The step of performing conflict detection on the set of service elements and the status of the waiters to generate a service path allocation weight table for the dining demand includes: Match the task requirements in the set of service elements with the status of the waiter; Calculate the expected path length, skill matching degree, and load impact for each server; The expected path length, the skill matching degree, and the load impact degree are integrated into a multi-dimensional conflict index for the dining demand. Based on the preset conflict resolution strategy, the multi-dimensional conflict indicators are comprehensively evaluated to generate the service path allocation weight for each service worker. The allocation weights of each waiter are integrated to generate a service path allocation weight table for the dining needs.
7. A method for responding to dining needs based on voice interaction as described in claim 6, characterized in that, The step of driving the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and simultaneously updating the dynamic load coefficient of the waiter status, includes: Parse the service path allocation weight table to obtain the waiter identifier, service path sequence, and expected service time for the dining demand; Based on the waiter's identifier and the service path sequence, a service guidance instruction for the dining request is generated, and the interactive terminal for the dining request is driven to output the service guidance instruction. Based on the expected path length, skill matching degree, and expected service time, calculate the additional load value for each waiter; Update the dynamic load coefficient of the waiter status based on the newly added load value.
8. A method for responding to dining needs based on voice interaction as described in claim 7, characterized in that, The real-time monitoring of the synergy between the dynamic load coefficient and the service path allocation weight table, triggering secondary demand response compensation decisions, includes: Based on the dynamic load coefficient and the service path allocation weight table, the system coordination deviation of the dining demand is calculated; When the system coordination deviation exceeds a preset first threshold, a secondary demand response compensation decision is triggered. The secondary demand response compensation decision includes initiating path reallocation for the dining demand and requesting auxiliary scheduling for the dining demand.
9. A method for responding to dining needs based on voice interaction as described in claim 8, characterized in that, The formula for calculating the system's cooperative deviation is: ; in, The system's cooperative deviation is... Index for waiters, showing the total number of waiters currently providing service. waiter Real-time dynamic load factor, waiter Cumulative path weights Let be the standard complexity constant of the stated dining requirements. It is an exponential function. As an exponential adjustment factor, waiter The degree to which the skill attributes match the requirements of the current task. This refers to real-time environmental interference factors.
10. A voice-interactive dining demand response system, used to implement the voice-interactive dining demand response system according to any one of claims 1-9, characterized in that, The system includes: The environmental perception module is used to obtain the coordinates of the user's sound source and the frequency domain characteristics of environmental noise when dining. The speech processing module is used to construct a dynamic filtering mask for the dining needs based on the frequency domain characteristics of the ambient noise, and extract the semantic segments of the dining needs by combining the user's sound source coordinates. The semantic parsing module is used to complete the structure of the semantic fragments of the demand based on a preset service conflict rule base, and generate a set of service elements for the dining demand. The scheduling optimization module is used to perform conflict detection on the set of service elements and the status of the waiters, and generate a service path allocation weight table for the dining demand. The instruction execution module is used to drive the interactive terminal of the dining demand to output service guidance instructions according to the service path allocation weight table, and at the same time update the dynamic load coefficient of the waiter status. The collaborative monitoring module is used to monitor the synergy between the dynamic load coefficient and the service path allocation weight table in real time, and trigger secondary demand response compensation decisions.