Coffee machine voice interaction method and system
By building and updating the kitchen acoustic space model in the coffee machine, the user's sound source is accurately located and signal enhancement and cancellation processing is performed, solving the noise and echo interference problems of the coffee machine in the complex kitchen environment, and achieving highly accurate and smooth voice interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AGZZX OPTOELECTRONICS TECH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-04-10
AI Technical Summary
Existing coffee machines offer a poor voice interaction experience in complex kitchen acoustic environments, are susceptible to noise and echo interference, and have low recognition accuracy.
By probing the kitchen space when there is no voice interaction with the coffee machine, an acoustic spatial model of the kitchen environment is constructed and dynamically updated to accurately locate the user's sound source. Spatial enhancement and active noise/echo cancellation processing are performed in parallel to separate the pure user's voice signal.
It significantly improves the accuracy of speech recognition and the smoothness of interaction, solves the noise and echo interference problems faced by traditional coffee machines in practical applications, and achieves high-precision and high-robustness voice interaction.
Smart Images

Figure CN121838765A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and more specifically, to a voice interaction method and system for a coffee machine. Background Technology
[0002] With the increasing popularity of smart home devices, users have higher demands for the human-computer interaction experience of coffee machines. Traditional coffee machine operation often relies on physical buttons or touchscreens, which can lead to inconvenience and response delays in daily use, especially when users have limited hands or in low light conditions. To solve these problems, voice interaction technology has been introduced into coffee machines, aiming to provide a more convenient and contactless operation method.
[0003] However, in a real home kitchen environment, the acoustic environment faced by a coffee machine is far from ideal. There are often multiple noise sources in a kitchen, such as the sound of appliances running and water flowing. These noises are not only high in intensity but also have different acoustic characteristics and may be superimposed on the user's voice signal.
[0004] For example, in a smart home environment, a coffee machine with voice interaction capabilities is placed on the kitchen counter. In an ideal quiet environment, the user can wake up the coffee machine with a simple voice command, such as "Hello, a latte, please," and have it automatically make coffee—the whole process is smooth and convenient. The coffee machine's built-in microphone array can clearly pick up the user's voice and accurately recognize the command through its internal voice processing unit, thereby controlling the grinding, heating, and pumping modules to work together.
[0005] However, in real-world home kitchen scenarios, the environment is far from ideal. Background noise, along with the user's voice signal, is picked up by the coffee machine's microphone. Although the system has some noise reduction capabilities, this continuous noise degrades the clarity of the voice signal, causing difficulties for the voice recognition unit in interpreting commands. The relative positions of the user, the coffee machine, and multiple noise sources in three-dimensional space become a crucial yet difficult-to-handle variable determining the success or failure of the interaction. Even with some sound source localization capabilities, the coffee machine's microphone system struggles to accurately pinpoint the user's mouth position and effectively suppress strong interference signals from other directions and reflection paths in this complex acoustic environment with multiple paths and interference sources. Ultimately, the system cannot extract clear and complete voice commands from the chaotic audio information, causing the entire voice interaction function to completely fail when the user most needs its convenient operation.
[0006] Therefore, providing a voice interaction method and system for coffee machines to solve the above problems is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] This application discloses a voice interaction method and system for a coffee machine, aiming to solve the problems of poor voice interaction experience, susceptibility to noise and echo interference, and low recognition accuracy in existing coffee machines in complex kitchen acoustic environments. The technical solution of this application is as follows:
[0008] In a first aspect, this application discloses a voice interaction method for a coffee machine, comprising the following steps:
[0009] When the coffee machine is not interacting with the user via voice, the system detects the kitchen space and listens to ambient background sounds to construct and dynamically update an acoustic space model of the kitchen environment. This acoustic space model at least characterizes the physical location and acoustic characteristics of noise sources within the kitchen space, as well as the sound reflection paths through which sound propagates within the kitchen space.
[0010] When a user's voice command is detected, the user's sound source is accurately located based on the sound reflection path represented in the acoustic space model of the kitchen environment, and the user's mouth position is obtained.
[0011] The following two processes are performed in parallel on the mixed audio signal containing the user's voice command: First, based on the user's mouth position, the speech signal component originating from the user's mouth position in the mixed audio signal is spatially enhanced to increase the proportion of the user's speech in the mixed audio signal; Second, based on the physical location and acoustic characteristics of the noise source and the sound reflection path represented in the acoustic spatial model of the kitchen environment, the noise signal component originating from the physical location of the noise source in the mixed audio signal is actively noise-cancelled, and the echo signal component introduced by the sound reflection path is actively echo-cancelled, so as to remove noise and echo from the mixed audio signal and separate the pure user speech signal.
[0012] The system performs voice recognition on the clean user voice signal and executes the corresponding coffee-making instructions based on the voice recognition results.
[0013] This technical solution effectively addresses noise and echo interference in the complex acoustic environment of a kitchen. By constructing an acoustic spatial model, accurately locating the user's sound source, and performing spatial enhancement and active cancellation in parallel, it significantly improves the accuracy of speech recognition and the fluency of interaction, solving the challenges faced by traditional coffee machine voice interaction in practical applications.
[0014] Furthermore, when the coffee machine is not interacting with the user via voice, the system can detect the kitchen space and listen to ambient background sounds to construct and dynamically update an acoustic space model of the kitchen environment. This includes the following steps:
[0015] When the coffee machine is not interacting with voice, it emits a sound wave into the kitchen space and receives the reflected signal of the sound wave after it is reflected by objects in the kitchen space through a microphone array.
[0016] The propagation time, intensity, and phase difference of the reflected signal to reach the microphone array are analyzed to identify the location of the reflective surface in the kitchen space and its acoustic reflection characteristics, and an acoustic reflection path map for subsequent localization and echo cancellation is constructed based on this.
[0017] The microphone array is used to monitor ambient background sounds, identify and record the acoustic characteristics of fixed noise sources in the environment, and determine the location of the fixed noise source based on the differences in the received sounds.
[0018] The location of the reflective surface and its acoustic reflection characteristics, along with the acoustic features of the fixed noise source, are integrated into the acoustic space model of the kitchen environment to achieve the construction or dynamic updating of the acoustic space model of the kitchen environment.
[0019] Through this technical solution, this application can comprehensively and accurately construct and dynamically update the acoustic space model of the kitchen environment by combining active detection and environmental monitoring. This provides a solid foundation of data for subsequent sound source localization, noise cancellation, and echo cancellation, thereby significantly improving the system's adaptability and processing capability to complex acoustic environments.
[0020] Based on this, when a user's voice command is detected, the user's sound source is accurately located based on the sound reflection path represented in the acoustic space model of the kitchen environment to obtain the user's mouth position, including the following steps:
[0021] Based on the detected user voice command, the arrival time difference of the user voice command to the microphone array is calculated by the sound source localization algorithm to obtain at least one candidate sound source direction;
[0022] By combining the acoustic reflection path diagram in the acoustic space model of the kitchen environment, the false sound source direction generated by the reflecting surface in the direction of the candidate sound source is identified;
[0023] The false sound source direction is eliminated from the candidate sound source direction, and the remaining direction is determined as the user's true sound source direction. The precise position of the user's mouth in three-dimensional space is calculated based on the true sound source direction.
[0024] Through this technical solution, this application can effectively distinguish between real and false sound sources by utilizing the reflection path information in the acoustic spatial model, thereby achieving accurate positioning of the user's mouth position. This avoids the positioning errors that are prone to occur in multipath environments by traditional sound source localization methods, and provides an accurate spatial reference for subsequent speech signal processing.
[0025] Furthermore, based on the user's mouth position, the speech signal component originating from the user's mouth position in the mixed audio signal is spatially enhanced. Specifically, based on the precise position of the user's mouth in three-dimensional space, the pickup mode of the microphone array is adjusted to perform super-directional beamforming to enhance the speech signal originating from the user's mouth direction and suppress sound from other directions.
[0026] Through this technical solution, this application can utilize precise user mouth position information to achieve super-directional beamforming, thereby maximizing the enhancement of the user's voice signal in space, while effectively suppressing noise and interference from non-user directions, and significantly improving the signal-to-noise ratio of the user's voice in the mixed audio signal.
[0027] Based on the above, after spatially enhancing the speech signal component originating from the user's mouth position in the mixed audio signal according to the user's mouth position, the following steps are also included: tracking the movement of the user's mouth position in real time during the duration of the user's voice command.
[0028] When a change in the user's mouth position is detected, the beamforming weighting coefficient and delay parameter of the spatial enhancement are recalculated based on the changed user's mouth position.
[0029] Furthermore, based on the change in the sound reflection path caused by the change in the user's mouth position, the filtering coefficient for active echo cancellation is recalculated.
[0030] Through this technical solution, this application can track the dynamic changes in the user's mouth position in real time and adjust the parameters of spatial enhancement and echo cancellation in a timely manner, ensuring that the system can still effectively enhance the user's voice signal and accurately cancel the echo during the user's movement, which greatly improves the robustness of voice interaction and the smoothness of user experience.
[0031] As an optional solution, before detecting the user's voice command, the following steps are also included: monitoring user activity and acoustic fluctuations near the coffee machine;
[0032] When abnormal user activity is detected and the sound environment is in a highly fluctuating state, or when the recognition confidence of the user's voice commands received continuously is lower than a preset threshold, it is judged that a potential interaction difficulty situation has been entered; where abnormal user activity means that the frequency or amplitude of the user's movement in front of the coffee machine exceeds the preset range of normal interaction.
[0033] In this potentially challenging interaction scenario, non-intrusive prompts are initiated, and the speech recognition strategy is adjusted to improve fault tolerance. Here, non-intrusive prompts refer to visual or auditory prompts that are delivered without interrupting the user's current operation.
[0034] Through this technical solution, this application can predict and identify potential interactive difficulties and take non-intrusive prompts and adjust speech recognition strategies. This proactively optimizes the interactive experience before the user explicitly expresses difficulty, effectively avoiding recognition failures caused by changes in environment or user state, and improving the system's intelligence and user-friendliness.
[0035] In one implementation, the step of adjusting the speech recognition strategy to improve fault tolerance includes one or more of the following:
[0036] Reduce the requirements of the acoustic model on the matching degree between input features and standard features during the speech recognition process;
[0037] In the language model upon which this speech recognition relies, the decoding priority of core instruction words related to coffee making is increased;
[0038] The confidence threshold of the speech recognition result is lowered from a first preset value to a second preset value; when the confidence of the speech recognition is lower than the second preset value for several consecutive times, guided interaction is initiated, and the user is given guidance to confirm or re-enter the command through voice questions or interface options.
[0039] Through this technical solution, this application can flexibly adjust the speech recognition strategy for potentially difficult interaction scenarios. By reducing the matching degree requirement, increasing the priority of core instruction words, lowering the confidence threshold, and initiating guided interaction, the system's speech recognition fault tolerance and user interaction success rate in complex or uncertain environments can be significantly improved.
[0040] In another embodiment, before parallel processing of the mixed audio signal, the following steps are included: upon detection of a user voice command, analyzing in real time the non-linguistic features of the initial audio signal carrying the user voice command, which is directly acquired by the microphone array; wherein the non-linguistic features include volume change rate, speech rate, and pitch change.
[0041] When the user's input pattern deviates from the normal pattern based on the non-linguistic feature, adaptive preprocessing is performed on the initial audio signal to obtain a preprocessed signal, which is then used as the mixed audio signal for subsequent parallel processing. The deviation from the normal pattern refers to one of the following situations:
[0042] The volume rise rate exceeds the preset threshold, the speech rate exceeds the preset ratio, or the pitch fluctuation range exceeds the preset multiple;
[0043] The adaptive preprocessing specifically involves selecting and executing a corresponding processing method from a variety of preset processing methods based on the category of the detected deviation from the normal pattern.
[0044] Through this technical solution, this application can monitor the non-linguistic features of user voice commands in real time and perform adaptive preprocessing based on abnormalities in user input methods. This effectively avoids the degradation of voice signal quality caused by sudden changes in user emotions or environment, thereby providing higher quality input for subsequent signal processing and speech recognition, and improving the robustness and adaptability of the system.
[0045] Preferably, the preset multiple processing methods include: when the volume rise rate is detected to exceed a preset threshold, adjusting the gain of the microphone channel to prevent the initial audio signal from being overloaded;
[0046] When the speech rate is detected to exceed the preset ratio, the time-frequency analysis parameters of the initial audio signal are adjusted to match the fast speech.
[0047] When the detected pitch fluctuation range exceeds a preset multiple, the spectrum of the initial audio signal is smoothed.
[0048] Through this technical solution, this application can adopt customized adaptive preprocessing measures, such as gain adjustment, time-frequency parameter adjustment and spectrum smoothing, for different types of abnormal user input patterns, thereby accurately optimizing the quality of the initial audio signal, effectively responding to the challenges brought about by changes in user emotions or environment, and further improving the stability and accuracy of voice interaction.
[0049] Secondly, this application also discloses a voice interaction system for a coffee machine, including: an acoustic space modeling module, used to detect and listen to ambient background sounds in the kitchen space when the coffee machine is not performing voice interaction, so as to construct and dynamically update the acoustic space model of the kitchen environment, wherein the acoustic space model of the kitchen environment at least characterizes the physical location and acoustic characteristics of noise sources in the kitchen space, as well as the sound reflection path of sound propagation in the kitchen space.
[0050] The sound source localization module is used to accurately locate the user's sound source based on the sound reflection path represented in the acoustic space model of the kitchen environment when a user's voice command is detected, so as to obtain the position of the user's mouth.
[0051] The signal processing module is used to perform the following two processes in parallel on the mixed audio signal containing the user's voice command: First, based on the user's mouth position, spatially enhance the speech signal component in the mixed audio signal originating from the user's mouth position; Second, based on the physical location and acoustic characteristics of the noise source represented in the acoustic spatial model of the kitchen environment and the sound reflection path, actively cancel the noise signal component in the mixed audio signal originating from the physical location of the noise source, and actively cancel the echo signal component introduced by the sound reflection path, so as to separate the pure user voice signal from the mixed audio signal.
[0052] A voice recognition module is used to recognize the clean user voice signal; an execution control module is used to execute the corresponding coffee-making instructions based on the voice recognition results. Through this technical solution, this application provides a coffee machine voice interaction system integrating acoustic space modeling, precise sound source localization, parallel signal processing, and intelligent execution. This system can work collaboratively to effectively overcome the challenges posed by the complex acoustic environment of a kitchen, achieving high-precision and highly robust voice interaction, and significantly improving the convenience and intelligence of user operation of the coffee machine.
[0053] Beneficial effects
[0054] This application discloses a voice interaction method for coffee machines. When the coffee machine is not interacting with the user, it detects and listens to ambient background sounds in the kitchen space, constructs and dynamically updates an acoustic spatial model of the kitchen environment. This model characterizes the physical location, acoustic characteristics, and sound reflection paths of noise sources. When a user's voice command is detected, the user's mouth position can be accurately located based on the sound reflection paths in the acoustic spatial model. Subsequently, spatial enhancement and active noise / echo cancellation processing are performed in parallel on the mixed audio signal containing the user's voice command, thereby separating the pure user's voice signal from the mixed audio signal and performing speech recognition to execute coffee-making instructions. This method effectively solves the problem of low speech recognition accuracy and poor interactive experience caused by noise and echo interference in complex kitchen acoustic environments in existing technologies. By constructing a refined acoustic spatial model, this application can accurately identify and locate noise sources and echo paths, and perform targeted signal processing, significantly improving the clarity and recognizability of the user's voice. This achieves high-precision and robust voice interaction in noisy environments, greatly improving the convenience and intelligence of operating the coffee machine. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 A flowchart illustrating a voice interaction method for a coffee machine provided in an embodiment of the present invention;
[0057] Figure 2 A flowchart of step S1 provided in an embodiment of the present invention;
[0058] Figure 3 A flowchart of step S2 provided in an embodiment of the present invention;
[0059] Figure 4 This is a flowchart provided by an embodiment of the present invention before detecting a user's voice command;
[0060] Figure 5 This is a schematic diagram of the structure of a coffee machine voice interaction system provided in an embodiment of the present invention. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] The embodiments of this invention are written in a progressive manner.
[0063] like Figure 1 As shown, a voice interaction method for a coffee machine includes the following steps:
[0064] S1. When the coffee machine is not interacting with the user via voice, the kitchen space is detected and the ambient background sounds are monitored in order to construct and dynamically update the acoustic space model of the kitchen environment; wherein, the acoustic space model of the kitchen environment at least represents the physical location and acoustic characteristics of the noise sources in the kitchen space, as well as the sound reflection path of the sound propagation in the kitchen space.
[0065] S2. When a user's voice command is detected, the user's sound source is accurately located based on the sound reflection path represented in the acoustic space model of the kitchen environment to obtain the user's mouth position.
[0066] S3. Perform the following two processes in parallel on the mixed audio signal containing user voice commands: First, based on the user's mouth position, spatially enhance the speech signal component originating from the user's mouth position in the mixed audio signal to increase the proportion of user speech in the mixed audio signal; Second, based on the physical location and acoustic characteristics of the noise source and the sound reflection path represented in the acoustic spatial model of the kitchen environment, actively cancel the noise signal component originating from the physical location of the noise source in the mixed audio signal, and actively cancel the echo signal component introduced by the sound reflection path, so as to remove noise and echo from the mixed audio signal and separate the pure user speech signal;
[0067] S4. Perform speech recognition on the clean user's voice signal and execute the corresponding coffee making instructions based on the speech recognition results.
[0068] The kitchen environmental acoustic space model is a comprehensive data structure that records the acoustic characteristics of the kitchen space in detail. Specifically, this model characterizes at least the physical location and acoustic features of various noise sources within the kitchen space, such as the hum of the refrigerator and the roar of the range hood, as well as the sound reflection paths that sound experiences as it propagates within the kitchen space, such as the path of sound emitted from the user's mouth, reflected by objects like walls and cabinets, and finally reaching the microphone. The construction and dynamic updating of this model are crucial for subsequent precise localization of user sound sources, spatial enhancement, active noise cancellation, and active echo cancellation.
[0069] The user's mouth position refers to the precise coordinates of their mouth in three-dimensional space when they issue a voice command. Obtaining this positional information helps the coffee machine system focus more accurately on the user's voice and effectively suppress interference from other directions.
[0070] Mixed audio signals refer to the raw audio signals collected by a microphone array in a kitchen environment, which include user voice commands, ambient noise, and echo signals caused by sound reflection paths.
[0071] A clean user voice signal refers to a signal that has undergone a series of signal processing steps to effectively remove noise and echoes from a mixed audio signal, retaining only the content of the user's voice commands.
[0072] In step S1, when the coffee machine is not interacting with the user via voice, the system continuously probes the kitchen space and listens to ambient background sounds to construct and dynamically update the kitchen's acoustic spatial model. For example, it can emit probe sound waves of a specific frequency into the kitchen space and use a microphone array to receive the signals reflected from objects within the kitchen. By analyzing the propagation time, intensity, and phase difference of the reflected signals, the location of reflective surfaces and their acoustic reflection characteristics can be identified, thus constructing an acoustic reflection path map. Simultaneously, the microphone array continuously listens to ambient background sounds, identifying and recording the acoustic characteristics of fixed noise sources in the environment, such as the operating sounds of appliances like refrigerators and dishwashers, and determining the location of these noise sources based on the differences in the received sounds. This information is then integrated into the kitchen's acoustic spatial model, enabling the model's construction or dynamic updating.
[0073] In step S2, when the system detects a user's voice command, it accurately locates the user's sound source based on the sound reflection paths represented in the acoustic space model of the kitchen environment to determine the user's mouth position. For example, the user's voice signal received by the microphone array can be combined with the sound reflection path information in the acoustic space model, and advanced sound source localization algorithms, such as multipath sound source localization algorithms, can be used to calculate the precise position of the user's mouth in three-dimensional space. This method can effectively distinguish between real sound sources and false sound sources caused by reflections, thereby improving the accuracy of localization.
[0074] In step S3, two key processes are then performed in parallel on the mixed audio signal containing the user's voice commands. First, based on the determined user's mouth position, the speech signal component originating from that position in the mixed audio signal is spatially enhanced. For example, the microphone array's pickup pattern can be dynamically adjusted based on the precise position of the user's mouth in three-dimensional space to perform super-directional beamforming. This beamforming technique can precisely align the microphone array's pickup focus with the user's mouth, thereby significantly enhancing the speech signal originating from that direction and effectively suppressing interference from other directions, increasing the proportion of the user's voice in the mixed audio signal. Second, based on the physical location and acoustic characteristics of the noise source, as well as the sound reflection path, as represented in the acoustic spatial model of the kitchen environment, active noise cancellation is performed on the noise signal component originating from the physical location of the noise source in the mixed audio signal, and active echo cancellation is performed on the echo signal component introduced by the sound reflection path. For example, for noise cancellation, the physical location and acoustic characteristics of the noise source recorded in the model can be used to generate a cancellation signal that is inversely phase to the noise signal through an adaptive filter, and this signal is superimposed on the mixed audio signal to achieve effective noise elimination. For echo cancellation, the characteristics of the echo signal can be predicted using the sound reflection path information recorded in the model, and an estimated echo signal can be generated and canceled using an adaptive echo cancellation algorithm. Through these two parallel processes, noise and echo can be removed from the mixed audio signal, separating the clean user speech signal.
[0075] In step S4, the final step is to perform speech recognition on the separated, clean user voice signal and execute the corresponding coffee-making instructions based on the speech recognition results. For example, if the recognition result is "make a latte," the coffee machine system will start the corresponding coffee-making program.
[0076] Compared with existing technologies, the advantages of this application are as follows: First, through a dynamically updated acoustic spatial model, it can achieve precise localization of the user's sound source, effectively distinguishing between real speech and reflected sound, thus solving the problem of insufficient accuracy of traditional sound source localization in multipath environments. Second, in the signal processing stage, this application employs spatial enhancement, active noise cancellation, and active echo cancellation technologies in parallel. Spatial enhancement can significantly improve the signal-to-noise ratio of the user's speech through super-directional beamforming based on the user's mouth position. Simultaneously, active noise cancellation and active echo cancellation precisely process the noise source and sound reflection path, respectively, fundamentally eliminating interference components in the mixed audio signal. This comprehensive and refined processing method results in a purer final separated user speech signal, greatly improving the accuracy and robustness of speech recognition. Therefore, the method of this application can ensure that the coffee machine provides a smooth and accurate voice interaction experience even in a noisy kitchen environment, significantly improving the convenience and satisfaction of user operation.
[0077] like Figure 2 As shown, step S1 includes the following steps:
[0078] A1. When the coffee machine is not interacting with voice, it emits detection sound waves into the kitchen space and receives the reflected signals of the detection sound waves after they are reflected by objects in the kitchen space through a microphone array;
[0079] A2. Analyze the propagation time, intensity, and phase difference of the reflected signal to reach the microphone array to identify the location of the reflective surface in the kitchen space and its acoustic reflection characteristics, and construct an acoustic reflection path map for subsequent localization and echo cancellation based on this.
[0080] A3. Listen to ambient background sounds through a microphone array, identify and record the acoustic characteristics of fixed noise sources in the environment, and determine the location of the fixed noise sources based on the differences in the received sounds;
[0081] A4. Integrate the location of the reflective surface and its acoustic reflection characteristics, along with the acoustic features of the fixed noise source, into the acoustic space model of the kitchen environment to achieve the construction or dynamic updating of the acoustic space model of the kitchen environment.
[0082] In step A1, the probe sound waves can be understood as a pre-set sound signal with a specific frequency and pattern, the purpose of which is to actively probe the physical structure and acoustic environment within the kitchen space. A microphone array is configured to receive signals reflected back from these probe sound waves after they encounter objects within the kitchen (such as walls, cabinets, appliances, etc.). By analyzing these reflected signals, key information about the geometry and acoustic characteristics of the kitchen space can be obtained.
[0083] In step A2, the propagation time, intensity, and phase difference of the reflected signal upon reaching the microphone array are analyzed to accurately identify the location and acoustic reflection characteristics of each reflective surface within the kitchen space. For example, by measuring the time difference between sound wave emission and reception, the distance to the reflective surface can be calculated; by analyzing the signal intensity attenuation, the sound absorption or reflection capability of the reflective surface can be assessed; and the phase difference helps determine the relative paths of sound waves reaching different microphones, thus more accurately depicting the geometry of the reflective surface. Based on these identification results, an acoustic reflection path map can be constructed for subsequent sound source localization and echo cancellation, which details the multipath propagation of sound within the kitchen space.
[0084] In step A3, simultaneously, ambient background sounds are continuously monitored using a microphone array. This allows for the identification and recording of the acoustic characteristics of persistent, fixed noise sources in the environment, such as refrigerators, range hoods, and air conditioners. Based on differences in the received sounds, such as the intensity or phase differences of noise signals received by different microphones, the location of these fixed noise sources can be further determined. The acoustic characteristics and location information of these noise sources are crucial for subsequent active noise cancellation.
[0085] In step A4, the locations and acoustic reflection characteristics of the identified reflective surfaces, along with the acoustic characteristics and directional information of the identified fixed noise sources, are effectively fused. This information is integrated into the acoustic space model of the kitchen environment, thereby enabling the model's construction or dynamic updating. This fusion ensures that the model can comprehensively and accurately characterize the physical locations and acoustic characteristics of noise sources within the kitchen space, as well as the acoustic reflection paths of sound propagation within the kitchen space.
[0086] like Figure 3 As shown, step S2 includes the following steps:
[0087] B1. Based on the detected user voice command, calculate the time difference of arrival of the user voice command to the microphone array through a sound source localization algorithm to obtain at least one candidate sound source direction;
[0088] B2. Based on the acoustic reflection path diagram in the acoustic space model of the kitchen environment, identify the false sound source directions generated by the reflecting surface in the direction of the candidate sound source;
[0089] B3. Eliminate false sound source directions from the candidate sound source directions, determine the remaining directions as the user's true sound source directions, and calculate the precise position of the user's mouth in three-dimensional space based on the true sound source directions.
[0090] Specifically, sound source localization algorithms can employ various existing technologies, such as algorithms based on principles like Time Difference of Arrival (TDOA), Angle of Arrival (DOA), or energy intensity. Their purpose is to initially determine the possible directions of the sound source. The Time Difference of Arrival (TDOA) can be understood as the time difference between the arrival times of sound emitted from the same sound source at different microphones in a microphone array. By analyzing these time differences, the location of the sound source can be inferred. In practical applications, candidate sound source directions refer to a set of possible sound source directions initially calculated by the sound source localization algorithm. These directions may include false directions caused by sound wave reflection within the kitchen space. Furthermore, the acoustic reflection path map is constructed using the above steps. Its purpose is to record in detail the location of reflective surfaces within the kitchen space, their acoustic reflection characteristics, and the sound propagation path in the space. False sound source directions refer to situations where, due to sound wave reflection on reflective surfaces, the reflected sound waves received by the microphone array are mistakenly identified as originating from a sound source in a certain direction, which is not the actual direction of the sound source. The true sound source direction refers to the direction in which the user actually emits a voice command. Ultimately, the precise location of the user's mouth in three-dimensional space is calculated using triangulation or other geometric positioning methods, combined with the actual sound source direction and the known location of the microphone array. The purpose is to provide the accurate spatial coordinates of the user's mouth.
[0091] In some preferred embodiments, it is assumed that the coffee machine is located in the center of the kitchen, and its microphone array is activated. When a user issues a voice command "I want a latte" from a location in the kitchen, the microphone array receives both the direct sound and the reflected sound from surfaces such as kitchen walls, refrigerators, and cabinets. Traditional sound source localization algorithms may incorrectly identify multiple false sound source directions based on these reflected sound waves. However, in the solution of this application, the system first calculates the arrival time difference of the sound to the microphone array based on the detected user voice command using a sound source localization algorithm such as Generalized Cross-Correlation (GCC-PHAT), thereby obtaining a set of candidate sound source directions containing both real and false sound sources. Subsequently, the system combines this with an acoustic reflection path map from a pre-constructed acoustic spatial model of the kitchen environment. This path map records in detail the location and acoustic reflection characteristics of each reflective surface in the kitchen. By analyzing the correspondence between these candidate sound source directions and the reflection path map, the system can identify which directions are false sound sources caused by sound waves reflecting off specific reflective surfaces (e.g., the tiled walls of the kitchen or the surface of the metal refrigerator). For example, if a candidate sound source direction closely matches the reflection path produced by a reflective surface, that direction is identified as a false sound source. After eliminating all identified false sound source directions, the remaining unique or most prominent direction is determined as the user's true sound source direction. Finally, based on this true sound source direction, the system accurately calculates the coordinates of the user's mouth in three-dimensional space, such as (x, y, z) meters, providing accurate localization information for subsequent speech signal processing.
[0092] The spatial enhancement process in step S3 can be understood as achieving directional capture and enhancement of the target speech signal by precisely controlling the pickup characteristics of the microphone array.
[0093] Based on the precise position of the user's mouth in three-dimensional space, the microphone array's pickup pattern is adjusted to perform super-directional beamforming, thereby enhancing the speech signal originating from the user's mouth and suppressing sounds from other directions.
[0094] Specifically, the pickup pattern of a microphone array refers to its sensitivity distribution to sound from different directions. By adjusting this pattern, the microphone array's pickup focus can be precisely aligned with the user's mouth. Superdirectional beamforming is an advanced signal processing technique that aims to create an extremely narrow sound beam that is highly focused in a specific direction (i.e., the direction of the user's mouth), thereby maximizing the pickup of speech signals from that direction while significantly attenuating or suppressing sound from non-target directions (such as other noise sources in the kitchen or reflected sound).
[0095] This application further proposes that, after spatially enhancing the speech signal component originating from the user's mouth position in the mixed audio signal according to the user's mouth position, the following steps are also included:
[0096] The movement of the user's mouth position is tracked in real time during the duration of the user's voice command;
[0097] When a change in the user's mouth position is detected, the beamforming weighting coefficients and delay parameters for spatial enhancement are recalculated based on the changed user's mouth position.
[0098] Furthermore, the filtering coefficients for active echo cancellation are recalculated based on the changes in the sound reflection path caused by the change in the user's mouth position.
[0099] Real-time tracking of the user's mouth position refers to the system continuously monitoring the location of the user's sound source as the user issues voice commands and engages in continuous voice interaction. This can be achieved, for example, by utilizing the sound source localization capabilities of a microphone array, periodically or event-drivenly re-executing the sound source localization algorithm at a preset frequency or when significant acoustic changes are detected. The goal is to ensure that the system always has the latest spatial information about the user's sound source.
[0100] Detecting a change in the user's mouth position can be understood as follows: when the spatial distance or directional angle between the real-time tracked user's mouth position and the previously determined position exceeds a preset threshold, the system determines that the user's mouth position has changed significantly. For example, a distance threshold (such as 5 centimeters) or an angle threshold (such as 5 degrees) can be set.
[0101] Recalculating the beamforming weighting coefficients and delay parameters based on the changed user mouth position means that once a change in the user's mouth position is detected, the system immediately uses the new three-dimensional spatial position information of the user's mouth to readjust the microphone array's pickup pattern. This includes updating the weighting coefficients used for hyperdirectional beamforming and the delay parameters of each microphone channel to ensure that the beam is always precisely pointed at the user's mouth, thereby continuously and effectively enhancing the user's voice signal and suppressing interference from other directions.
[0102] The recalculation of active echo cancellation filter coefficients, based on the altered sound reflection paths caused by changes in the user's mouth position, means that the direct path of sound from the user's mouth to the microphone array and the sound reflection path after reflection from objects within the kitchen space both change due to the altered mouth position. Therefore, the system needs to re-evaluate and update the filter coefficients used in the active echo cancellation algorithm based on the new user mouth position and the sound reflection path information represented in the acoustic space model of the kitchen environment. This ensures that the echo cancellation algorithm can adapt to the new acoustic environment and effectively cancel the echo signal components introduced by the new reflection paths.
[0103] In some preferred embodiments, a specific example is given below. Suppose a user issues a voice command, "A latte, please," at the coffee machine. While waiting for the coffee to be made, the user moves approximately 10 centimeters to the left. The coffee machine's built-in microphone array continuously monitors the user's voice. When the system detects that the user's mouth position has moved from its initial position to the new position, the sound source localization module immediately updates the precise position of the user's mouth in three-dimensional space. Subsequently, the signal processing module recalculates the beamforming weighting coefficients and delay parameters of the microphone array based on this new mouth position, re-aligning the superdirectional beam with the user's new mouth position, thereby continuing to effectively enhance the user's voice. Simultaneously, the system recalculates the active echo cancellation filter coefficients based on the changes in the sound reflection path caused by the change in the user's mouth position—for example, a slight change in the path of sound reflected from the kitchen countertop to the microphone array after being emitted from the new position—to ensure the accuracy of echo cancellation. Through this real-time adaptive adjustment, even if the user moves slightly within the kitchen space, the coffee machine can continue to accurately receive and recognize the user's voice commands, such as a subsequent command, "Add more sugar."
[0104] like Figure 4 As shown, this application further proposes a voice interaction method for a coffee machine that includes the following steps before detecting a user's voice command:
[0105] C1. Monitor user activity and ambient sound fluctuations near the coffee machine;
[0106] C2. When abnormal user activity is detected and the sound environment is in a state of high fluctuation, or when the recognition confidence of continuously received user voice commands is lower than the preset threshold, it is judged as entering a potential interaction difficulty situation; where abnormal user activity refers to the frequency or amplitude of the user's movement in front of the coffee machine exceeding the preset range during normal interaction.
[0107] C3. In situations where interaction is potentially difficult, non-intrusive prompts are activated, and the speech recognition strategy is adjusted to improve the fault tolerance rate. Non-intrusive prompts refer to visual or auditory prompts that are given without interrupting the user's current operation.
[0108] Monitoring user activity and acoustic fluctuations near the coffee machine can be understood as using sensors integrated into or around the coffee machine (e.g., infrared sensors, millimeter-wave radar, cameras, etc.) to sense information such as the user's movement and posture in front of the coffee machine, and continuously monitoring changes in ambient background sound, such as noise levels and the frequency of sound events, through a microphone array. The aim is to obtain real-time information on the user's interaction with the environment, providing a data foundation for identifying potential interaction difficulties.
[0109] When abnormal user activity is detected and the sound environment is highly fluctuating, or when the confidence level of consecutively received user voice commands is below a preset threshold, the system will determine that it has entered a potentially difficult interaction situation. Specifically, abnormal user activity refers to the frequency or amplitude of the user's movement in front of the coffee machine exceeding the preset range for normal interaction, such as the user walking quickly in front of the coffee machine or waving their hand dramatically. This may indicate that the user is busy or emotionally agitated, making regular voice interaction unsuitable. A highly fluctuating sound environment refers to a significantly increased ambient noise level or frequent sudden noises, such as the operation of kitchen appliances or the sound of tableware clattering. Furthermore, if the system repeatedly attempts to recognize user voice commands, but the confidence level of each recognition result is below a preset threshold, it also indicates that the current interaction is difficult. Meeting any of these conditions will trigger the system to enter a potentially difficult interaction situation.
[0110] In situations where interaction may be challenging, the system will initiate non-intrusive prompts and adjust its speech recognition strategy to improve fault tolerance. Non-intrusive prompts refer to visual or auditory cues that are delivered without interrupting the user's current operation. For example, the coffee machine screen can display gentle prompts such as "The environment is noisy, please wait" or "Please speak clearly," or use soft indicator flashing or low-volume prompts to inform the user that the current environment may be unfavorable for voice interaction without forcibly interrupting the user's workflow. Simultaneously, the system will adjust its speech recognition strategy, such as lowering the stringent requirements for speech recognition accuracy or prioritizing the recognition of core instructions related to coffee making, to increase the likelihood of correct instruction recognition even under unfavorable conditions.
[0111] In some preferred embodiments, a specific example is given below. Suppose a user is busy preparing breakfast in the kitchen, with the dishwasher running and generating continuous background noise, while the user is frequently moving around in front of the coffee machine looking for ingredients. At this moment, the user attempts to give the coffee machine a voice command, "Give me a latte."
[0112] According to the solution in this application, the coffee machine will first continuously monitor the frequency and amplitude of the user's movements in front of the coffee machine, and listen to the background noise generated by the dishwasher. When the system detects that the user's movement frequency exceeds the normal range (abnormal user activity) and the ambient background sound is in a state of high fluctuation (dishwasher noise), or when the system recognizes the confidence level is lower than a preset threshold after the user issues a command for the first time, the system will determine that it is currently in a potentially difficult interaction situation.
[0113] In this scenario, the coffee machine will not immediately perform conventional voice recognition processing, but will instead initiate a non-intrusive prompt. For example, the coffee machine screen may display a soft text prompt: "The environment is noisy, please speak clearly," while the indicator light on the side of the coffee machine will flash slowly. After seeing the prompt, the user may subconsciously move closer to the coffee machine and issue the command "I'd like a latte" again in a clearer and slower voice.
[0114] At the same time, the system will adjust its speech recognition strategy. For example, it may reduce the requirement for the matching degree between input features and standard features in the acoustic model, or increase the decoding priority of core command words such as "latte" and "coffee" in the language model. When the user issues the command again, even if the speech signal is still affected by some noise or the user speaks slightly faster, the adjusted recognition strategy can successfully recognize the command "get a latte" with a higher fault tolerance and execute the corresponding coffee making operation.
[0115] In this way, the solution proposed in this application avoids the frustration of users repeatedly trying commands in adverse environments, and improves the smoothness of voice interaction and user satisfaction.
[0116] This application further proposes that the steps for adjusting the speech recognition strategy to improve the fault tolerance rate include one or more of the following methods:
[0117] Reduce the matching degree between input features and standard features required by the acoustic model during speech recognition;
[0118] In the language model upon which speech recognition relies, increase the decoding priority of core instruction words related to coffee making;
[0119] The confidence threshold of the speech recognition results is lowered from the first preset value to the second preset value;
[0120] When the confidence level of voice recognition falls below the second preset value for multiple consecutive times, guided interaction is initiated, providing the user with guidance to confirm or re-enter the command through voice questions or interface options.
[0121] Reducing the matching requirement between input features and standard features in the acoustic model means allowing greater differences between the acoustic features of the input speech and the preset standard speech feature template during the speech recognition process, while still being judged as a match. The aim is to ensure that the system can still attempt recognition even when the user's speech is distorted due to environmental noise, changes in speech rate, or inaccurate pronunciation, thereby improving the recognition success rate under non-ideal conditions.
[0122] In the language model upon which speech recognition relies, prioritizing the decoding of core instructions related to coffee making can be understood as assigning higher weights to words related to the core functions of the coffee machine, such as "make a latte," "get an Americano," and "heat the milk," during the decoding process. Specifically, when the speech recognition engine selects from multiple possible word sequences, it prioritizes sequences containing these high-priority instructions, even if their acoustic matching is slightly lower than other non-core instruction sequences. The goal is to ensure that, when recognition is difficult, the user's most frequently used core instructions are recognized accurately and preferentially, reducing instruction failures due to misrecognition.
[0123] Lowering the confidence threshold for speech recognition results from a first preset value to a second preset value means relaxing the confidence standard required by the system when judging whether a recognition result is valid. For example, if a confidence level of 0.8 is required for acceptance in normal mode, it can be lowered to 0.6 in difficult scenarios. The purpose is to avoid frequently rejecting valid user commands due to overly strict confidence requirements when the recognition difficulty increases, thereby improving the command acceptance rate.
[0124] When the confidence level of voice recognition falls below the second preset value after multiple consecutive attempts, guided interaction is initiated. This means the system no longer passively waits for the user to re-enter, but actively intervenes, guiding the user to confirm or reselect through voice prompts (e.g., "What kind of coffee would you like to make?") or by providing preset options on the coffee machine's display screen (e.g., "Latte," "Americano," "Cappuccino"). The purpose is to reduce the user's interaction threshold and cognitive burden by proactively guiding them when the system's recognition capabilities are limited, quickly correcting recognition errors, and improving user experience and interaction efficiency.
[0125] In some preferred embodiments, suppose a user is busy in the kitchen while attempting to make coffee via voice commands. At this time, there may be continuous noise from the range hood operating in the kitchen, or the user's voice commands, uttered while moving, may contain background noise and an accent. According to the method described above, the coffee machine will first determine that the current situation is a potentially difficult interaction scenario.
[0126] Specifically, when a user says "I'd like a latte," if the system detects that the acoustic features of the user's speech match the standard pronunciation of "latte" slightly below the normal threshold, the system will still attempt to recognize it as "latte" because the acoustic model's matching requirement has been lowered. At the same time, since "latte" is a core instruction related to coffee making, its decoding priority in the language model is increased, further increasing the likelihood of correct recognition.
[0127] If the system identifies a confidence level of 0.65 for "latte", while the confidence level threshold in normal mode is 0.7, but has been lowered to 0.6, the instruction will still be accepted and executed.
[0128] However, if a user attempts to issue a command, such as "make a cappuccino," but the system's confidence level for each of the three consecutive recognitions is below 0.6 due to excessive ambient noise or unclear pronunciation, the system will initiate guided interaction. The coffee machine may then prompt with a voice message, "Sorry, I didn't hear you clearly. Do you want a cappuccino?" and display options such as "cappuccino," "Americano," and "latte" on the screen. The user simply needs to answer "yes" or tap the option on the screen to confirm the command, thus avoiding the dilemma of repeated attempts by the user and continuous misrecognition by the system, significantly improving interaction efficiency and user satisfaction.
[0129] This application further proposes a voice interaction method for coffee machines that includes the following steps before parallel processing of the mixed audio signals:
[0130] When a user's voice command is detected, the non-linguistic features of the initial audio signal carrying the user's voice command, which is directly acquired by the microphone array, are analyzed in real time. These non-linguistic features include volume change rate, speech rate, and pitch change.
[0131] When the user's input method deviates from the normal mode based on non-linguistic features, the initial audio signal is adaptively preprocessed to obtain the preprocessed signal, and the preprocessed signal is used as the mixed audio signal for subsequent parallel processing.
[0132] Among them, deviation from normal mode refers to one of the following situations: the volume rise rate exceeds the preset threshold, the speech rate exceeds the preset ratio, or the pitch fluctuation range exceeds the preset multiple;
[0133] Adaptive preprocessing specifically involves selecting and executing a corresponding processing method from a variety of preset processing methods based on the category of the detected deviation from the normal pattern.
[0134] Non-linguistic features refer to acoustic attributes that are unrelated to the speech content but reflect the user's speaking state and emotions. Volume change rate can be understood as how quickly sound intensity changes over time, reflecting the loudness dynamics of the user's speech; speech rate refers to the number of syllables or words pronounced per unit time, reflecting the speed of the user's speech; pitch change refers to the range of fluctuation of the fundamental frequency of the sound, reflecting the stability of the user's pitch. These features are typically obtained through time-domain and frequency-domain analysis of the initial audio signal.
[0135] The judgment of deviations from the normal pattern is achieved by comparing non-linguistic features obtained from real-time analysis with the feature range of the preset normal interaction pattern. For example, preset thresholds, preset ratios, and preset multiples can be derived from statistical analysis of a large amount of user interaction data to define what degree of volume, speech rate, or pitch change is considered abnormal.
[0136] The purpose of adaptive preprocessing is to make preliminary adjustments to the speech signal before it enters the subsequent spatial enhancement and noise echo cancellation modules, so that it better meets the input requirements of the speech recognition system. Preset processing methods may include, but are not limited to, gain adjustment, time-frequency analysis parameter adjustment, and spectral smoothing, each optimized for a specific type of deviation pattern.
[0137] In some preferred embodiments, suppose a user, in a kitchen environment, due to excitement or eagerness, suddenly commands the coffee machine to "make a latte" at a very high volume and extremely fast speed. When the microphone array detects the user's voice command and acquires the initial audio signal, the system immediately analyzes the non-linguistic features of the signal. Specifically, it will detect that the volume increase rate far exceeds a preset threshold, and the speech rate also exceeds a preset ratio. In this case, the system will determine that the user's input method deviates from the normal pattern.
[0138] To prevent clipping distortion of the initial audio signal due to excessive volume, the system adjusts the microphone channel gain to reduce the overall signal strength and keep it within a suitable dynamic range. Simultaneously, to better handle fast-moving speech, the system adjusts time-frequency analysis parameters, such as shortening frame length or increasing frame shift, to more precisely capture rapidly changing speech features. After this adaptive preprocessing, the resulting preprocessed signal is sent as a mixed audio signal to subsequent spatial enhancement and noise echo cancellation modules for further processing. This ensures that even with abnormal user input, high-quality, clean user speech signals can be separated, thereby improving the recognition success rate of the "make a latte" command.
[0139] Several preset processing methods are available, including:
[0140] When the volume rise rate is detected to exceed a preset threshold, the microphone channel gain is adjusted to prevent the initial audio signal from being overloaded.
[0141] When the speech rate is detected to exceed the preset ratio, the time-frequency analysis parameters of the initial audio signal are adjusted to match the fast speech.
[0142] When a pitch fluctuation range exceeds a preset multiple, the spectrum of the initial audio signal is smoothed.
[0143] When a user's input is characterized by a volume increase rate exceeding a preset threshold, it typically indicates that the user may have suddenly raised the volume due to emotional excitement or a noisy environment. In this case, to prevent clipping distortion or overload of the initial audio signal acquired by the microphone array, which could lead to loss of voice information, the gain of the microphone channel needs to be dynamically adjusted. Gain adjustment aims to control the signal amplitude within the effective input range of the microphone and analog-to-digital converter, ensuring signal integrity.
[0144] When a user's speaking speed exceeds a preset limit, it may indicate that the user is in a state of urgency or has a habitually fast speaking speed. Traditional speech recognition systems may experience a decrease in recognition accuracy when processing fast-paced speech due to a mismatch between the feature extraction window or the model training data. Therefore, it is necessary to adjust the time-frequency analysis parameters of the initial audio signal, such as shortening the frame length, increasing the frame shift, or adjusting the size of the Fourier transform window, to better capture rapidly changing speech features and thus match the characteristics of fast speech.
[0145] When a user's pitch fluctuation is detected to exceed a preset multiple, it may reflect significant emotional fluctuations or unusual pronunciation habits. Excessive pitch fluctuations can destabilize the fundamental frequency information of the speech signal, affecting the robustness of the acoustic model. Smoothing the spectrum of the initial audio signal, such as through low-pass filtering or mean filtering, can effectively suppress sharp peaks and drastic fluctuations in the spectrum, making the spectral characteristics more stable and thus reducing the negative impact of drastic pitch changes on speech recognition.
[0146] like Figure 5 As shown in the illustration, a specific embodiment of this application also discloses a voice interaction system for a coffee machine, comprising:
[0147] The acoustic space modeling module is used to detect and listen to ambient background sounds in the kitchen space when the coffee machine is not interacting with voice, so as to build and dynamically update the acoustic space model of the kitchen environment. The acoustic space model of the kitchen environment at least represents the physical location and acoustic characteristics of noise sources in the kitchen space, as well as the sound reflection path of sound propagation in the kitchen space.
[0148] The sound source localization module is used to accurately locate the user's sound source based on the sound reflection path represented in the acoustic space model of the kitchen environment when a user's voice command is detected, so as to obtain the position of the user's mouth.
[0149] The signal processing module is used to perform the following two processes in parallel on the mixed audio signal containing user voice commands: First, spatial enhancement is performed on the speech signal component originating from the user's mouth position in the mixed audio signal according to the user's mouth position; Second, active noise cancellation is performed on the noise signal component originating from the physical position of the noise source in the mixed audio signal according to the physical position and acoustic characteristics of the noise source and the sound reflection path represented in the acoustic spatial model of the kitchen environment, and active echo cancellation is performed on the echo signal component introduced by the sound reflection path, so as to separate the pure user voice signal from the mixed audio signal.
[0150] The speech recognition module is used to perform speech recognition on clean user speech signals;
[0151] The execution control module is used to execute the corresponding coffee-making instructions based on the voice recognition results.
[0152] This system aims to address the issue of poor voice interaction performance in the complex acoustic environment of a kitchen, a problem inherent in traditional coffee machines. Through a modular design, the system works collaboratively. First, an acoustic space modeling module constructs and dynamically updates the acoustic space model of the kitchen environment, providing foundational data for subsequent precise sound source localization and signal processing. Then, the sound source localization module uses this model to accurately locate the user's mouth. The signal processing module performs spatial enhancement, active noise cancellation, and active echo cancellation on the mixed audio signals in parallel, efficiently separating the clean user voice signal. Finally, the voice recognition module recognizes the clean speech, and the execution control module executes the corresponding coffee-making instructions, thus providing a smooth and accurate voice interaction experience in a noisy kitchen environment.
[0153] The steps and technical details of the coffee machine voice interaction method have been described in detail in the above embodiments, and will not be repeated here. It should be emphasized that the coffee machine voice interaction system proposed in this application achieves this by assigning these functions to specific modules, thereby forming an efficient and robust voice interaction architecture.
[0154] Specifically, the acoustic space modeling module can be understood as a component responsible for environmental perception and model building. Its implementation can include: dedicated hardware circuitry integrated within the coffee machine, such as a digital signal processor (DSP) or microcontroller, working with a microphone array and sound wave transmitter to execute acoustic detection and environmental sound monitoring algorithms; or, the module can run as software on the coffee machine's general-purpose processor, using underlying hardware interfaces to complete data acquisition and model updates. This module operates continuously to ensure the real-time performance and accuracy of the acoustic space model of the kitchen environment.
[0155] The sound source localization module is responsible for accurately locating the user's sound source using a pre-constructed acoustic space model upon detecting a user's voice command. This module can consist of a dedicated sound source localization chip or software running specific algorithms. For example, algorithms based on Time Difference of Arrival (TDOA) or Angle of Arrival (DOA), combined with sound reflection path information from the acoustic space model, can be used to calculate the precise position of the user's mouth in three-dimensional space. Its implementation can run on a separate processor or share computing resources with the signal processing module.
[0156] The signal processing module is the core processing unit of this system, responsible for optimizing the mixed audio signals in various ways. This module is typically implemented using a high-performance DSP or Field-Programmable Gate Array (FPGA) to meet the requirements of parallel processing and real-time performance. Its functions include spatial enhancement using super-directional beamforming based on the user's mouth position, and active noise cancellation and active echo cancellation based on an acoustic spatial model of the kitchen environment. These processes can be accomplished by a series of cascaded or parallel digital filters, adaptive algorithms, and gain controllers.
[0157] The speech recognition module is responsible for converting the processed, clean user speech signal into text commands. This module can employ an embedded speech recognition engine, implemented either by running a pre-trained neural network model on a dedicated AI acceleration chip or by running a lightweight speech recognition algorithm on the coffee machine's main control chip. The performance of this module directly impacts the accuracy of understanding user commands.
[0158] The execution control module is the system's command execution center. It is responsible for converting the commands output by the voice recognition module into specific hardware operation commands, such as controlling the start / stop and parameter adjustments of components like the water pump, heater, and grinder. This module ensures that voice commands are accurately and promptly translated into actual coffee-making actions.
[0159] Compared with existing technologies, the advantages of this system are as follows: First, the acoustic spatial modeling module can comprehensively perceive and characterize the acoustic environment of the kitchen, including the location of noise sources, acoustic characteristics, and sound reflection paths, thus laying the foundation for subsequent precise processing. Second, the sound source localization module can accurately locate the user's sound source based on this model, effectively distinguishing between real speech and reflected sound, solving the problem of insufficient accuracy of traditional sound source localization in multipath environments. Third, the signal processing module employs spatial enhancement, active noise cancellation, and active echo cancellation technologies in parallel. These technologies all rely heavily on the information provided by the acoustic spatial model, making the removal of interference from mixed audio signals more thorough and accurate. Spatial enhancement significantly improves the signal-to-noise ratio of the user's speech through super-directional beamforming, while active noise cancellation and active echo cancellation precisely process the noise source and sound reflection path, respectively. This integrated and refined processing flow results in a purer user speech signal, greatly improving the accuracy and robustness of the speech recognition module. Therefore, this system ensures that the coffee machine provides a smooth and accurate voice interaction experience even in a noisy kitchen environment, significantly improving user convenience and satisfaction.
[0160] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical, or other forms.
[0161] Furthermore, in the various embodiments of the present invention, each functional module can be fully integrated into a processor, or each module can be a separate device, or two or more modules can be integrated into a device; each functional module in the various embodiments of the present invention can be implemented in hardware or in the form of hardware plus software functional units.
[0162] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by program instructions and related hardware. The aforementioned program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, they perform the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0163] It should be understood that the use of terms such as "system," "device," "unit," and / or "module" in this application is merely one method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0164] As indicated in this application, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. An element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.
[0165] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0166] If a flowchart is used in this application, it is used to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0167] The above provides a detailed description of a voice interaction method and system for a coffee machine provided by the present invention. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice interaction method for a coffee machine, characterized in that, Includes the following steps: When the coffee machine is not interacting with the user via voice, the system detects the kitchen space and listens to ambient background sounds to construct and dynamically update an acoustic space model of the kitchen environment. The acoustic space model of the kitchen environment at least characterizes the physical location and acoustic characteristics of noise sources within the kitchen space, as well as the sound reflection path of sound propagation within the kitchen space. When a user's voice command is detected, the user's sound source is accurately located based on the sound reflection path represented in the acoustic space model of the kitchen environment, and the user's mouth position is obtained. The following two processes are performed in parallel on the mixed audio signal containing the user's voice command: First, based on the user's mouth position, spatial enhancement is performed on the speech signal component originating from the user's mouth position in the mixed audio signal to increase the proportion of the user's speech in the mixed audio signal; Second, based on the physical location and acoustic characteristics of the noise source represented in the acoustic spatial model of the kitchen environment and the sound reflection path, active noise cancellation is performed on the noise signal component originating from the physical location of the noise source in the mixed audio signal, and active echo cancellation is performed on the echo signal component introduced by the sound reflection path, so as to remove noise and echo from the mixed audio signal and separate the pure user speech signal; The clean user voice signal is subjected to speech recognition, and the corresponding coffee making instructions are executed based on the speech recognition results.
2. The voice interaction method for a coffee machine according to claim 1, characterized in that, The process of detecting the kitchen space and listening to ambient background sounds when the coffee machine is not interacting with the user via voice, in order to construct and dynamically update the acoustic space model of the kitchen environment, includes the following steps: When the coffee machine is not interacting with voice, it emits detection sound waves into the kitchen space and receives the reflected signals of the detection sound waves after they are reflected by objects in the kitchen space through a microphone array. The propagation time, intensity, and phase difference of the reflected signal upon arrival at the microphone array are analyzed to identify the location of the reflective surface and its acoustic reflection characteristics within the kitchen space, and an acoustic reflection path map is constructed based on this for subsequent localization and echo cancellation. The microphone array is used to monitor ambient background sounds, identify and record the acoustic characteristics of fixed noise sources in the environment, and determine the location of the fixed noise sources based on the differences in the received sounds. The location and acoustic reflection characteristics of the reflective surface and the acoustic characteristics of the fixed noise source are integrated into the acoustic space model of the kitchen environment to realize the construction or dynamic updating of the acoustic space model of the kitchen environment.
3. The voice interaction method for a coffee machine according to claim 2, characterized in that, When a user's voice command is detected, the user's sound source is accurately located based on the sound reflection path represented in the acoustic space model of the kitchen environment to obtain the user's mouth position, including the following steps: Based on the detected user voice command, the arrival time difference of the user voice command to the microphone array is calculated by the sound source localization algorithm to obtain at least one candidate sound source direction; By combining the acoustic reflection path diagram in the acoustic space model of the kitchen environment, the false sound source direction generated by the reflecting surface in the direction of the candidate sound source is identified; The false sound source directions are eliminated from the candidate sound source directions, and the remaining directions are determined as the user's true sound source directions. The precise position of the user's mouth in three-dimensional space is calculated based on the true sound source directions.
4. The voice interaction method for a coffee machine according to claim 3, characterized in that, The step of spatially enhancing the speech signal component originating from the user's mouth position in the mixed audio signal, based on the user's mouth position, specifically involves: Based on the precise position of the user's mouth in three-dimensional space, the microphone array's pickup pattern is adjusted to perform super-directional beamforming, thereby enhancing the speech signal originating from the user's mouth direction and suppressing sound from other directions.
5. The voice interaction method for a coffee machine according to claim 4, characterized in that, After spatially enhancing the speech signal component originating from the user's mouth position in the mixed audio signal based on the user's mouth position, the method further includes the following steps: During the duration of the user's voice command, the movement of the user's mouth position is tracked in real time; When a change in the user's mouth position is detected, the beamforming weighting coefficient and delay parameter of the spatial enhancement are recalculated based on the changed user's mouth position. Furthermore, the filtering coefficients for active echo cancellation are recalculated based on the change in the sound reflection path caused by the change in the user's mouth position.
6. The voice interaction method for a coffee machine according to claim 1, characterized in that, Before detecting the user's voice command, the following steps are also included: Monitor user activity and ambient sound fluctuations near the coffee machine; When abnormal user activity is detected and the sound environment is in a highly fluctuating state, or when the recognition confidence of the continuously received user voice commands is lower than a preset threshold, it is determined that a potential interaction difficulty situation has been entered; wherein, the abnormal user activity refers to the frequency or amplitude of the user's movement in front of the coffee machine exceeding the preset range during normal interaction. In the context of potential interaction difficulties, non-intrusive prompts are initiated, and the speech recognition strategy is adjusted to improve the fault tolerance rate. The non-intrusive prompts refer to visual or auditory prompts that are issued without interrupting the user's current operation.
7. The voice interaction method for a coffee machine according to claim 6, characterized in that, The step of adjusting the speech recognition strategy to improve the fault tolerance rate includes one or more of the following methods: Reduce the matching degree between input features and standard features required by the acoustic model during the speech recognition process; In the language model upon which the speech recognition relies, the decoding priority of core instruction words related to coffee making is increased; The confidence threshold of the speech recognition result is lowered from a first preset value to a second preset value; When the confidence level of the speech recognition is lower than the second preset value for several consecutive times, guided interaction is initiated, providing the user with guidance to confirm or re-enter the command through voice questions or interface options.
8. The voice interaction method for a coffee machine according to claim 2, characterized in that, Before performing parallel processing on the mixed audio signals, the following steps are also included: When a user's voice command is detected, the non-linguistic features of the initial audio signal carrying the user's voice command, which is directly acquired by the microphone array, are analyzed in real time; wherein, the non-linguistic features include volume change rate, speech rate, and pitch change. When the user's input method deviates from the normal mode based on the non-linguistic features, the initial audio signal is adaptively preprocessed to obtain the preprocessed signal, and the preprocessed signal is used as the mixed audio signal for subsequent parallel processing. The deviation from the normal mode refers to one of the following situations: the volume rise rate exceeds a preset threshold, the speech rate exceeds a preset ratio, or the pitch fluctuation range exceeds a preset multiple; The adaptive preprocessing specifically involves selecting and executing a corresponding processing method from a set of preset processing methods based on the category of the detected deviation from the normal mode.
9. The voice interaction method for a coffee machine according to claim 8, characterized in that, The preset processing methods include: When the volume rise rate is detected to exceed a preset threshold, the microphone channel gain is adjusted to prevent the initial audio signal from being overloaded. When the speech rate is detected to exceed a preset ratio, the time-frequency analysis parameters of the initial audio signal are adjusted to match the fast speech. When the detected pitch fluctuation range exceeds a preset multiple, the spectrum of the initial audio signal is smoothed.
10. A voice interaction system for a coffee machine, characterized in that, include: An acoustic space modeling module is used to detect and listen to ambient background sounds in the kitchen space when the coffee machine is not interacting with voice, so as to build and dynamically update the acoustic space model of the kitchen environment. The acoustic space model of the kitchen environment at least represents the physical location and acoustic characteristics of noise sources in the kitchen space, as well as the sound reflection path of sound propagation in the kitchen space. The sound source localization module is used to accurately locate the user's sound source based on the sound reflection path represented in the acoustic space model of the kitchen environment when a user's voice command is detected, so as to obtain the position of the user's mouth. The signal processing module is used to perform the following two processes in parallel on the mixed audio signal containing the user's voice command: First, spatial enhancement is performed on the speech signal component originating from the user's mouth position in the mixed audio signal according to the user's mouth position; Second, active noise cancellation is performed on the noise signal component originating from the physical position of the noise source in the mixed audio signal according to the physical position and acoustic characteristics of the noise source represented in the acoustic spatial model of the kitchen environment and the sound reflection path, and active echo cancellation is performed on the echo signal component introduced by the sound reflection path, so as to separate the pure user voice signal from the mixed audio signal. The speech recognition module is used to perform speech recognition on the clean user speech signal; The execution control module is used to execute the corresponding coffee-making instructions based on the voice recognition results.
Citation Information
Patent Citations
Voice screening method of range hood
CN109994123A
Voice processing method, voice processing device, computer equipment and storage medium
CN110232916A
Display device, server and voice instruction recognition method
CN118283339A
Full-automatic chef machine control method based on voice interaction
CN120015030A
Apparatus and method for processing speech in noise environment
KR1020120102306A