Voice interaction system, voice interaction method and water cup

By employing phased monitoring and edge-cloud collaborative processing, combined with microphone arrays and lightweight models, the accuracy and power consumption issues of voice interaction systems in water cup environments were resolved, achieving efficient user voice data recognition and interaction.

CN120977302AInactive Publication Date: 2025-11-18SHENZHEN RUNHENG SMART CITY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510947020.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a water cup environment, the voice interaction system suffers from poor accuracy in listening to and recognizing user voice data due to its compact structure, complex environment, and limited computing resources, which affects the user's interactive experience.

Method used

A phased monitoring approach is adopted to listen to environmental voice data. User voice data is processed by combining end-cloud collaboration modules and edge modules. Microphone arrays and beamforming algorithms are used to improve monitoring accuracy. Lightweight models and cloud resources are used for data processing, and power consumption management is optimized.

Benefits of technology

It improves the accuracy of user voice data recognition and interaction, reduces power consumption limitations, and enhances the user's interaction experience with the water cup.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977302A_ABST
    Figure CN120977302A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of audio data processing, and provides a voice interaction system, a voice interaction method and a water cup, and the system comprises a voice recognition module which is used for monitoring environment voice data of an environment where the water cup is located through a staged monitoring mode, and recognizing user voice data from the environment voice data; the end-cloud collaboration module is used for sending the user voice data to the edge end module or sending the user voice data to the cloud end module; the edge end module is used for processing the user voice data to obtain a first processing result; the voice feedback module is used for receiving the voice processing result and playing voice based on the voice processing result; wherein the voice processing result comprises a first processing result, and / or a second processing result fed back by the cloud module based on the user voice data. The accuracy of interaction between the user and the water cup through voice can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of audio data processing, and particularly relates to a voice interaction system, a voice interaction system and a cup. BACKGROUND

[0002] With the rapid development of artificial intelligence and Internet of Things technologies, voice interaction has become an important means for intelligent devices to improve user experience and realize natural human-computer interaction. For example, introducing smooth and intelligent voice interaction capabilities in intelligent devices such as intelligent cups can remind users to drink water in a timely manner and improve user experience.

[0003] At present, although voice interaction systems have been widely used in intelligent speakers, mobile phones and other devices, when applied to cups, due to the complexity of the cup use environment and the limitation of computing resources, voice interaction between the cup and the user is relatively difficult. SUMMARY

[0004] The embodiments of the application provide a voice interaction system, device and electronic equipment, which can improve the accuracy of user interaction with the cup through voice.

[0005] In a first aspect, the embodiments of the application provide a voice interaction system applied to a video monitoring system including a plurality of monitoring devices, and the voice interaction system includes:

[0006] A voice recognition module is configured to listen to environmental voice data of an environment in which a cup is located through a staged listening manner, and recognize user voice data from the environmental voice data; wherein the power consumption of different stages in the staged listening manner is different.

[0007] An edge-cloud collaboration module is configured to send the user voice data to an edge module or send the user voice data to a cloud module.

[0008] An edge module is configured to process the user voice data to obtain a first processing result.

[0009] A voice feedback module is configured to receive a voice processing result and perform voice playing based on the voice processing result; wherein the voice processing result includes the first processing result and / or a second processing result fed back by the cloud module based on the user voice data.

[0010] In a second aspect, the embodiments of the application provide a voice interaction method, including:

[0011] Listening to environmental voice data of an environment in which a cup is located through a staged listening manner, and recognizing user voice data from the environmental voice data; wherein the power consumption of different stages in the staged listening manner is different.

[0012] sending the user voice data to an edge module, or sending the user voice data to a cloud module;

[0013] processing the user voice data according to the edge module to obtain a first processing result;

[0014] receiving a voice processing result and playing voice based on the voice processing result; wherein the voice processing result comprises the first processing result and / or a second processing result fed back by the cloud module based on the user voice data.

[0015] In a third aspect, an embodiment of the present application provides a water cup, comprising: a cup body and a detachable base, wherein the detachable base is provided with the voice interaction system of the first aspect.

[0016] In a fourth aspect, an embodiment of the present application provides a detachable base, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of the voice interaction system of the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the voice interaction system of the first aspect.

[0018] In a sixth aspect, an embodiment of the present application provides a computer program product, and when the computer program product is executed on the detachable base, the detachable base implements the steps of the voice interaction system of the first aspect.

[0019] Compared with the prior art, the embodiment of the present application has the following beneficial effects:

[0020] In the embodiment of the present application, the voice recognition module listens to the environmental voice data of the environment where the cup is located through a phased listening manner, and recognizes the user voice data from the environmental voice data. Since the power consumption corresponding to the above different phase listening manners is different, the listening manner with different power consumption can be used to listen to the environmental voice data, thereby improving the applicability and accuracy of the environmental voice data listening, and further improving the accuracy of the user voice data recognition; the edge-cloud cooperation module sends the user voice data to the edge module or sends the user voice data to the cloud module, so that the user voice data can be task allocated according to the actual situation, thereby improving the accuracy of the user voice data processing; the edge module processes the user voice data to obtain a first processing result, and then the voice feedback module performs voice playing based on the received first processing result or the second processing result fed back by the cloud module, thereby completing the interaction with the user voice data. Therefore, the above voice interaction system can improve the accuracy of the user interaction through voice and cup. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 is a structural schematic diagram of a voice interaction system provided by an embodiment of the present application;

[0023] Figure 2 is a flowchart of a voice interaction method provided by an embodiment of the present application;

[0024] Figure 3 is a structural schematic diagram of a detachable base provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.

[0026] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0027] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0028] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0029] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0031] With the development of artificial intelligence and Internet of Things technologies, voice interaction has been widely used in devices such as smart speakers and smartphones. However, when these technologies are directly applied to water cups, which have special structures, complex environments, and limited resources, the user's interaction experience with the water cup is often unsatisfactory, mainly due to the following aspects:

[0032] First, as a frequently used utensil, water cups are typically placed on a table for use. There is a considerable distance between the user's mouth and the water cup on the table (typically 0.5 to 1.5 meters, or even further). The environment in which the water cup is located (such as the tabletop environment) has a large number of sound reflection surfaces (including the tabletop, walls, monitor, etc.), and may also include various noises (such as the sound of air conditioning running, keyboard typing, lights, etc.). Therefore, it is quite difficult to listen to and recognize the user's voice data.

[0033] Secondly, in order to enable interaction between the user and the water cup, the water cup often integrates a microphone (for collecting user voice data) and a speaker (for playing voice feedback or prompts). Due to the compact structure of the water cup, the microphone and speaker are often very close to each other, the acoustic coupling path is short and direct, and there may be phenomena such as structural vibration transmission, nonlinear distortion, and near-end noise interference, which makes echo cancellation more difficult than in devices such as mobile phones and smart speakers, resulting in poor accuracy of user voice data recognition.

[0034] Finally, since the primary function of a water cup is water storage, the performance of its processor, the size of its memory, and the amount of its storage space are all strictly limited, making it difficult to accurately process user voice data. In addition, the water cup needs to be able to respond to user voice requests in a timely manner, which means that the water cup needs to maintain an "always-on" listening state. Therefore, there are also more stringent requirements for low power consumption when listening to voice data.

[0035] Therefore, when traditional voice interaction systems are applied to water cup scenarios, the accuracy of user interaction with the water cup via voice is poor, which seriously affects the user's interaction experience with the water cup.

[0036] To improve the synchronization of audio data between different control systems, this application provides a voice interaction system. In this system, a voice recognition module listens to the ambient voice data of the environment in which the water cup is located through a phased listening method, identifies user voice data from the ambient voice data, and an edge-cloud collaboration module sends the user voice data to an edge module. If the edge module returns an error, the user voice data is sent to a cloud module. The edge module processes the user voice data to obtain a first processing result. Finally, a voice feedback module plays the voice based on the first processing result or the second processing result returned by the cloud module.

[0037] Figure 1This illustration shows a structural diagram of a voice interaction system provided in an embodiment of this application. The voice interaction system includes a voice recognition module 11, an end-to-cloud collaboration module 12, an edge module 13, a voice feedback module 14, and a cloud module. The cloud module and the end-to-cloud collaboration module 12 are communicatively connected, and the cloud module and the edge module 13 are also communicatively connected. The voice recognition module 11, the end-to-cloud collaboration module 12, the edge module 13, and the voice feedback module 14 can be deployed on a water cup or its base. The base can be a detachable base. The detachable base may further include sensors (e.g., a temperature sensor, a humidity sensor, a motion sensor, and a cup recognition unit), a microphone array, a communication unit (e.g., Wi-Fi), and an output unit (e.g., a speaker, an LED light, and a vibration motor).

[0038] The voice recognition module 11 is used to listen to the environmental voice data of the environment in which the water cup is located through a phased listening method, and to identify the user's voice data from the environmental voice data; wherein, the power consumption of different stages of the phased listening method is different.

[0039] It should be understood that a microphone array may be provided on the side or base of the water cup. The layout of the microphone array may be designed based on beamforming algorithms and connected through a digital interface, thereby reducing structural obstruction and vibration coupling of the microphone array and improving the accuracy of monitoring and acquiring environmental voice data.

[0040] For example, multiple microphones (e.g., MEMS microphones) with high signal-to-noise ratios (e.g., SNR>65dB) and high acoustic overload points (e.g., AOP>120dB SPL) can be installed on the base of the water cup and connected via digital interfaces (e.g., PDM or I2S) to reduce analog signal interference.

[0041] It should be understood that the aforementioned phased monitoring method includes monitoring methods at different stages, with the power consumption of the monitoring device increasing sequentially in each stage. That is, a low-power monitoring method is used first, and then, according to the monitoring needs, a higher-power monitoring method is used to monitor the aforementioned environmental voice data. The aforementioned user voice data is voice data that reflects user needs. For example, the aforementioned user voice data may include: "How much water did you drink today?", "Set a water drinking reminder", "What is the current water temperature?", etc.

[0042] Specifically, the speech recognition module 11 can first continuously monitor the ambient speech data of the environment where the water cup is located using a low-power monitoring method. When the monitored ambient speech data may contain user speech activity, a second-stage monitoring method is activated. If it is determined that the ambient speech data monitored in the second stage contains user speech activity, a third-stage monitoring method is activated to obtain real-time ambient speech data. Then, the first speech recognition algorithm is used to recognize the user's speech data from the real-time ambient speech data. Optionally, the first speech recognition algorithm can be a lightweight speech recognition model (Automatic Speech Recognition, ASR) to recognize user speech data with low power consumption.

[0043] In addition, targeted low-power technologies can be used for each stage of the monitoring method, including but not limited to: Dynamic Voltage and Frequency Scaling (DVFS), gated clock, memory hibernation, and algorithm optimization (such as using fixed-point arithmetic and model quantization), thereby further reducing power consumption.

[0044] In this embodiment, ambient voice data is monitored by a monitoring method with gradually increasing power consumption. A low-power monitoring method can continuously monitor the ambient voice data of the environment where the water cup is located, thereby responding to the user's voice activities in a timely manner. When the user makes a voice activity, a higher-power monitoring method is activated to obtain real-time ambient voice data. While ensuring continuous monitoring, ambient voice data that may contain user voice activities can be processed more accurately, thereby improving the accuracy of user voice data recognition.

[0045] The edge-cloud collaboration module 12 is used to send the aforementioned user voice data to the edge module, or to send the aforementioned user voice data to the cloud module.

[0046] Specifically, the aforementioned edge-cloud collaboration module 12 can use a pre-trained request classifier to classify the user voice data. When the classification result instructs the edge module to process the data, the user voice data is sent to the edge module; when the classification result instructs the cloud module to process the data, it is sent to the cloud module via the network. The pre-trained request classifier can be a neural network model that classifies data based on keywords, syntactic structure, or preliminary intent. Furthermore, to improve the timeliness of voice interaction, the user voice data can be sent to the edge module first, and then sent to the cloud module only when the edge module returns an error.

[0047] For example, when the classification result instructs the cloud module to process the data, the user's voice data can be sent to the cloud module via HTTPS or a secure WebSocket call to its API interface. The cloud module then uses its acoustic model (e.g., an ASR engine) to further improve the quality of the voice data, and finally performs deep inference using a large language model (e.g., LLM) in the cloud to obtain the second processing result. This cloud module can be a cloud-based intelligent platform, including but not limited to Google Cloud AI, AWS AI, Azure AI, or other self-built platforms.

[0048] In this embodiment, sending user voice data to an edge module or a remote module improves the flexibility of user voice data processing. Furthermore, the cloud module has more powerful computing resources than the edge module, which can further improve the accuracy of user voice data processing.

[0049] Edge module 13 is used to process the above user voice data to obtain the first processing result.

[0050] Specifically, the edge module 13 can process the aforementioned user voice data using a lightweight language model, then extract key information (e.g., user intent, keywords, etc.), execute dialogue management logic directly related to local functions, and output the corresponding first processing result. The lightweight language model can be a language model customized for a specific domain (e.g., water cup control, simple health queries) (e.g., a small N-gram or neural network language model), or it can be a rule-based intent recognition or lightweight machine learning model (e.g., SVM, FastText, etc.).

[0051] In this embodiment, the edge module can respond to user voice data in a timely manner, thereby improving the user experience.

[0052] The voice feedback module 14 is used to receive the voice processing result and play the voice based on the voice processing result; wherein the voice processing result includes the first processing result or the second processing result fed back by the cloud module based on the user's voice data.

[0053] It should be understood that a speaker may be installed on the side or bottom of the water cup. The above-mentioned speech processing results may include processed audio streams or text.

[0054] Specifically, the aforementioned voice feedback module 14 can receive voice processing results from the edge module and / or the cloud module, then convert them into speech to be played through voice conversion technology, and play them through a speaker. The voice conversion technology can be speech synthesis technology (e.g., Text-to-Speech, TTS). It should be noted that when the voice processing result includes both the first and second processing results, playback can be prioritized according to a preset priority. For example, the first processing result from the edge module can be played first, or a specific type of voice processing result (e.g., remaining water volume) can be played first.

[0055] In this embodiment, the voice recognition module listens to the ambient voice data of the water cup's environment in a phased listening manner and identifies the user's voice data from the ambient voice data. Since the power consumption corresponding to the different listening methods in each phase is different, different power consumption listening methods can be used to listen to the ambient voice data, improving the applicability and accuracy of ambient voice data listening, and thus improving the accuracy of user voice data recognition. The edge-cloud collaboration module sends the user's voice data to the edge module or to the cloud module, thereby allocating tasks based on the actual situation and improving the accuracy of user voice data processing. The edge module processes the user's voice data to obtain a first processing result, and then the voice feedback module plays the voice based on the received first processing result or the second processing result fed back by the cloud module, completing the interaction with the user's voice data. Therefore, the above voice interaction system can improve the accuracy of user interaction with the water cup via voice.

[0056] In some embodiments, the above-mentioned speech recognition module includes:

[0057] The first monitoring module is used to monitor the environmental voice data of the environment where the water cup is located in real time.

[0058] The second monitoring module is used to detect whether the above-mentioned environmental voice data includes a preset wake word;

[0059] The third monitoring module is used to identify the user's voice data from the environmental voice data when the preset wake word is detected.

[0060] Specifically, the first monitoring module is used to execute the first stage of monitoring, continuously monitoring the ambient voice data of the environment where the water cup is located through a preset low-power detection method. Then, when the monitored ambient voice data may contain user voice activity, the second monitoring module is activated. The second monitoring module is used to execute the second stage of monitoring, identifying whether the current ambient voice data includes a preset wake word through a wake word model. If the preset wake word is not detected, the detection continues, or the monitoring returns to the first stage. If the preset wake word is detected, the third monitoring module is activated. The third monitoring module identifies the user voice data from the ambient voice data through a preset audio front-end processing algorithm stack.

[0061] The aforementioned preset low-power detection methods may include: monitoring using an analog or digital sound activity detector (SAD) with extremely low power consumption (e.g., microamp level), or monitoring using a minimalist hardware / software (e.g., voice activity detection software (VAD)). The aforementioned wake word engine may be a low-power wake word engine (WWE), for example, a model based on convolutional neural networks (CNN), deep neural networks (DNN), or a more lightweight structure. The aforementioned preset audio front-end processing algorithm stack may be an algorithm stack integrating data augmentation algorithms, noise processing algorithms, echo cancellation algorithms, etc.

[0062] It should also be understood that the aforementioned different monitoring modules can precisely control the power state of each hardware module (such as MCU core frequency, peripheral clock, sensor power supply, wireless module, etc.) through a real-time operating system (RTOS) or state machine to perform power management, that is, activate the corresponding function only when woken up, thereby achieving precise power consumption management.

[0063] In this embodiment, by using a phased monitoring method and phased power management, low-power data monitoring and high-power data processing can be achieved, reducing the conflict between power consumption limitations and "always-on" monitoring, and improving the real-time performance and accuracy of user voice data monitoring.

[0064] In some embodiments, the third monitoring module includes:

[0065] The data enhancement module is used to enhance the speech data in the target direction of the above-mentioned environmental speech data to obtain enhanced environmental speech data; wherein, the target direction is the direction pointed to by the microphone array set in the water cup.

[0066] The noise processing module is used to remove noise from the enhanced environmental speech data to obtain optimized environmental speech data.

[0067] The liveness detection module is used to identify the user's voice data from the optimized environment voice data.

[0068] The target direction mentioned above refers to the direction in which the microphone array points towards the user, such as the direction corresponding to the sector formed by the microphone array. The target direction can be a specific direction set by the microphone array (at least 2), such as the sector direction facing the handle or drinking spout; or, the target direction can be a direction automatically identified by beamforming algorithms (such as fixed beam, adaptive beamforming (such as General Eigenvalue Decomposition (GEV), Minimum Variance Distortionless Response (MVDR)) etc.) and sound source localization (SSL) algorithms.

[0069] Specifically, the aforementioned data augmentation module can enhance speech data from the target direction in the environmental speech data using beamforming algorithms, while suppressing speech data from non-target directions, to obtain augmented environmental speech data. Then, the aforementioned noise processing module separates environmental noise data from the augmented environmental speech data and performs mixing and echo cancellation to obtain optimized environmental speech data. Finally, the liveness detection module identifies user speech data from the optimized environmental speech data using a liveness detection model (e.g., a VAD model). Furthermore, to further improve the quality of user speech data, automatic gain control (AGC) can be used for volume variation compensation, and speech signal enhancement technology can be used for speech signal enhancement processing.

[0070] In this embodiment, since the environment in which the water cup is located may contain various noises, data augmentation is performed first, followed by noise cancellation through a noise processing module, and finally, the user's voice data is identified through a liveness detection module, which can improve the accuracy of user voice data recognition.

[0071] In some embodiments, the noise processing module includes:

[0072] The first noise processing module is used to separate the target environmental noise from the enhanced environmental speech data according to the noise separation model to obtain processed environmental speech data; wherein, the target environmental noise is the ambient sound data of the environment in which the water cup is located.

[0073] The second noise processing module is used to eliminate environmental reflection noise and environmental echo noise in the above-mentioned processed environmental speech data to obtain the above-mentioned optimized environmental speech data.

[0074] The noise separation model described above is trained based on the target noise data. The target environmental noise includes, but is not limited to: keyboard typing, mouse clicking, object placement sounds, water flow sounds, air conditioning sounds, and outdoor traffic sounds. The environmental reflection noise mentioned above refers to the reverberation noise caused by sound reflections from surfaces such as desktops, walls, and monitors in the environment where the cup is located. The environmental echo noise mentioned above refers to the echo noise generated when the propagation path of the user's voice data is affected by various factors such as the shape of the desktop and the cup itself.

[0075] Specifically, the first noise processing module takes the enhanced environmental speech data as input to the noise separation model and outputs the data separated from the target environmental noise to obtain the processed environmental speech data; then the second noise processing module eliminates the environmental reflection noise in the environmental speech data through the reverberation cancellation model and eliminates the environmental echo noise through the acoustic echo canceller (AEC) algorithm to obtain the optimized environmental speech data.

[0076] Optionally, the noise separation model described above can be a deep learning-based model, including but not limited to: network structures based on RNN, CNN, Transformer, or combinations thereof, such as a real-time speech enhancement neural network model (Dual-SignalTransformation LSTM Network, DTLN), a deep learning model combining convolutional neural networks (CNN) and recurrent neural networks (RNN) (CRNN), and other joint models. The reverberation cancellation model described above can be an algorithm capable of estimating the room impulse response (RIR) and performing inverse filtering, or a DNN-based dereverberation model.

[0077] It should be understood that the above noise separation model can be trained through the following steps: First, acquire training data, which includes a large amount of target noise data (e.g., keyboard typing, mouse clicking, object placement, water flow, common household appliance noise, etc.) and real or simulated data under different reverberation conditions; then, perform data augmentation on the training data (e.g., add noise with different signal-to-noise ratios, simulate reverberation in different rooms); use the data-augmented training data to train a deep neural network to obtain the trained model; then, quantize (e.g., INT8), prune, and convert the trained model into a format suitable for efficient operation in the first noise processing module (e.g., TensorFlow Lite Micro, ONNX Runtime Mobile, etc.) to obtain the noise separation model, thereby reducing computational load and memory usage. The training process of the above DNN-based dereverberation model is similar and will not be repeated here.

[0078] In this embodiment, a complete processing pipeline is formed by integrating beamforming, echo cancellation (AEC), and reverberation cancellation models, which can optimize the computation process and improve the accuracy of user voice data recognition.

[0079] In some embodiments, the system further includes:

[0080] The data acquisition module is used to acquire the status data of the water cup;

[0081] The aforementioned edge module further includes: using a first language model to identify the user intent in the aforementioned user voice data, filling a preset local intent template according to the aforementioned user intent and the aforementioned water cup status data, and obtaining the aforementioned first processing result, wherein the aforementioned first language model is a lightweight model set in the aforementioned edge module.

[0082] The aforementioned water cup status data refers to data reflecting the real-time status of the water cup, including water volume, water temperature, current state (e.g., picked up or put down), ambient temperature, and ambient humidity. The aforementioned preset local intent templates refer to the dialogue templates corresponding to the locally associated dialogue management logic, including temperature intent templates, drinking intent templates, etc. The aforementioned preset local intent templates include different slots (e.g., time, water volume, symptom description, health concept, etc.) for filling.

[0083] Specifically, the water cup or water cup base can be equipped with different types of sensors (such as gravity sensors, temperature sensors, humidity sensors, etc.). The data acquisition module acquires the data collected by the sensors in real time to obtain the water cup status data. Then, the edge module identifies the user intent through the first language model, finds the corresponding local intent template based on the user intent, and fills the local intent template with the water cup status data corresponding to the user intent to obtain the first processing result.

[0084] For example, if the first language model recognizes the user's intent in the above user voice data as "Is the water still hot?", the system finds the corresponding local intent template and fills the corresponding template with the current temperature data to get "The current water temperature is 40°C, it can be drunk directly"; or, if the first language model recognizes the user's intent in the above user voice data as "I am thirsty", the system finds the corresponding local intent template and combines it with the current water volume to get "You still have 200ml of warm water in your cup, which is just enough to replenish it".

[0085] In this embodiment, the first language model and the locally associated local intent template can respond to user needs in a timely manner and improve the processing efficiency of user voice data.

[0086] In some embodiments, the system further includes:

[0087] The aforementioned data acquisition module is also used to acquire user status data;

[0088] The aforementioned end-to-cloud collaboration module is also used to send the aforementioned water cup status data and the aforementioned user status data to the aforementioned cloud module;

[0089] The aforementioned voice feedback module is also used to receive the status processing results fed back by the aforementioned cloud module based on the aforementioned water cup status data and / or the aforementioned user status data, and to play voice based on the aforementioned status processing results.

[0090] The aforementioned user status data may include information such as age, weight, health goals, allergy history, set drinking plans, exercise data, and sleep reports.

[0091] Specifically, the aforementioned data acquisition module can acquire user status data uploaded by the user through an application (such as an app on a mobile phone), or it can acquire user status data from authorized devices (such as fitness trackers). The aforementioned end-to-cloud collaboration module sends the aforementioned water cup status data and user status data to the aforementioned cloud module, so that the models in the cloud module can iterate and update based on the aforementioned water cup status data and user status data, and then output the status processing result determined based on the water cup status data and / or user status data through the updated model. For example, the cloud module inputs health goals and current water volume into a second language model (such as GPT series, Claude, Gemini, etc. large language models), and outputs the status processing result "According to your weight loss plan, it is recommended to drink water before meals" and "It was detected that you exercised a lot today, please pay attention to increasing your water intake." It should also be noted that the cloud module can continuously optimize the dialogue process based on the aforementioned water cup status data and user status data, thereby further improving the accuracy of processing user voice data. For example, when a user says, "I have a headache today," the cloud module not only records it but can also ask follow-up questions such as, "Is it persistent or intermittent? Are there any other discomforts? Did you drink enough water today?" thereby continuously guiding the user and providing accurate personalized suggestions.

[0092] In this embodiment, the cloud module can enrich the user profile by combining water cup status data and user status data, thereby improving the accuracy of cloud module processing and enhancing the user interaction experience.

[0093] In some embodiments, the above-mentioned end-to-cloud collaboration module further includes:

[0094] The encryption module is used to encrypt the aforementioned user voice data and user status data, and then send the encrypted data to the aforementioned cloud module.

[0095] Specifically, to prevent user data leakage and improve user data security, user voice data and user status data can be encrypted using a preset lightweight encryption algorithm before being transmitted to the cloud module. Simultaneously, the cloud module can anonymize or de-identify the aforementioned user voice data and user status data during processing or transmission to prevent the leakage of user privacy. For example, the preset lightweight encryption algorithm could be AES-128 encryption, PRESENT encryption, etc.

[0096] In some embodiments, the system further includes:

[0097] The synchronization module is used to synchronize the data between the aforementioned edge module and the aforementioned cloud module.

[0098] Specifically, after the edge module and / or cloud module are updated, the aforementioned synchronization module can synchronize the updated content to the cloud module and / or edge module based on the model context synchronization mechanism. For example, when the edge module determines a change in user state (e.g., the user has just finished drinking water), it can synchronize this change to the cloud module. After updating the user profile, the cloud module can synchronize the updated user profile to the edge module, thereby guiding the update of the first language model in the edge module. Of course, since the computing resources of the edge module are relatively limited, the synchronization module can synchronize the data from the cloud module to the edge module only after receiving the update instruction sent by the edge module. The aforementioned model context synchronization mechanism can be the Model Context Protocol (MCP).

[0099] In this embodiment, by synchronizing the data of the edge module and the cloud module, the performance of the edge module and the cloud module can be continuously improved, thereby improving the accuracy of user voice data processing.

[0100] Corresponding to the voice interaction system described in the above embodiments, Figure 2 A flowchart illustrating a voice interaction method provided in an embodiment of this application is shown. This method can be applied to a water cup, and is described in detail below:

[0101] S21. The ambient voice data of the environment where the water cup is located is monitored by a phased monitoring method, and the user voice data is identified from the ambient voice data; wherein, the power consumption of different phases of the phased monitoring method is different.

[0102] S22. Send the user voice data to the edge module, or send the user voice data to the cloud module.

[0103] S23. The user voice data is processed by the edge module to obtain a first processing result.

[0104] S24. Receive the voice processing result and play the voice based on the voice processing result; wherein, the voice processing result includes the first processing result and / or, the second processing result fed back by the cloud module based on the user's voice data.

[0105] In this embodiment, environmental voice data of the water cup's environment is monitored in stages, and user voice data is identified from this environmental voice data. Since the power consumption of each stage of the monitoring method differs, monitoring methods with varying power consumption can be used to monitor environmental voice data, improving the applicability and accuracy of environmental voice data monitoring, and consequently, improving the accuracy of user voice data recognition. The user voice data is then sent to an edge module or a cloud module. Task allocation for the user voice data can be performed based on actual conditions, improving the accuracy of user voice data processing. The edge module processes the user voice data to obtain a first processing result. Based on the received first processing result or the second processing result fed back by the cloud module, voice playback is performed, completing the interaction with the user's voice data. Therefore, the above voice interaction method can improve the accuracy of user interaction with the water cup via voice.

[0106] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0107] Corresponding to the voice interaction system described in the above embodiments, this application embodiment also provides a water cup, including: a cup body and a detachable base, wherein the detachable base is provided with the above-mentioned voice interaction system.

[0108] Figure 3 This is a schematic diagram of the structure of a detachable base provided in one embodiment of this application. Figure 3 As shown, the detachable base 3 of this embodiment includes: at least one processor 30 ( Figure 3 Only one is shown in the diagram. A memory 31 and a computer program 32 stored in the memory 31 and executable on the at least one processor 30 are also shown. When the processor 30 executes the computer program 32, it implements the steps of any of the modules disposed in the detachable base 3.

[0109] The detachable base may include, but is not limited to, the processor 30 and the memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of the detachable base 3 and does not constitute a limitation on the detachable base 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, the detachable base may also include an input transmitting device, a network access device, a bus, etc.

[0110] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0111] In some embodiments, the memory 31 may be an internal storage unit of the removable base 3, such as a hard drive or memory of the removable base 3. The memory 31 may also be an external storage device of the removable base 3, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card mounted on the removable base 3. Furthermore, the memory 31 may include both internal storage units and external storage devices of the removable base 3. The memory 31 is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of computer programs. The memory 31 can also be used to temporarily store data that has been sent or will be sent.

[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the division of the functional units and modules is only described as an example. In practical applications, the functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0113] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the various method embodiments.

[0114] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments.

[0115] This application provides a computer program product that, when run on a detachable base, enables the detachable base to perform the steps described in the various method embodiments.

[0116] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the embodiments described in this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the camera / detachable base, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0117] In the embodiments described, each embodiment has its own emphasis. For parts not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0118] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0119] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0121] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A voice interaction system, characterized in that, include: The voice recognition module is used to listen to the environmental voice data of the environment in which the water cup is located through a phased listening method, and to identify the user's voice data from the environmental voice data; wherein, the power consumption of different stages of the phased listening method is different. The edge-cloud collaboration module is used to send the user's voice data to the edge module, or to send the user's voice data to the cloud module; An edge module is used to process the user's voice data and obtain a first processing result; A voice feedback module is used to receive voice processing results and play voice based on the voice processing results; wherein, the voice processing results include the first processing result and / or, the second processing result fed back by the cloud module based on the user's voice data.

2. The voice interaction system as described in claim 1, characterized in that, The speech recognition module includes: The first monitoring module is used to monitor the environmental voice data of the environment in which the water cup is located in real time. The second monitoring module is used to detect whether the environmental voice data includes a preset wake word; The third monitoring module is used to identify the user's voice data from the environmental voice data when the preset wake word is detected.

3. The voice interaction system as described in claim 2, characterized in that, The third monitoring module includes: The data enhancement module is used to enhance the speech data in the target direction of the environmental speech data to obtain enhanced environmental speech data; wherein, the target direction is the direction pointed to by the microphone array set in the water cup; A noise processing module is used to remove noise from the enhanced environmental speech data to obtain optimized environmental speech data. A liveness detection module is used to identify the user's voice data from the optimized environment voice data.

4. The voice interaction system as described in claim 3, characterized in that, The noise processing module includes: The first noise processing module is used to separate the target environmental noise from the enhanced environmental speech data according to the noise separation model to obtain processed environmental speech data; wherein, the target environmental noise is the ambient sound data of the environment in which the water cup is located; The second noise processing module is used to eliminate environmental reflection noise and environmental echo noise in the processed environmental speech data to obtain the optimized environmental speech data.

5. The voice interaction system as described in any one of claims 1-4, characterized in that, Also includes: The data acquisition module is used to acquire the status data of the water cup; The edge module further includes: using a first language model to identify user intent in the user voice data, filling a preset local intent template according to the user intent and the water cup status data, and obtaining the first processing result, wherein the first language model is a lightweight model set in the edge module.

6. The voice interaction system as described in claim 5, characterized in that, Also includes: The data acquisition module is also used to acquire user status data; The end-to-cloud collaboration module is also used to send the water cup status data and the user status data to the cloud module; The voice feedback module is also used to receive the status processing result fed back by the cloud module based on the water cup status data and / or the user status data, and to play voice based on the status processing result.

7. The voice interaction system as described in claim 6, characterized in that, The endpoint-cloud collaboration module also includes: An encryption module is used to encrypt the user's voice data and the user's status data, and send the encrypted data to the cloud module.

8. The voice interaction system as described in claim 6 or 7, characterized in that, Also includes: A synchronization module is used to synchronize data between the edge module and the cloud module.

9. A voice interaction method, characterized in that, include: The ambient voice data of the water cup's environment is monitored in a phased manner, and the user's voice data is identified from the ambient voice data; wherein, the power consumption of different phases of the phased monitoring method is different. The user voice data is sent to the edge module, or the user voice data is sent to the cloud module; The user voice data is processed by the edge module to obtain a first processing result; The system receives voice processing results and plays voice content based on the voice processing results; wherein the voice processing results include the first processing result and / or the second processing result fed back by the cloud module based on the user's voice data.

10. A water cup, characterized in that, include: The cup body and the detachable base, wherein the detachable base is provided with the voice interaction system according to any one of claims 1-8.