Voice interaction device, voice interaction method and electronic equipment
By using a millimeter-wave radar sensing module and a gradually brightening/dimming LCD display, the problems of false wake-up and high power consumption in existing voice interaction devices have been solved, achieving natural interaction and low power consumption management, and improving the user experience of the device.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-03
AI Technical Summary
Existing voice interaction devices are susceptible to interference in complex acoustic environments, resulting in false wake-up or wake-up failure, unnatural interaction, high power consumption, and poor coordination, failing to achieve natural interaction and efficient power consumption management.
Millimeter-wave radar is used as the sensing module to achieve contactless wake-up. Combined with the gradual brightening and dimming control of the voice processing module and the LCD display module, the main control module coordinates the management of wake-up, interaction and sleep states, and utilizes offline voice recognition and ambient light sensors to optimize the interactive experience.
It enhances the naturalness and convenience of interaction, reduces system standby power consumption, achieves low power management in unattended states, and provides a smooth closed-loop interactive experience.
Smart Images

Figure CN121789672A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of voice interaction technology, specifically relating to a voice interactive device, a voice interaction method, and an electronic device. Background Technology
[0002] With the rapid development of artificial intelligence and the Internet of Things (IoT) technologies, voice interaction has become a core interaction method in smart homes, smart offices, and commercial displays. Current mainstream voice interaction devices, such as smart speakers and smart terminals with screens, typically rely on keyword wake-up or physical touch for activation. However, in complex acoustic environments, wake words are easily interfered with, leading to false wake-ups or wake-up failures; users must memorize specific commands or find touch points, making the interaction less intuitive and natural.
[0003] Some high-end devices have attempted to improve the user experience by introducing auxiliary sensors (such as infrared or cameras), but shortcomings remain. For example, infrared sensors cannot detect stationary human bodies, while cameras raise privacy concerns and are greatly affected by lighting conditions. Meanwhile, the display functions of existing screen-equipped devices are often independent of the wake-up logic, failing to achieve intelligent linkage with the user's approach or departure, resulting in unnecessary energy consumption and visual interference. Millimeter-wave radar technology has attracted attention in the sensing field due to its advantages such as strong penetration, high accuracy, no optical privacy issues, and the ability to detect minute movements. However, current technologies mostly use it for simple motion-triggered switches, failing to deeply integrate it into a complete system-level solution aimed at natural interaction, especially lacking refined collaborative management with voice and display modules in terms of power consumption control and response timing.
[0004] Therefore, in order to address the aforementioned technical problems, it is necessary to provide a voice interactive device, a voice interactive method, and an electronic device.
[0005] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a voice interactive device, voice interactive method, and electronic device that can solve the problems of unnatural interaction triggering, high power consumption, and poor coordination in existing voice interactive devices.
[0007] To achieve the above objectives, a specific embodiment of the present invention provides the following technical solution:
[0008] A voice-interactive device includes a main control module, a sensing module, a voice processing module, a liquid crystal display module, and a power management module. The main control module is used for system control and data coordination. The sensing module is connected to the main control module and is used to detect human targets and generate a wake-up signal. The voice processing module is connected to the main control module and is used for voice signal pickup, recognition, and synthesis. The liquid crystal display module is connected to the main control module and includes a liquid crystal panel and a backlight control circuit. The power management module is connected to the main control module and the sensing module and is used to switch the device's power state according to the wake-up signal. The main control module, based on the output of the sensing module, coordinates the start / stop and operation of the voice processing module and the liquid crystal display module.
[0009] In one or more embodiments of the present invention, the backlight control circuit is a gradual brightening and dimming control circuit, used to control the backlight brightness of the liquid crystal screen to change smoothly within a preset time.
[0010] In one or more embodiments of the present invention, the sensor of the sensing module is a millimeter-wave radar, the millimeter-wave radar operates at a frequency of 60 GHz, has a detection range of 0.5-5 meters, and is capable of detecting stationary human targets.
[0011] In one or more embodiments of the present invention, the voice processing module includes an offline voice recognition unit and an online voice recognition interface, wherein the offline voice recognition unit is used to recognize predefined voice commands in the absence of a network connection.
[0012] In one or more embodiments of the present invention, the device further includes an ambient light sensor connected to the main control module, the ambient light sensor being used to collect ambient light intensity, and the main control module adjusting the backlight brightness of the liquid crystal display module according to the ambient light intensity.
[0013] A voice interaction method, applied to the aforementioned voice interactive device, the method comprising the following steps:
[0014] S1. The sensing module monitors the environment in low-power mode and triggers wake-up when a human target is detected;
[0015] S2. In response to wake-up, activate the device and initiate voice activity detection;
[0016] S3. Based on the voice activity detection results, if valid voice is found, then recognition and demand identification are performed.
[0017] S4. Generate voice and graphic feedback content based on the identification results;
[0018] S5. Collaborate on voice broadcasting and control the screen to display graphic content in a gradually brightening manner;
[0019] S6. After the termination condition is met, control the screen to dim and cause the device to return to low power mode.
[0020] In one or more embodiments of the present invention, in S3, if no valid human voice is detected within a preset time window after wake-up, the control device executes a default non-voice task or directly returns to the low-power mode.
[0021] In one or more embodiments of the present invention, during demand identification, matching is first performed in the local instruction library; if local matching fails, the identification information is uploaded to the cloud server for identification.
[0022] An electronic device includes the aforementioned voice-interactive device, and further includes at least one processor and at least one memory. The memory stores a computer program. When the computer program is executed by the processor, the aforementioned voice interaction method is implemented.
[0023] In one or more embodiments of the present invention, the electronic device further includes a modular housing having detachable independent cavities for accommodating the voice processing module and the sensing module.
[0024] Compared with existing technologies, the voice interaction device, voice interaction method and electronic device of the present invention achieve contactless and non-specific wake-up word-free intelligent perception and triggering through millimeter-wave radar, which greatly improves the naturalness and convenience of interaction. By adopting a hierarchical power management strategy with millimeter-wave radar as the primary wake-up source, the core components are in deep sleep when no one is present, resulting in extremely low system standby power consumption, which is suitable for application scenarios that require long-term power supply. The radar perception, voice interaction and visual display are deeply coordinated in terms of timing and logic through the main control, realizing a smooth closed-loop experience of "perception is wake-up, interaction is screen-on, and leaving is sleep", which enhances the information transmission effect. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the system structure of a voice interactive device according to an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of a voice interaction method in one embodiment of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the technical solutions in this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.
[0029] like Figure 1 As shown, an embodiment of the present invention provides a voice-interactive device, comprising a main control module, a sensing module, a voice processing module, a liquid crystal display module, and a power management module. The main control module is used for system control and data coordination. The sensing module is connected to the main control module and is used to detect human targets and generate wake-up signals. The voice processing module is connected to the main control module and is used for voice signal pickup, recognition, and synthesis. The liquid crystal display module is connected to the main control module and includes a liquid crystal panel and a backlight control circuit. The power management module is connected to both the main control module and the sensing module and is used to switch the device's power state according to the wake-up signal. The main control module, based on the output of the sensing module, coordinates the start-up, shutdown, and operation of the voice processing module and the liquid crystal display module.
[0030] The main control module runs the operating system, schedules tasks, processes data, and manages power consumption. The sensing module operates in low-power mode, non-contactly detecting the presence, movement, and location of a human body within a preset space, and generating a wake-up signal based on the detection results. The voice processing module is activated upon receiving the wake-up signal, responsible for picking up audio, recognizing user intent, and generating voice feedback. The LCD module displays graphical information in response to main control commands and can smoothly turn the backlight on or off.
[0031] In this embodiment, the main control module uses a quad-core processor based on the ARM Cortex-M0 architecture. It communicates with the sensing module via the SPI bus to acquire raw data or preprocessed results such as target distance, speed, and angle. The millimeter-wave radar uses a 60GHz FMCW radar chip, which integrates a low-power MCU and runs presence detection and micro-motion recognition algorithms. The power management module uses a multi-output power management chip. The voice processing module uses a dedicated processing chip that supports hybrid online and offline operation. It can integrate WiFi or connect to an external WiFi module, and features high integration, powerful computing power, and excellent scalability. Its advantages include low cost, strong main control capabilities, and support for natural language interaction. The chip incorporates noise reduction, echo cancellation, AEC playback interruption, dual-microphone enhancement, and directional sound pickup functions. It also supports local voiceprint recognition and self-learning capabilities, significantly improving the voice interaction experience. Furthermore, the chip also includes a dual-microphone array audio interface, a hardware VAD, and an offline recognition engine supporting hundreds of local command words. The LCD module uses a 6.86-inch IPS LCD screen, and its backlight driving circuit integrates a constant current gradually brightening and dimming circuit controlled by the main control PWM signal.
[0032] Among them, the backlight control circuit is a gradual brightening and dimming control circuit, which is used to control the backlight brightness of the LCD screen to change smoothly within a preset time (such as 0.5-5 seconds).
[0033] In addition, the LCD module's display screen supports both standard rectangular and cylindrical physical shapes. The cylindrical screen uses flexible OLED or special curved LCD technology, combined with optical bonding technology, to achieve a surround visual experience.
[0034] Preferably, the sensor in the sensing module is a millimeter-wave radar, operating at a frequency of 60 GHz with a detection range of 0.5-5 meters, capable of detecting stationary human targets. This invention utilizes the high precision and micro-motion detection capabilities of 60 GHz millimeter-wave radar to achieve highly reliable stationary presence sensing, expanding effective interaction scenarios.
[0035] like Figure 1 As shown, the voice processing module includes an offline voice recognition unit and an online voice recognition interface. The offline voice recognition unit is used to recognize predefined voice commands in the absence of a network connection. The online voice recognition interface can connect to a cloud server to upload complex sentences for processing.
[0036] This invention ensures rapid response and privacy security for core commands through an offline speech recognition unit, and provides powerful natural language understanding capabilities by connecting to a cloud server through an online speech recognition interface.
[0037] like Figure 1As shown, the device is also equipped with an ambient light sensor, which is connected to the main control module to collect ambient light levels in real time. The main control module dynamically adjusts the backlight brightness of the LCD module based on the light data, thereby achieving adaptive adjustment of the screen brightness. This not only optimizes the visual experience under different lighting conditions but also reduces energy consumption in low-light environments, thus saving energy.
[0038] In addition, the device can obtain time and brightness information via Wi-Fi to further refine the screen display. During voice interaction, the device can recognize the user's gender and age and automatically match an appropriate brightness mode accordingly, thereby improving viewing comfort.
[0039] The device also considers time and lighting conditions, triggering corresponding voice broadcasts after recognizing voice commands in specific scenarios. For example, when receiving a voice command in the morning, the system will broadcast positive messages such as "The early bird catches the worm, keep it up, master," enhancing the pleasantness and atmosphere of the interaction.
[0040] In this example, the ambient light sensor is connected to the main control module via an I2C bus to report the ambient illuminance value in real time.
[0041] like Figure 2 As shown, a voice interaction method is applied to the aforementioned voice interactive device, and the method includes the following steps:
[0042] S1. In low-power standby mode, the environment is monitored by the millimeter-wave radar of the sensing module. When the sensing module detects a human target that meets the preset conditions, it generates and sends a wake-up signal to the power management module.
[0043] S2. The power management module responds to the wake-up signal, switches the device to the working mode, and activates the voice activity detection function of the voice processing module.
[0044] S3. Within the preset time window after activation, determine whether to perform voice interaction based on the voice activity detection results: if there is a valid human voice, perform voice pickup, recognition and demand identification; if not, perform the default non-voice task or return to standby; during voice interaction, the LCD screen will also display the corresponding information, and the home appliances will perform the corresponding actions.
[0045] S4. Based on the identification results, generate corresponding voice feedback content and graphic display content;
[0046] S5 drives the voice processing module to broadcast voice feedback and controls the LCD display module to display graphic content in a backlight gradually brightening manner. At the same time, when the touch function is used on the LCD display module, its function or purpose can also be broadcast by voice.
[0047] S6. After the preset interaction termination conditions are met, control the LCD display module to turn off the display by gradually dimming the backlight, and control the device to return to low-power standby mode.
[0048] In S1, during low-power standby mode, the main control module, voice processing module, and LCD display module are powered off, while only the sensing module and power management module operate at extremely low power. The millimeter-wave radar scans periodically at a frequency of 1Hz.
[0049] In addition, in S3, during demand identification, matching is first performed in the local instruction library; if local matching fails, the identification information is uploaded to the cloud server for identification.
[0050] An electronic device includes the aforementioned voice-interactive device, and further includes at least one processor and at least one memory. The memory stores a computer program. When the computer program is executed by the processor, the aforementioned voice interaction method is implemented.
[0051] In addition, the electronic device also includes a modular housing with detachable independent cavities for accommodating the voice processing module and the sensing module.
[0052] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0053] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0055] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0056] It will be apparent to those skilled in the art that this disclosure is not limited to the details of the exemplary embodiments described above, and that this disclosure can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of this disclosure is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this disclosure. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0057] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A voice-interactive device, characterized in that, include: The main control module is used for system control and data coordination. A sensing module, connected to the main control module, is used to detect human targets and generate a wake-up signal; The voice processing module, connected to the main control module, is used for voice signal pickup, recognition, and synthesis. A liquid crystal display module is connected to the main control module, and the liquid crystal display module includes a liquid crystal panel and a backlight control circuit. The power management module, connected to the main control module and the sensing module, is used to switch the power state of the device according to the wake-up signal. The main control module, based on the output of the sensing module, coordinates the start-up, shutdown, and operation of the voice processing module and the liquid crystal display module.
2. The voice interactive device according to claim 1, characterized in that, The backlight control circuit is a gradual brightening and dimming control circuit, used to control the backlight brightness of the LCD screen to change smoothly within a preset time.
3. The voice interactive device according to claim 1, characterized in that, The sensor in the sensing module is a millimeter-wave radar, which operates at a frequency of 60 GHz and has a detection range of 0.5-5 meters, enabling it to detect stationary human targets.
4. The voice interactive device according to claim 1, characterized in that, The voice processing module includes an offline voice recognition unit and an online voice recognition interface. The offline voice recognition unit is used to recognize predefined voice commands in the absence of a network connection.
5. A voice interactive device according to claim 1, characterized in that, The device also includes an ambient light sensor connected to the main control module. The ambient light sensor is used to collect ambient light intensity, and the main control module adjusts the backlight brightness of the liquid crystal display module according to the ambient light intensity.
6. A voice interaction method, applied to the voice interactive device according to any one of claims 1-5, characterized in that, The method includes the following steps: S1. The sensing module monitors the environment in low-power mode and triggers wake-up when a human target is detected; S2. In response to wake-up, activate the device and initiate voice activity detection; S3. Based on the voice activity detection results, if valid voice is found, then recognition and demand identification are performed. S4. Generate voice and graphic feedback content based on the identification results; S5. Collaborate on voice broadcasting and control the screen to display graphic content in a gradually brightening manner; S6. After the termination condition is met, control the screen to dim and cause the device to return to low power mode.
7. A voice interaction method according to claim 6, characterized in that, In S3, if no valid human voice is detected within the preset time window after wake-up, the control device will execute the default non-voice task or directly return to the low-power mode.
8. The voice interaction method according to claim 7, characterized in that, When identifying a requirement, the system prioritizes matching within the local command library; if local matching fails, the identification information is uploaded to the cloud server for further verification.
9. An electronic device comprising the voice interactive device according to any one of claims 1-5, characterized in that, Also includes: At least one processor; At least one memory that stores a computer program; When the computer program is executed by the processor, it implements the voice interaction method as described in any one of claims 6-8.
10. An electronic device according to claim 9, characterized in that, It also includes a modular housing with detachable, separate cavities for accommodating the voice processing module and the sensing module.