AI voice dialogue system and method based on single 4G module

By integrating main control, audio processing, and communication functions into a single 4G module, a simplified end-to-end cloud collaborative hardware architecture is constructed, solving the stability and cost issues of AI voice devices in environments without high-speed networks. This achieves low-latency, low-power voice interaction, making it suitable for portable terminals.

CN121815233APending Publication Date: 2026-04-07AIWO TECHNOLOGY (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing AI voice devices suffer from unstable connections, high latency, high hardware costs, high power consumption, and large size in environments without high-speed network coverage, making them unsuitable for portable devices or low-cost terminals.

Method used

It adopts a single 4G module to integrate main control, audio processing and communication functions, and builds a simplified terminal network cloud collaborative hardware architecture. Combined with the wide coverage of 4G network, it can achieve stable and low-latency voice interaction.

Benefits of technology

Achieve stable AI voice interaction in environments without high-speed networks, reduce hardware costs by 40%-60%, shrink device size by 50%, achieve high voice recognition accuracy, smooth response, low power consumption and long battery life, making it suitable for portable terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815233A_ABST
    Figure CN121815233A_ABST
Patent Text Reader

Abstract

The invention discloses an AI voice dialogue system based on a single 4G module. The AI voice dialogue system comprises the single 4G module, a cloud AI voice service platform and a peripheral interaction component, wherein the single 4G module integrates a main control sub-module, an audio processing sub-module and a communication sub-module; the single 4G module can independently complete the functions of communication, voice preprocessing, data transmission, instruction analysis, peripheral interaction component driving and the like, and an additional chip is not needed. The design characteristic is that a practical end-side AI function can be realized under extremely low power consumption, and the method is suitable for Internet of Things equipment and AI intelligent terminals (AIoT) which are sensitive to cost and power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI voice interaction and wireless communication technology, specifically to an AI voice dialogue system and method based on a single 4G module, applicable to scenarios such as smart homes, industrial control, and portable terminals that require intelligent voice interaction in environments without high-speed networks. Background Technology

[0002] With the rapid development of AI voice technology, voice interaction has become one of the core interaction methods for smart devices, widely used in smart homes, industrial control, smart terminals, and many other fields. However, existing mainstream AI voice devices suffer from the following technical shortcomings: 1. Strong network dependence. Most existing AI voice devices rely on high-speed networks such as WiFi for cloud data transmission and processing. In remote areas, industrial plants, and outdoor mobile scenarios where there is no high-speed network coverage, problems such as unstable connection, high latency, and poor interactive experience exist. 2. Complex hardware architecture. To address the issue of reliance on high-speed networks, existing technologies employ multiple modules, such as 4G / 5G + WiFi. While this approach improves network compatibility, it suffers from drawbacks such as high hardware costs, high power consumption, and large size, making it unsuitable for portable devices or low-cost terminals.

[0003] For example, Chinese patent application CN111402882A discloses an intelligent voice module and its operating method, including: at least one microphone, a voice module, and a WiFi module, all mounted on a PCB board; at least one microphone, used to send a voice command to the voice module after acquiring a voice command; the voice module, used to send the voice command from the microphone to the WiFi module; and the WiFi module, used to send the voice command to a target device when it receives a voice command from the voice module, so that the target device performs the operation corresponding to the voice command.

[0004] Chinese patent application CN112331212A discloses a voice control system and method for smart devices, relating to the field of smart device control technology. It addresses how to effectively utilize mobile networks to compensate for the impact of unstable home Wi-Fi signals on the voice control of smart devices. The system includes: a controller, a voice cloud, a device cloud, and a smart device. The controller acquires voice commands and transmits them to the voice cloud via a 5G mobile network hotspot. The voice cloud identifies the control intent corresponding to the voice command and sends the intent to the device cloud. The device cloud generates control commands based on the control intent and sends them to the smart device. The smart device receives the control commands and executes their corresponding control operations. Utilizing the 5G mobile network hotspot to transmit voice commands solves the problem of unstable home Wi-Fi signals affecting the voice control of smart devices.

[0005] Therefore, how to achieve stable and low-latency AI voice dialogue interaction in environments without high-speed network coverage, while ensuring low cost, low power consumption, and miniaturization, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0006] To address the aforementioned problems in existing technologies, this invention provides an AI voice dialogue system and method based on a single 4G module. It integrates three core functions: main control, 4G communication, and audio processing. It eliminates the need for additional independent chips and constructs a simplified end-to-end cloud collaborative hardware architecture of "module-cloud platform-peripheral components." By combining the wide coverage of 4G networks with AI voice technology, it achieves efficient and stable voice interaction.

[0007] This invention provides an AI voice dialogue system based on a single 4G module, comprising: The system includes a single 4G module, a cloud-based AI voice service platform, and peripheral interactive components; wherein the single 4G module integrates a main control submodule, an audio processing submodule, and a communication submodule. The main control submodule is used to coordinate the linkage of various functional modules, make decisions, and issue control commands. The audio processing submodule is used to receive the raw voice signal collected by the peripheral interaction component, preprocess the raw voice signal, and transmit the preprocessed voice signal to the communication submodule; the audio processing submodule is also used to drive the peripheral interaction component to respond to the control commands of the main control submodule; The communication submodule is used to upload data from a single 4G module to a cloud-based AI voice service platform via 4G communication, and to receive data transmitted from the cloud-based AI voice service platform to the single 4G module. The cloud-based AI voice service platform is used to receive voice data transmitted from the single 4G module, and to perform voice recognition, semantic understanding and dialogue logic generation, and output structured dialogue response results.

[0008] Furthermore, the cloud-based AI voice service platform includes an automatic speech recognition engine, a natural language understanding engine, a dialogue management module, and a text-to-speech engine.

[0009] Furthermore, the communication submodule 30 has 4G communication functionality and supports LTE-CAT1 or LTE-CAT4 communication standards.

[0010] Furthermore, the communication submodule includes a network status monitoring module, which is used to monitor the signal strength, signal-to-noise ratio and network speed of the single 4G module in real time, and to initiate a disconnection reconnection mechanism when the network is interrupted, and to automatically resume the communication link after recovery.

[0011] Furthermore, the communication submodule includes a data transmission optimization module that supports data fragmentation transmission and adapts to the module's data processing bandwidth.

[0012] Furthermore, the preprocessed speech signal is converted from an analog speech signal to a 16kHz sampling rate, 16-bit bit-depth digital signal through encoding and decoding to achieve data standardization; and then compressed.

[0013] This invention also provides an AI voice dialogue method based on a single 4G module, which mainly includes the following steps: The single 4G module completes network registration and parameter configuration through its own communication interface and protocol, and establishes a two-way data connection with the cloud AI voice service platform; The single 4G module acquires the user's voice signal through peripheral interaction components, and preprocesses the voice signal through a built-in audio processing algorithm to obtain standardized voice data. The single 4G module compresses the standardized voice data and uploads it to the cloud AI voice service platform at a set frame rate using low-latency data transmission capabilities. The communication submodule of the single 4G module receives the structured dialogue response result data transmitted back from the cloud AI voice service platform. The single 4G module parses the dialogue response result data through its built-in decoding module and drives the peripheral interactive components to respond.

[0014] Furthermore, the communication submodule includes a data transmission optimization module that supports data fragmentation transmission and adapts to the module's data processing bandwidth.

[0015] Furthermore, the preprocessed speech signal is encoded and decoded to convert the analog speech signal into a 16kHz sampling rate, 16-bit bit-depth digital signal, thereby achieving data standardization.

[0016] The equipment and method described in this invention have the following beneficial effects: 1. Solved the problem of dependence on high-speed networks: Relying on the wide coverage of 4G networks, stable AI voice interaction can still be achieved in remote areas, outdoor scenarios, and other scenarios without high-speed networks such as WiFi, thus expanding the application scope of AI voice technology.

[0017] 2. Simplified hardware architecture: It adopts a single 4G module, integrating three core functions, eliminating the need for independent main control chips and signal processing chips, reducing hardware costs by 40%-60% and device size by more than 50%, thus reducing hardware costs and device size and adapting to the needs of portable terminals and low-cost, miniaturized terminals.

[0018] 3. Enhanced interactive performance: Through integrated data processing and optimized transmission solutions within the module, the risk of data latency and loss is reduced. Combined with lightweight encoding and segmented transmission technology, transmission efficiency is further improved, ensuring real-time voice interaction. The voice recognition accuracy can reach 95% in quiet environments and 83% in normal environments, with smooth response.

[0019] 4. Lowering the barrier to entry and providing strong scalability: It adopts a standardized interface design, supports custom configuration by manufacturers, and allows manufacturers to add their own voice content, device control command library and frequently asked questions. Hardware assembly and software settings are simple, making it easy for mass production and integrated applications.

[0020] 5. Low power consumption and long battery life: The integrated design reduces power consumption for data interaction between chips. A single 4G module paired with a 1200mAh lithium battery can support no less than 5 hours of continuous use. It has low standby power consumption and is suitable for portable device needs. Attached Figure Description

[0021] To gain a more complete understanding of the invention, reference will now be made to the following description taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the structure of an AI voice dialogue system based on a single 4G module according to Embodiment 1 of the present invention. Detailed Implementation

[0022] To clearly illustrate the purpose, technical details, and effective applications of this invention, and to facilitate understanding and implementation by those skilled in the art, a further detailed description will be provided below in conjunction with the embodiments and accompanying drawings. Obviously, the embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention. Example 1

[0023] Embodiment 1 of the present invention provides an AI voice dialogue system based on a single 4G module, see [link to documentation]. Figure 1 The system specifically includes: a main control submodule 10, an audio processing submodule 20, and a communication submodule 30; the main control submodule 10, the audio processing submodule 20, and the communication submodule 30 are all integrated within a single 4G module 100; the system also includes a cloud-based AI voice service platform 200 and peripheral interaction components 300.

[0024] The single 4G module 100 can independently perform functions such as communication, voice preprocessing, data transmission, command parsing, and driving peripheral interactive components without the need for additional chips. Its design feature is the ability to achieve practical edge AI functions with extremely low power consumption, making it suitable for cost- and power-sensitive IoT devices and AI smart terminals (AIoT). The single 4G module 100 preferably uses the EG800Z 4G module. Based on the characteristics of this module, this invention can accurately solve the pain points of voice interaction in weak coverage scenarios, while simplifying hardware design, reducing costs, and improving operation and maintenance efficiency.

[0025] The single 4G module 100's main processor adopts a dual-core processor architecture with a main frequency of up to 306MHz. It integrates a Vector Processing Unit (VPU) for processing various audio, video, and AI algorithms. It also integrates vector instruction extensions to support INT8 / INT16 / INT32 integer operation acceleration, which is the main driving force for its AI computing. It can perform high-performance, low-latency multimedia and edge AI computing processing. Furthermore, it can complete core processing functions such as Automatic Noise Reduction (ANS) and AEC echo cancellation (AEC) in high-definition voice calls and AI voice interaction.

[0026] The main control submodule 10 is used to coordinate the linkage of various functional modules, make decisions and issue control commands, and at the same time handle abnormal states and feedback interactions.

[0027] The audio processing submodule 20 can collect raw voice signals through devices such as the microphone of the peripheral interaction component 300, and can drive devices such as the speaker of the peripheral interaction component 300 to play voice.

[0028] The user triggers the system's voice acquisition function through dialogue, and the microphone acquires the user's original voice signal; the audio processing submodule 20 collects the original voice signal and performs noise reduction processing on the original voice signal through built-in audio processing algorithms, including echo cancellation and signal amplification, to obtain a pre-processed voice signal.

[0029] The audio processing submodule 20 transmits the preprocessed voice signal to the communication submodule 30.

[0030] The communication submodule 30 is used for data transmission between the single 4G module 100 and the cloud AI voice service platform 200.

[0031] The communication submodule 30 has 4G communication capabilities and supports LTE-CAT1 or LTE-CAT4 communication standards.

[0032] The communication submodule 30 converts the preprocessed analog voice signal into a 16kHz sampling rate, 16-bit depth OPUS format digital signal to achieve data standardization and compress it.

[0033] The compressed voice signal is uploaded to the cloud AI voice service platform 200 at a set frame rate and high priority via 4G communication, utilizing low-latency data transmission capabilities; simultaneously, it is used to receive data transmitted from the cloud AI voice service platform 200.

[0034] Furthermore, the communication submodule 30 also includes a network status monitoring module, which is used to collect the signal strength (RSRP), signal-to-noise ratio (SINR) and network rate parameters of a single 4G module in real time, and to start a disconnection reconnection mechanism when the network is interrupted, and to automatically resume the voice dialogue link after recovery.

[0035] Furthermore, the communication submodule 30 also includes a data transmission optimization module, which adopts the OPUS lightweight data compression algorithm to meet the real-time requirements of voice data; supports data fragmentation transmission, splits large voice data packets into small data packets, reduces data loss rate, and adapts to the module's data processing bandwidth.

[0036] The cloud-based AI voice service platform 200 is the core functional unit, comprising an Automatic Speech Recognition (ASR) engine, a Natural Language Understanding (NLU) engine, a dialogue management module, and a Text-to-Speech (TTS) engine, responsible for speech recognition, semantic understanding, and dialogue response generation. The ASR engine supports real-time recognition of Mandarin Chinese and other languages, with an accuracy rate of no less than 95%. The NLU engine, based on a deep learning model, can parse user intents such as device control, information query, and task execution, and extract key parameters. The dialogue management module can generate logically coherent response strategies based on user intent and context.

[0037] After receiving the compressed standardized voice data uploaded by the single 4G module 100, the cloud-based AI voice service platform 200 converts the received data into text using its ASR engine, and its NLU engine performs semantic analysis on the text content to parse user intent and key parameters. The dialogue management module calls the corresponding business logic to generate response text based on the user intent. The TTS engine synthesizes the response text into voice data of a specified style and transmits it back to the single 4G module 100 via the 4G network.

[0038] The cloud-based AI voice service platform 200 has a built-in device control command library, frequently asked questions and answers, and supports manufacturers to add custom voice content, device control command library and frequently asked questions and answers, which can adapt to the personalized needs of different scenarios.

[0039] The communication submodule 30 receives the voice data and transmits it to the main control submodule 10 for decoding; the main control submodule 10 drives the speakers and other devices of the peripheral interaction component 300 to play the voice through the audio processing submodule 20, thus completing an AI voice interaction.

[0040] The peripheral interaction component 300 includes a high-sensitivity microphone for voice acquisition, a power amplifier speaker for voice playback, and an LED indicator for 3-state prompts.

[0041] The single 4G module 100 provided by this invention integrates the main control submodule 10, the audio processing submodule 20, and the communication submodule 30. Through the integrated data processing and optimized transmission scheme within the module, the risk of data delay and loss is reduced. Combined with lightweight encoding and segmented transmission technology, the transmission efficiency is further improved, ensuring the real-time performance of voice interaction. The voice recognition accuracy can reach 95% in a quiet environment and 83% in a normal environment, with smooth response.

[0042] When the system provided by this invention is working, it is necessary to ensure that the single 4G module 100 and the peripheral interaction components 300 are in a working state. Specifically, after the single 4G module 100 is powered on, it completes 4G network search and registration according to its preset startup sequence. The single 4G module completes network registration and parameter configuration through its own communication interface and protocol, establishes a two-way data connection with the cloud AI voice service platform, and synchronously initializes peripheral components to ensure that the module and components are time-matched.

[0043] The bidirectional data connection is established using the WebSocket protocol.

[0044] The system provided by this invention can be easily embedded into various terminal devices such as dolls and desk lamps with a single 4G module, exhibiting strong hardware compatibility. Multi-scenario testing of the system showed excellent results. When the mechanism box is embedded in a doll device, a child saying "Little helper, sing a nursery rhyme" will cause the device to start playing the nursery rhyme within 5 seconds; saying "Stop playing" will immediately stop playback, demonstrating accurate voice recognition and timely response. When the mechanism box is embedded in a desk lamp device, a user saying "Little helper, turn on the desk lamp" will successfully start the lamp; saying "Turn it brighter" will automatically increase the light brightness, resulting in smooth and lag-free interaction. Example 2

[0045] Based on the system described in Embodiment 1, the present invention also provides an AI voice dialogue method based on a single 4G module, which mainly includes the following steps: Step S01: The single 4G module completes network registration and parameter configuration through its own communication interface and protocol, and establishes a two-way data connection with the cloud AI voice service platform; Step S02: The single 4G module acquires the user's voice signal through the peripheral interaction component, and preprocesses the voice signal through the built-in audio processing algorithm to obtain standardized voice data.

[0046] Step S03: The single 4G module compresses the standardized voice data and uploads it to the cloud AI voice service platform at a set frame rate using low-latency data transmission capabilities.

[0047] Step S04: The communication submodule of the single 4G module receives the structured dialogue response result data transmitted back by the cloud AI voice service platform.

[0048] Step S05: The single 4G module parses the dialogue response result data through its built-in decoding module and drives the peripheral interaction components to respond.

[0049] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the description of the embodiments above. Therefore, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the system claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. An AI voice dialogue system based on a single 4G module, comprising: The system includes a single 4G module, a cloud-based AI voice service platform, and peripheral interactive components; wherein the single 4G module integrates a main control submodule, an audio processing submodule, and a communication submodule. The main control submodule is used to coordinate the linkage of various functional modules, make decisions, and issue control commands. The audio processing submodule is used to receive the raw voice signal collected by the peripheral interaction component, preprocess the raw voice signal, and transmit the preprocessed voice signal to the communication submodule; the audio processing submodule is also used to drive the peripheral interaction component to respond to the control commands of the main control submodule; The communication submodule is used to upload data from the single 4G module to the cloud AI voice service platform via 4G communication, and to receive data transmitted from the cloud AI voice service platform to the single 4G module. The cloud-based AI voice service platform is used to receive voice data transmitted from the single 4G module, perform voice recognition, semantic understanding and dialogue logic generation, output structured dialogue response results, and transmit them to the single 4G module.

2. The system according to claim 1, characterized in that, The cloud-based AI voice service platform includes an automatic speech recognition engine, a natural language understanding engine, a dialogue management module, and a text-to-speech engine.

3. The system according to claim 1, characterized in that, The communication submodule has 4G communication capabilities and supports LTE-CAT1 or LTE-CAT4 communication standards.

4. The system according to claim 1, characterized in that, The communication submodule includes a network status monitoring module, which is used to monitor the signal strength, signal-to-noise ratio and network speed of the single 4G module in real time, and to activate the disconnection reconnection mechanism when the network is interrupted, and to automatically resume the communication link after recovery.

5. The system according to claim 1, characterized in that, The communication submodule includes a data transmission optimization module, which supports data fragmentation transmission and adapts to the module's data processing bandwidth.

6. The system according to claim 1, characterized in that, The preprocessed speech signal is converted from an analog speech signal to a 16kHz sampling rate, 16-bit bit depth digital signal through encoding and decoding to achieve data standardization; and then compressed.

7. An AI voice dialogue method based on a single 4G module, mainly including the following steps: The single 4G module completes network registration and parameter configuration through its own communication interface and protocol, and establishes a two-way data connection with the cloud AI voice service platform; The single 4G module acquires the user's voice signal through peripheral interaction components, and preprocesses the voice signal through a built-in audio processing algorithm to obtain standardized voice data. The single 4G module compresses the standardized voice data and uploads it to the cloud AI voice service platform at a set frame rate using low-latency data transmission capabilities. The communication submodule of the single 4G module receives the structured dialogue response result data transmitted back from the cloud AI voice service platform. The single 4G module parses the dialogue response result data through its built-in decoding module and drives the peripheral interactive components to respond.

8. The method according to claim 7, characterized in that, The communication submodule includes a data transmission optimization module, which supports data fragmentation transmission and adapts to the module's data processing bandwidth.

9. The method according to claim 7, characterized in that, The preprocessed speech signal is encoded and decoded to convert the analog speech signal into a 16kHz sampling rate, 16-bit bit-depth digital signal, thereby achieving data standardization.

Citation Information

Patent Citations

  • Intelligent voice module and operation method thereof

    CN111402882A

  • Intelligent equipment voice control system and method

    CN112331212A