Voice processing method and device based on miniature wireless AI module, electronic equipment and storage medium
Through the voice interaction of the micro wireless AI module and multi-layer large-scale model analysis technology, the problem of complex text input is solved, convenient AI applications in multiple scenarios are realized, and user experience is improved.
Patent Information
- Application Number
- CN202510542664.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-02
AI Technical Summary
In the prior art, the use of AI requires complex text input and is difficult to easily apply in outdoor or mobile office scenarios.
The micro wireless AI module is used to collect voice stream data through voice interaction, perform offline matching and optimization processing. If it fails, it will be uploaded to the cloud-based large model for analysis, output natural language audio stream, and combine multi-layer large model for voice processing.
It realizes the rapid and convenient use of AI tools in various environments, simplifies operational processes and improves user experience.
Smart Images

Figure CN120581007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice processing method, device, electronic device and storage medium based on a micro wireless AI module. Background Art
[0002] At present, the development of AI technology has gradually penetrated into all aspects of our lives. It can not only handle complex operations, but also help us quickly analyze and locate problems in our work and life.
[0003] In related technologies, when using AI, one often needs to take out a computer or mobile phone and type in complex and detailed text to get AI assistance. This is inconvenient for some outdoor or mobile office scenarios and environments that require timely AI assistance, and the process is relatively cumbersome. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a voice processing method, device, electronic device and storage medium based on a miniature wireless AI module, which can quickly and conveniently apply AI technology in more environments without the need for complex text input and tools, but only simple voice interaction.
[0005] The object of the present invention is achieved through the following technical solutions:
[0006] The first aspect of the present application provides a voice processing method based on a micro wireless AI module, comprising: collecting voice stream data, identifying and processing the voice stream data, and outputting keyword data; matching the keyword data with offline voice data, and if the match fails, uploading the keyword data to a large cloud model for parsing and processing, and outputting a natural language audio stream.
[0007] The matching process of the keyword data with the offline voice data includes: optimizing the keyword data and outputting preprocessed keyword data; identifying the preprocessed keyword data and outputting text data; extracting keyword feature data in the text data, wherein the keyword feature data includes word domain information, part of speech information and lexical information; pairing the keyword feature data with the offline voice data, and if the pairing is successful, generating reply text information based on the keyword feature data, converting the reply text information into audio data and outputting it; if the pairing fails, uploading the keyword data to a large cloud model for parsing and outputting a natural language audio stream.
[0008] The uploading of the keyword data to the cloud-based large model for parsing and processing includes: uploading the keyword data to the first large model through an interactive protocol, the first large model performing parsing and processing on the keyword data, and outputting natural language text; receiving the natural language text, sending the natural language text to the second large model, the second large model performing parsing and processing on the natural language text, and outputting a reply text; receiving the reply text, sending the reply text to the third large model, the third large model performing conversion and processing on the reply text, and outputting the natural language audio stream.
[0009] The method further includes: receiving a voice wake-up word through an external microphone device to end the sleep state.
[0010] The second aspect of the present application provides a voice processing device based on a micro wireless AI module, including: a recognition module for collecting voice stream data, identifying and processing the voice stream data, and outputting keyword data; an output module for matching the keyword data with offline voice data. If the match fails, the keyword data is uploaded to a large cloud model for parsing and processing, and a natural language audio stream is output.
[0011] The output module includes a matching unit, which is used to optimize the keyword data and output preprocessed keyword data; identify the preprocessed keyword data and output text data; extract keyword feature data in the text data, and the keyword feature data includes word domain information, part of speech information and lexical information; pair the keyword feature data with the offline voice data, and if the pairing is successful, generate reply text information based on the keyword feature data, convert the reply text information into audio data and output it; if the pairing fails, upload the keyword data to a large cloud model for parsing and output a natural language audio stream.
[0012] The output module includes a parsing unit, which is used to optimize the keyword data and output preprocessed keyword data; identify the preprocessed keyword data and output text data; extract keyword feature data in the text data, and the keyword feature data includes word domain information, part of speech information and lexical information; pair the keyword feature data with the offline voice data, and if the pairing is successful, generate reply text information based on the keyword feature data, convert the reply text information into audio data and output it; if the pairing fails, upload the keyword data to a large cloud model for parsing and output a natural language audio stream.
[0013] It also includes a wake-up unit, which is used to receive a voice wake-up word through an external microphone device to end the sleep state.
[0014] A third aspect of the present application provides an electronic device, including:
[0015] processor; and
[0016] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.
[0017] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.
[0018] Compared with the prior art, the present invention has at least the following advantages:
[0019] This application uses a miniature wireless AI module that can help solve users' problems simply by collecting their voice. It does not require complex text input and tools, but only simple voice interaction, allowing users to use AI tools at all times, which is very convenient to operate. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments.
[0021] Figure 1 This is a flow chart of a method for voice processing based on a micro wireless AI module in one embodiment of the present invention;
[0022] Figure 2 This is a method flow chart of another implementation of a voice processing method based on a micro wireless AI module in one embodiment of the present invention;
[0023] Figure 3 This is a method flow chart of another implementation of a voice processing method based on a micro wireless AI module in one embodiment of the present invention;
[0024] Figure 4 This is a method flow chart of another implementation of a voice processing method based on a micro wireless AI module in one embodiment of the present invention;
[0025] Figure 5 This is a functional module diagram of a voice processing device based on a micro wireless AI module in one embodiment of the present invention;
[0026] Figure 6 FIG. 4 is a schematic structural diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0028] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0029] Unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," and the like should be interpreted broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; and internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0030] When using AI, you often need to take out a computer or mobile phone and type in complex and detailed text to get AI assistance. This is inconvenient for some outdoor or mobile office scenarios and environments that require timely AI help, and the process is relatively cumbersome.
[0031] To address the above problems, the embodiments of the present application provide a voice processing method, device, electronic device and storage medium based on a micro wireless AI module, which can quickly and conveniently apply AI technology in more environments without the need for complex text input and tools, but only simple voice interaction.
[0032] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0033] Figure 1 This is a flow chart of a voice processing method based on a micro wireless AI module shown in an embodiment of the present application.
[0034] See also Figure 1 , a voice processing method based on a micro wireless AI module, comprising:
[0035] Step S101: collect voice stream data, perform recognition processing on the voice stream data, and output keyword data.
[0036] It should be noted that the miniature wireless AI module of this application is only the size of a car key. It collects the user's voice stream data, where the voice stream data is generally the user's questions. It then recognizes and extracts the voice stream data and outputs the keywords in the voice stream data.
[0037] Step S102: Match the keyword data with the offline voice data. If the match fails, upload the keyword data to the cloud-based large model for parsing and outputting a natural language audio stream.
[0038] It should be noted that the micro wireless AI module has built-in offline voice data, which stores answers to some common questions. If the keyword data fails to successfully match the offline voice data, the keyword data will be uploaded to the backend cloud for analysis and finally sent back to the AI module to output a natural language audio stream. The natural language audio stream here is natural language.
[0039] See also Figure 2 , Figure 2 A more detailed implementation of a voice processing method based on a miniature wireless AI module includes:
[0040] Step S201: Collect voice stream data, perform recognition processing on the voice stream data, and output keyword data.
[0041] The description here can refer to step 101, which will not be repeated here.
[0042] Step S202: Optimize the keyword data and output pre-processed keyword data;
[0043] Perform recognition processing on the pre-processed keyword data and output text data;
[0044] Extracting keyword feature data from text data, the keyword feature data including word domain information, part of speech information, and lexical information;
[0045] Pair the keyword feature data with the offline voice data. If the pairing is successful, generate a reply text message based on the keyword feature data, convert the reply text message into audio data, and output it.
[0046] If the pairing fails, the keyword data will be uploaded to the cloud-based large model for parsing and processing, and a natural language audio stream will be output.
[0047] It should be noted that keyword data often contains noise, so optimization operations such as noise reduction and gain adjustment are first required to improve audio quality. Digital signal processing techniques, such as filtering and Fourier transform, can be used to remove noise and optimize the audio's spectral characteristics, ensuring more accurate subsequent processing. An offline speech recognition engine then recognizes the preprocessed keyword data and converts it into text. Natural language processing techniques, such as lexical analysis and part-of-speech tagging, are then used to analyze the text and identify key words. Finally, the extracted keywords are matched against those in the offline speech data. Offline speech data is typically a database or file containing various keywords and related information. String matching algorithms, such as exact matching and fuzzy matching, can be used to compare the extracted keywords with those in the offline speech data. During the matching process, considerations such as capitalization, synonyms, and near-synonyms should be taken into account to improve matching accuracy and flexibility.
[0048] If the final match is successful, the audio data is directly output. If the match is unsuccessful, the keyword data is uploaded to the cloud-based large model for parsing and processing, specifically including: uploading the keyword data to the first large model through the interactive protocol, the first large model solves the keyword data and outputs natural language text; receiving the natural language text, sending the natural language text to the second large model, the second large model parses the natural language text and outputs the reply text; receiving the reply text, sending the reply text to the third large model, the third large model converts the reply text and outputs the natural language audio stream.
[0049] It should be noted that the interaction protocol can be WebSocket, a network communication protocol. The first model can be a voice Alexa, which is used to convert keyword data into natural language text. After conversion to natural language text, it will be replied to the AI wireless module. The AI wireless module then sends the natural language text to the second model, which can be an LLM model or a DeepSeek model. After processing by the second model, the reply text is output to the AI wireless module. The AI wireless module uploads the reply text to the third model, which can be a TTS model, which converts the reply text into a personalized human natural language audio stream. The entire process is very fast, and users can interact with the AI wireless module directly through voice without having to type in complex problem descriptions.
[0050] A voice processing method based on a micro wireless AI module also includes: receiving a voice wake-up word through an external microphone device to end the sleep state.
[0051] It is understandable that the AI wireless module in dormant state can be awakened by voice wake-up words.
[0052] Furthermore, in another embodiment, the AI wireless module is limited by computing power, power consumption, and storage, and lacks the ability to detect vulnerabilities and provide real-time defense, making it vulnerable to attacks. A voice processing method based on a micro wireless AI module also includes:
[0053] Step S301: collect real-time operation data, and generate a fixed-length behavior sequence according to the real-time operation data.
[0054] It should be noted that the real-time operation data includes API call sequences, memory access logs, and network traffic packets, and a fixed-length behavior sequence S = {s1, s2, …, sn} is generated through a sliding window.
[0055] Step S302: Perform lightweight processing on the fixed-length behavior sequence and output a behavior feature vector.
[0056] It should be noted that a lightweight autoencoder can be used to compress the feature dimensions of a fixed-length behavior sequence to generate a 128-dimensional behavior feature vector, thereby reducing the amount of model calculation.
[0057] Step S303: Collect the behavior feature vector at a preset time threshold, calculate and process the behavior feature vector, and output the attack probability.
[0058] It should be noted that the preset time threshold is typically 100ms. Behavioral feature vectors Vt are collected at 100ms intervals and input into the local detection model to predict the risk level R∈[0,1]. The CVE vulnerability library and fuzz testing can be combined to annotate normal behavior and vulnerability exploitation behavior, construct a training dataset, and train the local detection model.
[0059] Step S304: When the attack probability is within a first threshold range, execute an instruction to enable memory address space randomization and output a first exception log; when the attack probability is within a second threshold range, execute an instruction to terminate the suspicious process, load a pre-stored differential hot patch to repair the vulnerable function, and output a second exception log; when the attack probability is within a third threshold range, execute an instruction to cut off the wireless connection and output a third exception log.
[0060] It should be noted that the first threshold range can be 0.5≤R<0.8. Since the first threshold range is still at a relatively safe stage, memory address space randomization can be enabled to increase the difficulty of vulnerability exploitation. The second threshold range can be 0.8≤R<0.9. At this time, the instruction to terminate the suspicious process can be executed and the pre-stored differential hot patch can be loaded to repair the vulnerable function. Differential hot patching is a software update method that combines differential technology with a hot repair mechanism. It aims to achieve vulnerability repair or function upgrade of the runtime system with a minimum amount of data. Its core principle is to generate lightweight patches by analyzing the differences between the new and old versions of the code, and dynamically load the application without restarting the device, thereby achieving rapid repair in low-power, low-bandwidth scenarios. The third threshold range can be 0..95≤R. At this time, all wireless connections will be cut off and the system will enter safe mode to prevent attacks.
[0061] Step S305: construct a local training set according to the first abnormality log, the second abnormality log, and the third abnormality log.
[0062] It should be noted that a local training set is constructed based on the first abnormality log, the second abnormality log, and the third abnormality log to further train the local detection model.
[0063] Through the above method, by setting up multi-dimensional behavioral data collection and a three-level active defense mechanism, accurate identification and rapid response to vulnerability exploitation behaviors can be achieved under low computing power and low memory conditions, significantly improving the security and reliability of the AI wireless module, making it suitable for more usage scenarios.
[0064] See Figure 4 In another embodiment, when the AI wireless module experiences network fluctuations, the data transmission failure rate is high, thereby affecting the user experience. A voice processing method based on a micro wireless AI module also includes:
[0065] Step S401: monitor network indicator data in real time, calculate and process the network indicator data using a first algorithm, and generate a network scoring model.
[0066] It should be noted that the network indicator data includes but is not limited to signal strength, signal-to-noise ratio, round-trip delay, throughput, and packet loss rate. The first algorithm can be a hierarchical analysis method. The network indicator data is calculated by the first algorithm. The specific formula is:
[0067] Q=α×RTT+β×Throughput+λ×(1-PLR)+δ×RSSI
[0068] Where Q is the network scoring model score, RTT is the round-trip delay, Throughput is the throughput, PLR is the packet loss rate, and RSSI is the signal strength. The weight coefficients α = 0.3, β = 0.4, γ = 0.2, and δ = 0.1 can be dynamically adjusted according to the task type.
[0069] Step S402: outputting a first network operation strategy, a second network operation strategy, and a third network operation strategy according to the network scoring model;
[0070] The first network operation strategy includes: operating the primary network and placing the backup network in a standby state;
[0071] The second network operation strategy includes: using the primary network and the backup network for parallel transmission, slicing the keyword data, and uploading the keyword data to the cloud-based large model through multiple channels;
[0072] The third operation strategy includes: slicing and caching the keyword data, outputting the first part of the keyword data and the second part of the keyword data, and uploading the first part of the keyword data to the cloud-based large model; parsing the timestamp of the first part of the keyword data, and retransmitting the second part of the keyword data according to the timestamp after the network is restored.
[0073] It should be noted that the wireless AI module supports four types of networks running simultaneously: cellular network (4G / 5G), wireless LAN (Wi-Fi 2.4G / 5G), short-range communication (Bluetooth BLE 5.2), and near-field communication (NFC). When the first network operation strategy is Q≥0.7, the network condition is good, and the wireless LAN can be used alone at this time. When the second network operation strategy is 0.5≤Q<0.7, the network speed drops sharply at this time, and the user may be in the subway or elevator. The collection of voice stream data can be transmitted synchronously via Bluetooth, while the upload of keyword data will be fragmented, preferably 200KB / piece, and uploaded via cellular network 5G and Bluetooth respectively. When the third operation strategy is Q<0.5, the network is completely disconnected. Before the network is completely disconnected, the keyword data will be partially cached and uploaded in advance, and then this part of the keyword data will be finally converted into a natural language audio stream output so that the user can get a rough answer. After the network responds, the remaining second part of the keyword data will be transmitted according to the timestamp.
[0074] In this way, dimensional network quality assessment and intelligent switching strategies are used to solve functional reliability issues caused by network instability, achieving low-latency switching and high data reliability in complex network environments.
[0075] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a voice processing device, electronic device and corresponding embodiments based on a micro wireless AI module.
[0076] Figure 5 This is a functional module diagram of a voice processing device based on a micro wireless AI module shown in an embodiment of the present application.
[0077] See also Figure 5 A speech processing device based on a micro wireless AI module includes: a recognition module 100 and an output module 200. The recognition module 100 is used to collect voice stream data, recognize and process the voice stream data, and output keyword data; the output module 200 is used to match the keyword data with offline voice data. If the match fails, the keyword data is uploaded to a large cloud model for parsing and processing, and the natural language audio stream is output.
[0078] In one embodiment, the output module 200 includes a matching unit, which is used to optimize the keyword data and output preprocessed keyword data; identify the preprocessed keyword data and output text data; extract keyword feature data in the text data, the keyword feature data including word domain information, part of speech information and lexical information; pair the keyword feature data with the offline voice data, and if the pairing is successful, generate reply text information based on the keyword feature data, convert the reply text information into audio data and output it; if the pairing fails, upload the keyword data to the cloud-based large model for parsing and outputting a natural language audio stream.
[0079] In one embodiment, the output module 200 includes a parsing unit, which is used to optimize the keyword data and output preprocessed keyword data; identify the preprocessed keyword data and output text data; extract keyword feature data in the text data, the keyword feature data including word domain information, part of speech information and lexical information; pair the keyword feature data with the offline voice data, and if the pairing is successful, generate reply text information based on the keyword feature data, convert the reply text information into audio data and output it; if the pairing fails, upload the keyword data to the cloud-based large model for parsing and outputting a natural language audio stream.
[0080] In one embodiment, a wake-up unit is further included, which is used to receive a voice wake-up word through an external microphone device to end the sleep state.
[0081] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0082] Figure 6It is a structural diagram of an electronic device shown in an embodiment of the present application.
[0083] See also Figure 6 , the electronic device 1000 includes a memory 1010 and a processor 1020.
[0084] The processor 1020 may be a central processing unit (CPU), or an integrated circuit composed of other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be any conventional processor that can run the Linux kernel.
[0085] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage may be a readable and writable storage device. The permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (e.g., a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0086] The memory 1010 stores executable codes. When the executable codes are processed by the processor 1020 , the processor 1020 may execute part or all of the above-mentioned methods.
[0087] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0088] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium) on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.
[0089] The scheme of the present application has been described in detail above with reference to the accompanying drawings. In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. Those skilled in the art should also be aware that the actions and modules involved in the description are not necessarily required for this application. In addition, it is understood that the steps in the method of the embodiment of the present application can be adjusted in sequence, merged and deleted according to actual needs, and the modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0090] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A voice processing method based on a micro wireless AI module, characterized in that: include: Collecting voice stream data, performing recognition processing on the voice stream data, and outputting keyword data; The keyword data is matched with the offline voice data. If the match fails, the keyword data is uploaded to the cloud-based large model for parsing and outputting a natural language audio stream.
2. The voice processing method based on the micro wireless AI module according to claim 1, characterized in that: The matching process of the keyword data with the offline voice data includes: Optimizing the keyword data and outputting pre-processed keyword data; Performing recognition processing on the pre-processed keyword data and outputting text data; Extracting keyword feature data from the text data, wherein the keyword feature data includes word domain information, part of speech information, and lexical information; The keyword feature data is paired with the offline voice data. If the pairing is successful, a reply text message is generated based on the keyword feature data, and the reply text message is converted into audio data and output; if the pairing fails, the keyword data is uploaded to the cloud-based large model for parsing and processing, and a natural language audio stream is output.
3. The voice processing method based on the micro wireless AI module according to claim 1, characterized in that: The uploading of the keyword data to the cloud-based large model for parsing includes: Uploading the keyword data to the first large model through an interactive protocol, the first large model performing a calculation on the keyword data and outputting a natural language text; receiving the natural language text, and sending the natural language text to the second model, wherein the second model parses the natural language text and outputs a reply text; The reply text is received and sent to a third model, the third model converts the reply text and outputs the natural language audio stream.
4. The voice processing method based on the micro wireless AI module according to claim 1, characterized in that: The method further comprises: Receive the voice wake-up word through the external microphone device to end the sleep state.
5. A voice processing device based on a miniature wireless AI module, characterized in that: include: A recognition module, configured to collect voice stream data, perform recognition processing on the voice stream data, and output keyword data; The output module is used to match the keyword data with the offline voice data. If the match fails, the keyword data is uploaded to the cloud-based large model for parsing and outputting a natural language audio stream.
6. The voice processing device based on the micro wireless AI module according to claim 5, characterized in that: The output module includes a matching unit, which is used to optimize the keyword data and output pre-processed keyword data; Performing recognition processing on the pre-processed keyword data and outputting text data; Extracting keyword feature data from the text data, wherein the keyword feature data includes word domain information, part of speech information, and lexical information; Pairing the keyword feature data with the offline voice data, and if the pairing is successful, generating a reply text message based on the keyword feature data, converting the reply text message into audio data, and outputting the audio data; If the pairing fails, the keyword data is uploaded to the cloud-based large model for parsing and processing, and a natural language audio stream is output.
7. The voice processing device based on the micro wireless AI module according to claim 5, characterized in that: The output module includes a parsing unit, which is used to optimize the keyword data and output pre-processed keyword data; Performing recognition processing on the pre-processed keyword data and outputting text data; Extracting keyword feature data from the text data, wherein the keyword feature data includes word domain information, part of speech information, and lexical information; Pairing the keyword feature data with the offline voice data, and if the pairing is successful, generating a reply text message based on the keyword feature data, converting the reply text message into audio data, and outputting the audio data; If the pairing fails, the keyword data is uploaded to the cloud-based large model for parsing and processing, and a natural language audio stream is output.
8. The voice processing device based on the micro wireless AI module according to claim 5, characterized in that: It also includes a wake-up unit, which is used to receive a voice wake-up word through an external microphone device to end the sleep state.
9. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 4.
10. A computer-readable storage medium having executable code stored thereon, wherein when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 4.