Intelligent Interaction System and Method Based on Far-Field Voice Module

Through preprocessing and big data training, response command data is formed, and the local voice recognition engine is awakened, which solves the loss of function and privacy of the far-field voice interaction system under network instability or disconnection, and realizes the convenience of voice interaction in offline voice control and privacy environments.

CN119811390BActive Publication Date: 2025-07-18NANJING JIAHAO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510265287.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The existing far-field voice interaction system loses its function in an unstable network or disconnected environment, and its privacy cannot be guaranteed in a networked environment, resulting in inconvenient voice control.

Method used

By capturing voice input signals for preprocessing, combining acoustics and language model recognition, processing text is formed, and response instruction data is formed based on big data training, overlapping data is stored to wake up the local voice recognition engine, establish offline voice control logic, and realize offline voice interaction and voice control in a privacy environment.

Benefits of technology

Achieving purposeful voice interaction in an out-of-net environment and realizing voice control in a privacy environment, improving the portability and privacy of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811390B_ABST
    Figure CN119811390B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent interaction system and method based on a far-field voice module, which relates to the technical field of speech recognition. The system captures the voice input signal of the user, forms different processing feedback mechanisms according to different network states; combines an acoustic model and a language model to identify and optimize the preprocessed voice input signal to form a final processed text; activates a self-protection mechanism, and forms the first response instruction data for the next cycle based on big data training; uses the cycle mode set by the system to infer and form the second response instruction data for the next cycle; sets up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, so as to form a voice response prediction in an offline state, which can not only ensure purposeful voice interaction in a network-disconnected environment and voice control in a privacy environment, but also reduce the storage of offline voice configuration and improve the portability of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and particularly to an intelligent interaction system and method based on a far-field speech module. Background Art

[0002] Far-field speech refers to the technology of collecting and processing speech at a relatively long distance (usually more than 1 meter). With the continuous development of technologies such as artificial intelligence, the Internet of Things, and edge computing, the high scalability of far-field speech technology enables it to be combined with various advanced technologies, thus supporting a wider range of application scenarios and more complex functions. In the fields of intelligent factories, healthcare, and security, far-field speech interaction can

[0003] bring users a more intelligent and convenient interaction experience.

[0004] However, at present, the access of various Internet devices, although greatly enriching the functions and playability of speech interaction, has also brought endless privacy problems and network problems. For example, in an environment with unstable or disconnected network, various speech interaction devices lose their basic functions or have a huge delay; or in scenarios such as security, intelligent factories, and indoor homes, the excessive access of Internet devices leads to a large amount of data being transmitted on the Internet, and the privacy cannot be guaranteed. How to freely configure offline voice control in a networked environment, be able to achieve targeted speech interaction in a disconnected network environment, and achieve voice control in a privacy environment is one of the problems faced currently. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent interaction system and method based on a far-field speech module to solve the problems raised in the prior art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: An intelligent interaction method based on a far-field speech module, the method comprising the following steps:

[0007] Capture the voice input signal of the user, perform preliminary processing on the voice input signal to form a preprocessed voice input signal, detect the network state, and form different processing feedback mechanisms according to different network states;

[0008] Receive the preprocessed voice input signal, combine an acoustic model and a language model, identify and optimize the preprocessed voice input signal to form a final processed text;

[0009] Call the final processed text, activate the self-protection mechanism, use the system-set periodic method to call the historical response instruction data in the large database, and form the first response instruction data for the next period based on big data training;

[0010] Obtain the current time data, and use the cycle mode set by the system to infer and form the second response instruction data for the next cycle;

[0011] Set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish a local control logic, and form an offline voice control;

[0012] Perform intelligent interactions on the processing texts in different states respectively.

[0013] According to the above technical solution, the preliminary processing of the voice input signal includes:

[0014] Perform noise reduction on the received voice input signal to remove background noise, and at the same time normalize the audio signal to the same volume level to form a preprocessed voice input signal;

[0015] The detection of the network state includes: testing the network state, including the online network and the offline network.

[0016] According to the above technical solution, the recognition and optimization of the preprocessed voice input signal include:

[0017] Based on the acoustic model, model the relationship between voice features and phonemes;

[0018] Based on the language model, analyze and predict the word sequence probability to form the final processed text.

[0019] According to the above technical solution, the formation of the first response instruction data for the next cycle based on big data training includes:

[0020] Obtain the response instruction data of adjacent cycles. In the historical data, use the response instruction data of every two adjacent cycles as a set of analysis data;

[0021] Select any one response instruction data of the current cycle, denote any one response instruction data as data A, search for data A in the analysis data of the historical data, select all the next cycle response instruction data of data A in the historical data, and make a first mark;

[0022] Select all the response instruction data of the current cycle as a response data group, search for the response data group in the analysis data of the historical data, select all the next cycle response instruction data with the response data group in the historical data, and make a first mark;

[0023] Make a second mark on all the marked data. The marks include:

[0024] For any set of analysis data of the historical data, calculate the occupancy data ratio of data A or the response data group: or , where e represents the occupied data ratio, represents the total number of current cycle response instruction data; represents the total number of response instruction data of the response data group;

[0025] Mark the occupied data ratio e on the corresponding selected next cycle response instruction data, adjust all the next cycle response instruction data, set the same next cycle response instruction data into the same group, and at the same time calculate the sum of the marked e within each group as the output value of the next cycle response instruction data of the same group. The system sets a threshold, extracts the groups whose output values exceed the threshold, and selects the next cycle response instruction data corresponding to the extracted groups as the first response instruction data of the next cycle.

[0026] According to the above technical solution, the speculation to form the second response instruction data of the next cycle includes:

[0027] Obtain time information data based on the clock module, and based on the response instruction data of each time information data under historical data, select the time period where the current time information data is located. Taking each day as a cycle range, obtain the response instruction data under the same time period, classify the same response instruction data, and for any category, record the total number of response instruction data in the current category as the category output. If there is , then judge The corresponding response instruction data is the second response instruction data speculated to form the next cycle, where represents the category output of the i-th category; represents the total number of all response instruction data under the same time period; h represents the number of categories formed under the same time period.

[0028] According to the above technical solution, the awakening of the local voice recognition engine, establishing local control logic, and forming offline voice control include:

[0029] Embed an open-source voice recognition engine inside the system, call the control logic according to the recognition result, and call the corresponding hardware control code according to the control logic to achieve offline voice control.

[0030] For the intelligent interaction system based on the far-field voice module, the intelligent interaction system includes:

[0031] A voice collection and preprocessing module, which is responsible for capturing the user's voice input signal, performing preliminary processing on the voice input signal to form a preprocessed voice input signal, detecting the network state, and forming different processing feedback mechanisms according to different network states;

[0032] A post - processing and feedback module, which is used to receive the pre - processed voice input signal, combine an acoustic model and a language model, identify and optimize the pre - processed voice input signal, and form the final processed text;

[0033] A feature expansion module, which is used to call the final processed text, activate the self - protection mechanism, use the system - set periodic method to call the historical response instruction data in the large database, and form the first response instruction data for the next period based on big data training;

[0034] A time speculation module, which is used to obtain the current time data and speculate to form the second response instruction data for the next period using the system - set periodic method;

[0035] An offline voice control module, which is used to set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish local control logic, and form offline voice control;

[0036] An intelligent interaction module, which is used to form intelligent interaction based on the processed text.

[0037] According to the above - mentioned technical solution, the voice acquisition and pre - processing module includes:

[0038] Automatic network status recognition, including online network and offline network. When the network is normal, it works with the online network. In case of network anomalies, the online network entrance is closed and it works with the offline network;

[0039] Manual network status recognition, which is used for the administrator to set the corresponding password and can manually adjust to the offline network.

[0040] According to the above - mentioned technical solution, the time speculation module is also connected to a clock module, and the clock module provides time information data;

[0041] The clock module also includes a clock database, and the clock database stores the time information data corresponding to any one of the response instructions in the historical data.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: In the era of the rapid development of voice interaction, this application focuses on protecting privacy and can also solve the problem of offline voice control caused by network disconnection or network interference. It greatly enriches the functions and playability of voice interaction and can be applied to fields such as intelligent factories, privacy - focused homes, and intelligent security. Through continuous speculation in the network - connected environment, voice response instruction predictions for each period are formed, thus forming voice response predictions in the offline state. This can not only ensure purposeful voice interaction in the offline environment and voice control in the privacy environment, but also reduce the storage of offline voice configuration and improve the portability of voice interaction. Brief Description of the Drawings

[0043] Figure 1 This is a schematic diagram of the module connection of the intelligent interaction system based on the far-field voice module of the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Embodiment: The present invention provides an intelligent interaction method based on a far-field voice module. The method includes the following steps:

[0046] Capture the voice input signal of the user, perform preliminary processing on the voice input signal to form a preprocessed voice input signal, detect the network status, and form different processing feedback mechanisms according to different network statuses;

[0047] Receive the preprocessed voice input signal, combine the acoustic model and the language model, identify and optimize the preprocessed voice input signal to form a final processed text;

[0048] Call the final processed text, activate the self-protection mechanism, use the system-set cycle method to call the historical response instruction data in the large database, and form the first response instruction data for the next cycle based on big data training;

[0049] Obtain the current time data, and use the system-set cycle method to infer and form the second response instruction data for the next cycle;

[0050] Set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish a local control logic, and form an offline voice control;

[0051] Perform intelligent interaction on the processed text in different states respectively.

[0052] The preliminary processing of the voice input signal includes:

[0053] Receive a voice signal using a corresponding voice detection device, sample and quantize the received voice input signal. The first-step processing can be achieved by frame division (for example, dividing the signal into short-time frames, usually each frame being 20 - 40 ms) and windowing (reducing frame edge effects); then estimate the noise using the minimum statistic or recursive averaging method, perform noise reduction, and remove background noise. During the noise reduction process, various algorithms can be used, such as spectral subtraction or Wiener filtering method for implementation. If the signal data is in a non-stationary state, adaptive filtering or wavelet transform can be used.

[0054] At the same time, normalize the audio signal to the same volume level to form a preprocessed voice input signal;

[0055] The detection of the network state includes: testing the network state, including online network and offline network.

[0056] The recognition and optimization of the preprocessed voice input signal include:

[0057] Model the relationship between speech features and phonemes based on an acoustic model;

[0058] Analyze and predict the probability of word sequences based on a language model to form the final processed text.

[0059] The formation of the first response instruction data for the next cycle based on big data training includes:

[0060] Obtain the response instruction data of adjacent cycles. In the historical data, take the response instruction data of every two adjacent cycles as a set of analysis data;

[0061] Select any one response instruction data of the current cycle, denote any one response instruction data as data A, search for data A in the analysis data of the historical data, select all the next-cycle response instruction data of data A in the historical data, and make a first mark;

[0062] Select all the response instruction data of the current cycle as a response data group, search for the response data group in the analysis data of the historical data, select all the next-cycle response instruction data in the historical data where the response data group exists, and make a first mark;

[0063] Make a second mark on all the marked data. The marks include:

[0064] For any set of analysis data of historical data, calculate the occupancy data ratio of data A or the response data group: or , where e represents the occupancy data ratio, represents the total number of response instruction data in the current cycle; The total number of response instruction data representing the response data group;

[0065] For example, take any response data instruction "turn on the light". Then, in the analysis data, for any analysis data that had the instruction "turn on the light" in the previous cycle, all its response instructions in the next cycle are marked; the response data group uses the team marking method. For example, if there were instructions such as "turn on the light" and "turn on the air conditioner" in the previous cycle, the instructions are packaged and analyzed in a team. If there are the same groups in the historical data, all the corresponding response instructions in the next cycle are marked.

[0066] Mark the occupancy data ratio e on the corresponding selected response instruction data in the next cycle, adjust all the response instruction data in the next cycle, set the same response instruction data in the next cycle to the same group, and at the same time calculate the sum of the marked e within each group as the output value of the response instruction data in the same group in the next cycle. The system sets a threshold, extracts the groups whose output values exceed the threshold, and selects the response instruction data in the next cycle corresponding to the extracted groups as the first response instruction data in the next cycle.

[0067] The speculation to form the second response instruction data in the next cycle includes:

[0068] Based on the clock module to obtain time information data, based on the response instruction data of each time information data in the historical data, select the time period where the current time information data is located. Taking each day as a cycle range, obtain the response instruction data in the same time period, classify the same response instruction data, and for any category, record the total number of response instruction data in the current category as the category output. If there exists , then judge The corresponding response instruction data as the second response instruction data speculated to form in the next cycle, where represents the category output in the i-th category; represents the total number of all response instruction data in the same time period; h represents the number of categories formed in the same time period.

[0069] The awakening of the local voice recognition engine, establishing local control logic, and forming offline voice control includes:

[0070] Embed an open-source voice recognition engine inside the system, call the control logic according to the recognition result, and call the corresponding hardware control code according to the control logic to achieve offline voice control.

[0071] In the second embodiment, as Figure 1 shown, it further includes an intelligent interaction system based on the far-field voice module. The intelligent interaction system includes:

[0072] The voice acquisition and preprocessing module is responsible for capturing the user's voice input signal, performing preliminary processing on the voice input signal to form a preprocessed voice input signal, detecting the network status, and forming different processing feedback mechanisms according to different network statuses;

[0073] The post-processing and feedback module is used to receive the preprocessed voice input signal, combine the acoustic model and the language model, identify and optimize the preprocessed voice input signal to form the final processed text;

[0074] The feature expansion module is used to call the final processed text, activate the self-protection mechanism, use the system-set cycle method to call the historical response instruction data in the large database, and form the first response instruction data for the next cycle based on big data training;

[0075] The time prediction module is used to obtain the current time data and predict and form the second response instruction data for the next cycle using the system-set cycle method;

[0076] The offline voice control module is used to set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish local control logic, and form offline voice control;

[0077] The intelligent interaction module is used to form intelligent interaction based on the processed text.

[0078] The voice acquisition and preprocessing module includes:

[0079] Automatic network status recognition, including online network and offline network. When the network is normal, it works in the online network. In the case of network anomalies, the online network entrance is closed and it works in the offline network;

[0080] Manual network status recognition is used for the administrator to set the corresponding password and can manually adjust to the offline network.

[0081] The time prediction module is also connected to a clock module, and the clock module provides time information data;

[0082] The clock module also contains a clock database, and the clock database stores the time information data corresponding to any one of the response instructions in the historical data.

[0083] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. An intelligent interaction method based on a far-field voice module, characterized in that: The method includes the following steps: Capture the user's voice input signal, perform preliminary processing on the voice input signal to form a preprocessed voice input signal, detect the network status, and form different processing feedback mechanisms according to different network statuses; Receive the preprocessed voice input signal, combine the acoustic model and the language model, identify and optimize the preprocessed voice input signal to form the final processed text; Call the final processed text, activate the self-protection mechanism, use the system-set periodic method to call the historical response instruction data in the large database, and form the first response instruction data for the next period based on big data training; Obtain the current time data, use the system-set periodic method to infer and form the second response instruction data for the next period; Set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish local control logic, and form offline voice control; Form intelligent interactions for the processed text in different states respectively; The inferring and forming the second response instruction data for the next period includes: Obtain time information data based on the clock module. Based on the response instruction data of each time information data under historical data, select the time period where the current time information data is located. Taking each day as a cycle range, obtain the response instruction data under the same time period, classify the same response instruction data. For any category, record the total number of response instruction data under the current category as the category output. If there is , then judge The corresponding response instruction data is the second response instruction data speculated to form the next cycle. Among them, represents the category output under the i-th category; represents the total number of all response instruction data under the same time period; h represents the number of categories formed under the same time period.

2. The intelligent interaction method based on a far-field voice module according to claim 1, wherein: The performing preliminary processing on the voice input signal includes: Perform noise reduction on the received voice input signal, remove background noise, and at the same time normalize the audio signal to the same volume level to form a preprocessed voice input signal; The detecting the network status includes: testing the network status, including the online network and the offline network.

3. The intelligent interaction method based on a far-field voice module according to claim 1, wherein: The identifying and optimizing the preprocessed voice input signal includes: Model the relationship between voice features and phonemes based on the acoustic model; Analyze and predict the word sequence probability based on the language model to form the final processed text.

4. The intelligent interaction method based on a far-field voice module according to claim 1, wherein: The forming the first response instruction data for the next period based on big data training includes: Obtain the response instruction data of adjacent periods, and in the historical data, use the response instruction data of every two adjacent periods as a set of analysis data; Select any one response instruction data of the current period, denote any one response instruction data as data A, search for data A in the analysis data of the historical data, select all the next-period response instruction data of data A in the historical data, and make a first mark; Select all the response instruction data of the current period as a response data group, search for the response data group in the analysis data of the historical data, select all the next-period response instruction data in the historical data where the response data group exists, and make a first mark; Make a second mark on all the marked data, and the marks include: For the analysis data of any set of historical data, calculate the occupancy data ratio of data A or the response data set: , where e represents the occupancy data ratio, represents the total number of response instruction data in the current cycle; represents the total number of response instruction data in the response data set; Mark the occupancy data ratio e on the corresponding selected next-period response instruction data, adjust all the next-period response instruction data, set the same next-period response instruction data to the same group, and at the same time calculate the sum of the marked e within each group as the output value of the same group of next-period response instruction data. The system sets a threshold, extracts the groups whose output values exceed the threshold, and selects the next-period response instruction data corresponding to the extracted groups as the first response instruction data for the next period.

5. The intelligent interaction method based on a far-field voice module according to claim 1, characterized in that: The waking up the local voice recognition engine, establishing local control logic, and forming offline voice control includes: Embed an open-source speech recognition engine inside the system, call the control logic according to the recognition result, and call the corresponding hardware control code according to the control logic to achieve offline voice control.

6. An intelligent interaction system based on a far-field voice module, which is used to implement the intelligent interaction method based on a far-field voice module as described in claim 1, and is characterized in that: The intelligent interaction system includes: A voice acquisition and preprocessing module, which is responsible for capturing the user's voice input signal, preliminarily processing the voice input signal to form a preprocessed voice input signal, detecting the network status, and forming different processing feedback mechanisms according to different network statuses; A post-processing and feedback module, which is used to receive the preprocessed voice input signal, combine the acoustic model and the language model, identify and optimize the preprocessed voice input signal to form the final processed text; A feature expansion module, which is used to call the final processed text, activate the self-protection mechanism, use the cycle mode set by the system to call the historical response instruction data in the large database, and form the first response instruction data for the next cycle based on big data training; A time prediction module, which is used to obtain the current time data and predict and form the second response instruction data for the next cycle using the cycle mode set by the system; An offline voice control module, which is used to set up a storage space to store the overlapping response instruction data in the first response instruction data and the second response instruction data, wake up the local voice recognition engine, establish the local control logic, and form offline voice control; An intelligent interaction module, which is used to form intelligent interaction based on the processed text.

7. The intelligent interaction system based on a far-field voice module according to claim 6, characterized in that: The voice acquisition and preprocessing module includes: Automatic network status recognition, including online network and offline network. When the network is normal, it works with the online network. In case of network anomalies, the online network entrance is closed and it works with the offline network; Manual network status recognition, which is used for the administrator to set the corresponding password and can manually adjust to the offline network.

8. The intelligent interaction system based on a far-field voice module according to claim 7, characterized in that: The time prediction module is also connected to a clock module, and the clock module provides time information data; The clock module also includes a clock database, and the clock database stores the time information data corresponding to any one of the response instructions in the historical data.

Citation Information

Patent Citations

  • Method and system for controlling water heater

    CN110044073A

  • Instrument data management system and method based on intelligent voice recognition

    CN117275474A