Embedded voice recognition control module and its device interaction system
By integrating deep learning and natural language processing technologies through the embedded speech recognition control module, the network latency and privacy security issues of traditional speech recognition systems are solved, and low-power, efficient voice control and device interaction are achieved.
Patent Information
- Application Number
- CN202411678546.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Traditional voice recognition systems rely on cloud computing, which has problems such as network latency, privacy and security, high power consumption, and complex device interaction. They cannot meet the user's needs for intelligence and miniaturization when their hands are occupied.
It adopts an embedded speech recognition control module, integrating deep learning speech recognition technology and natural language processing technology, including speech acquisition, preprocessing, recognition and command parsing, control instruction generation and feedback modules, and improves system adaptability through adaptive learning and personalized sub-modules.
It achieves low-power voice control, reduces dependence on the network, improves response speed and privacy protection, and enhances the convenience and intelligence of device interaction.
Smart Images

Figure CN119763564B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of speech recognition and device control, and in particular to an embedded speech recognition control module and a device interaction system thereof. Background Art
[0002] With the development of the Internet of Things, people are increasingly demanding smarter and more compact devices. Voice control is crucial in many scenarios where people's hands are occupied, such as when driving, cooking, or simply carrying something. People want to be able to complete various tasks, such as querying information and controlling devices, through voice commands without having to put down their hands to manually operate them.
[0003] Compared to manually operating devices, voice control can complete tasks more quickly. Traditional voice recognition systems typically rely on cloud servers for computation, which poses issues such as network latency, privacy concerns, high power consumption, and complex device interactions. Therefore, this paper proposes an embedded voice recognition control module and its device interaction system. Summary of the Invention
[0004] The present invention provides an embedded voice recognition control module and a device interaction system thereof, which integrates deep learning voice recognition technology and natural language processing technology into a compact embedded voice recognition control module, so that voice control command recognition and feedback can be completed with only low power consumption. At the same time, it also reduces the dependence of voice control on the network, improves the response speed, and can better protect the privacy of users.
[0005] The present invention provides an embedded speech recognition control module, characterized by comprising:
[0006] A voice acquisition submodule consisting of one or more microphone arrays, used to collect the user's voice signal;
[0007] The speech preprocessing submodule is used to preprocess the speech signal to obtain high-quality speech signals;
[0008] The speech recognition and command parsing submodule is used to convert high-quality speech signals into text data based on a deep learning speech recognition model, and to perform semantic parsing of the text data through natural language processing technology to determine user control commands;
[0009] The control instruction generation submodule is used to generate corresponding device control instructions based on user control commands and send them to the lower-level devices;
[0010] The actuator and feedback submodule is used to generate corresponding feedback voice based on the control instruction execution results of the actuator of the lower device;
[0011] The adaptive learning and personalization submodule is used to adaptively train the deep learning speech recognition model based on the user's text usage habits and accent.
[0012] Preferably, in an embedded speech recognition control module, the speech acquisition submodule includes:
[0013] The direction sensing unit is used to determine the source direction of each sound source signal based on the time difference and phase difference of the sound signals received by each microphone in the microphone array, and to determine the effective coverage range corresponding to each channel signal according to the size of the different channel signals received by each microphone;
[0014] a sound data analysis unit, configured to determine an effective coverage area corresponding to a human voice signal as a target effective coverage range, and to confirm a target microphone array area corresponding to the target effective coverage range;
[0015] Acquire a first target human voice signal collected by each microphone in each target microphone array area, perform target sound fusion based on the first target human voice signal to obtain a second target human voice signal, and send the second target human voice signal as the collected voice signal to the voice preprocessing submodule;
[0016] a sensitivity control unit, configured to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value, obtain an acquisition error rate, and adjust the sensitivity of the microphone corresponding to each target human voice signal based on the acquisition error rate and the speaker position information;
[0017] When any microphone is within the target microphone array area corresponding to two or more channel signals, the final sensitivity of the microphone is based on the sensitivity corresponding to the channel signal with greater sensitivity.
[0018] Preferably, in an embedded speech recognition control module, the speech preprocessing submodule includes:
[0019] A first preprocessing unit is used to perform preliminary processing on the speech signal based on Gaussian filtering and normalized least mean square algorithm to obtain a denoised speech signal;
[0020] The second preprocessing unit is used to obtain the volume characteristics corresponding to each denoised speech signal, and perform volume normalization processing on the denoised speech signal based on the volume characteristics to obtain a high-quality speech signal.
[0021] Preferably, in an embedded speech recognition control module, the sensitivity control unit includes:
[0022] The first intelligent control sub-unit is used to compare the total coverage area of all target effective coverage areas with the full coverage area of the microphone array, determine the vocal invalid area, and adjust the sensitivity value of the microphone corresponding to the vocal invalid area to a minimum threshold;
[0023] The speaker localization subunit is used to obtain the distribution characteristics of the microphones in the target microphone array area corresponding to each human voice signal, and determine the relative distance between the speaker and the embedded speech recognition control module based on the distribution characteristics and the sound propagation characteristics;
[0024] The second intelligent control subunit is configured to compare the relative distance with a preset standard value to obtain a position difference, and to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value to obtain an acquisition error rate;
[0025] Determine the microphone sensitivity adjustment direction and the corresponding adjustment coefficient based on the positive and negative corresponding position difference, and obtain the microphone sensitivity adjustment amplitude based on the acquisition error and the standard sensitivity of the microphone;
[0026] An actual control value is obtained according to the sensitivity adjustment amplitude and the adjustment coefficient corresponding to the sensitivity adjustment direction, and the sensitivity of the target microphone is adjusted based on the actual control value and according to the sensitivity adjustment direction corresponding to the microphone.
[0027] Preferably, in an embedded speech recognition control module, the speech recognition and command parsing submodule includes:
[0028] The model recognition unit is used to identify high-quality voice signals based on a deep learning voice recognition model and convert them into text data. It also performs semantic analysis on the text data based on natural language technology to determine user semantics and user intent.
[0029] The command screening unit is used to screen high-quality voice signals based on user semantics and user intent, combined with a preset control command data set, to obtain effective control voice;
[0030] The command determination unit is used to determine the user control command according to the corresponding relationship between the effective control voice and the control vocabulary in the preset control instruction data set and the control instruction corresponding to the control vocabulary.
[0031] Preferably, in an embedded speech recognition control module, the control instruction generation submodule includes:
[0032] An instruction generation unit, configured to generate corresponding control instructions based on user commands and in combination with the communication protocol between the embedded speech recognition control module and the lower-level device;
[0033] The instruction sending unit is used to send control instructions to the lower-level device through a preset communication channel.
[0034] Preferably, in an embedded speech recognition control module, the adaptive learning and personalization submodule includes:
[0035] A user data processing unit is configured to obtain all valid historical wake-up records of the voice recognition control module, define the period from the initial wake-up to the successful recognition of the control command by the voice recognition control module as one wake-up cycle, divide the voice recognition control module into multiple successful control wake-up cycles, and obtain multiple successful control wake-up cycle data sets;
[0036] Respectively obtaining invalid voice contents of the user in a plurality of successful control wake-up cycles, and when the number of invalid voice contents is zero, eliminating the successful control wake-up cycle;
[0037] A data identification unit is configured to generate a corresponding instruction tag based on the successful control instruction corresponding to the successful voice wake-up cycle and add it to the successful control wake-up cycle data set, and monitor invalid voices, or true invalid voices and pseudo invalid voices, based on the keywords of the successful control instruction corresponding to the successful control wake-up cycle data set;
[0038] A usage habit analysis unit is used to respectively obtain unique voice features corresponding to each pseudo-invalid voice, compare the unique voice features with the voice features of the standard control instruction corresponding to the control command, and obtain the accent difference between the user's actual voice and the pre-selected voice and the user's text usage habits;
[0039] Based on the voiceprint features of the voice signals in different successful control wake-up cycles, the successful control wake-up cycles are clustered to obtain the voice control data clusters corresponding to the same user;
[0040] Compare the accent differences corresponding to the pseudo-invalid speech within the same voice control data cluster and the user's text usage habits to obtain the current user's personalized accent and usage habits, and compare the personalized accents and usage habits of different users to obtain user voice similarity;
[0041] a data merging unit, configured to merge voice control data clusters whose user voice similarity is greater than a preset value to obtain a merged data set, and generate different batches of training sets based on the merged data set and the unmerged voice control data clusters;
[0042] The adaptive training unit is used to input the corresponding training set into the deep learning speech recognition model for adaptive training according to the training set batches until all batches of training are completed and the adaptive training is ended.
[0043] Preferably, in an embedded speech recognition control module, the user data processing unit includes:
[0044] The voiceprint file creation subunit is used to create a voiceprint file based on the voiceprint features corresponding to the voice signal. When a new voiceprint feature is detected, a new voiceprint file is automatically generated.
[0045] When the current user's voice is confirmed to be a voiceprint object, determine whether the current user's voice control command is a single feedback command. If so, add an invalid label to the historical wake-up record of the current user;
[0046] Otherwise, add a valid tag to the historical wakeup record of the current user.
[0047] Preferably, in an embedded speech recognition control module, the user data processing unit further includes:
[0048] An intelligent comparison sub-unit is used to compare the current voice segment collected by the profiled user in the current successful control wake-up cycle with the historical voice segment in the invalid voiceprint file corresponding to the latest similar control command of the profiled user when a pseudo-invalid voice appears during the voice control command recognition process of the profiled user, and determine the usage habit and accent differences of the profiled user based on the comparison results;
[0049] The audio corresponding to the current speech segment is highlighted according to the usage habit difference or the pronunciation difference, and a training control group is generated based on the highlighted speech segment and the historical speech segment. The training control group is used to perform adaptive training on the deep learning speech recognition model.
[0050] The present invention provides a device interaction system of an embedded voice recognition control module, characterized in that a communication connection is established between any one of the embedded voice recognition control modules and its corresponding lower-level device.
[0051] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0052] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0054] Figure 1This is a structural diagram of an embedded speech recognition control module in an embodiment of the present invention;
[0055] Figure 2 This is a structural diagram of a voice acquisition submodule of an embedded voice recognition control module in an embodiment of the present invention;
[0056] Figure 3 This is a structural diagram of a speech preprocessing submodule of an embedded speech recognition control module in an embodiment of the present invention;
[0057] Figure 4 This is a structural diagram of a speech recognition and command parsing submodule of an embedded speech recognition control module in an embodiment of the present invention;
[0058] Figure 5 This is a structural diagram of a control instruction generation submodule of an embedded speech recognition control module in an embodiment of the present invention;
[0059] Figure 6 This is a structural diagram of an adaptive learning and personalization submodule of an embedded speech recognition control module in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0061] Example 1:
[0062] The present invention provides an embedded speech recognition control module, such as Figure 1 Shown, including:
[0063] A voice acquisition submodule consisting of one or more microphone arrays, used to collect the user's voice signal;
[0064] The speech preprocessing submodule is used to preprocess the speech signal to obtain high-quality speech signals;
[0065] The speech recognition and command parsing submodule is used to convert high-quality speech signals into text data based on a deep learning speech recognition model, and to perform semantic parsing of the text data through natural language processing technology to determine user control commands;
[0066] The control instruction generation submodule is used to generate corresponding device control instructions based on user control commands and send them to the lower-level devices;
[0067] The actuator and feedback submodule is used to generate corresponding feedback voice based on the control instruction execution results of the actuator of the lower device;
[0068] The adaptive learning and personalization submodule is used to adaptively train the deep learning speech recognition model based on the user's text usage habits and accent.
[0069] In this embodiment, the feedback voice can be a simple operation confirmation (such as "The light is on") or a prompt indicating an operation failure (such as "The curtain opening failed, please check the device status"). In this way, the interactive experience between the user and the device is enhanced, and the user can intuitively know whether the device has operated as desired.
[0070] In this embodiment, the lower-level devices include but are not limited to home appliances, automobiles, industrial automation equipment, and wearable devices.
[0071] The beneficial effects of the above technical solution are as follows: the present invention effectively improves the quality of voice signal acquisition through a voice acquisition submodule composed of one or more microphone arrays. The microphone array can utilize the spatial position relationship between multiple microphones to enhance the reception of sounds in a specific direction while suppressing interference noise from other directions, effectively improving the clarity of the collected voice signal; the voice preprocessing submodule preprocesses the voice signal to obtain high-quality voice signals, which is conducive to improving the accuracy of voice recognition; the voice recognition and command parsing submodule converts the high-quality voice signals into text data using a deep learning voice recognition model. The deep learning model has powerful feature extraction capabilities and can automatically learn complex features in the voice signal. Then, the semantics of the text data can be accurately determined by natural language processing technology to understand the user's control command and understand the user's intention. The control instruction generation submodule can generate corresponding device control instructions based on the user control command and send them to the lower-level device. The precise instruction generation and transmission mechanism ensures that the user's voice command can be accurately converted into the actual operation of the device, improving the convenience and intelligence of device control; the actuator and feedback submodule generates corresponding feedback voice based on the control instruction execution result of the lower-level device actuator, so that the user can timely understand the operation status of the device. Finally, the adaptive learning and personalization submodule adaptively trains the deep learning speech recognition model based on the user's text usage habits and accent, enabling the speech recognition system to better adapt to the characteristics of each user. As the user's usage increases, the system becomes increasingly familiar with the user's accent, vocabulary usage habits, and so on, thereby continuously improving the accuracy of speech recognition. This invention integrates deep learning speech recognition technology and natural language processing technology into a compact embedded speech recognition control module, allowing voice control command recognition and feedback to be completed with only low power consumption. It also reduces the voice control's reliance on the network, improves response speed, and better protects user privacy.
[0072] Example 2:
[0073] On the basis of Example 1, the voice collection submodule, such as Figure 2 Shown, including:
[0074] The direction sensing unit is used to determine the source direction of each sound source signal based on the time difference and phase difference of the sound signals received by each microphone in the microphone array, and to determine the effective coverage range corresponding to each channel signal according to the size of the different channel signals received by each microphone;
[0075] a sound data analysis unit, configured to determine an effective coverage area corresponding to a human voice signal as a target effective coverage range, and to confirm a target microphone array area corresponding to the target effective coverage range;
[0076] Acquire a first target human voice signal collected by each microphone in each target microphone array area, perform target sound fusion based on the first target human voice signal to obtain a second target human voice signal, and send the second target human voice signal as the collected voice signal to the voice preprocessing submodule;
[0077] a sensitivity control unit, configured to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value, obtain an acquisition error rate, and adjust the sensitivity of the microphone corresponding to each target human voice signal based on the acquisition error rate and the speaker position information;
[0078] When any microphone is within the target microphone array area corresponding to two or more channel signals, the final sensitivity of the microphone is based on the sensitivity corresponding to the channel signal with greater sensitivity.
[0079] The beneficial effects of the above technical solution: the present invention determines the effective coverage range of each sound signal through the direction perception unit, and then determines the effective coverage range of the human voice signal through the sound data analysis unit and confirms the target microphone array area corresponding to the target effective coverage range, and at the same time, generates a second target human voice signal with higher clarity by fusing the human voice signals collected by different microphones. At the same time, in the process of human voice signal collection, the sensitivity of the microphone is adjusted according to the signal strength of the first target human voice signal collected by each microphone, thereby enhancing the microphone array's reception of sound in a specific direction and suppressing interference noise from other directions, effectively improving the quality of the human voice signal, and providing a basic guarantee for the accuracy of speech recognition.
[0080] Example 3:
[0081] On the basis of Example 1, the speech preprocessing submodule, such as Figure 3 Shown, including:
[0082] A first preprocessing unit is used to perform preliminary processing on the speech signal based on Gaussian filtering and normalized least mean square algorithm to obtain a denoised speech signal;
[0083] The second preprocessing unit is used to obtain the volume characteristics corresponding to each denoised speech signal, and perform volume normalization processing on the denoised speech signal based on the volume characteristics to obtain a high-quality speech signal.
[0084] The beneficial effects of the above technical solution are as follows: The present invention uses a first preprocessing unit to perform denoising and echo cancellation on the corresponding voice signal, effectively reducing the interference of noise on the human voice signal while also reducing the interference between human voice signals. This can effectively improve the module's voice recognition efficiency and accuracy, thereby improving the responsiveness of user voice control. The second preprocessing unit then obtains the volume characteristics corresponding to each denoised voice signal and, based on the volume characteristics, performs volume normalization processing on the denoised voice signal to eliminate differences between the individual audio signals and obtain a high-quality voice signal.
[0085] Example 4:
[0086] Based on Example 2, the sensitivity control unit includes:
[0087] The first intelligent control sub-unit is used to compare the total coverage area of all target effective coverage areas with the full coverage area of the microphone array, determine the vocal invalid area, and adjust the sensitivity value of the microphone corresponding to the vocal invalid area to a minimum threshold;
[0088] The speaker localization subunit is used to obtain the distribution characteristics of the microphones in the target microphone array area corresponding to each human voice signal, and determine the relative distance between the speaker and the embedded speech recognition control module based on the distribution characteristics and the sound propagation characteristics;
[0089] The second intelligent control subunit is configured to compare the relative distance with a preset standard value to obtain a position difference, and to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value to obtain an acquisition error rate;
[0090] Determine the microphone sensitivity adjustment direction and the corresponding adjustment coefficient based on the positive and negative corresponding position difference, and obtain the microphone sensitivity adjustment amplitude based on the acquisition error and the standard sensitivity of the microphone;
[0091] An actual control value is obtained according to the sensitivity adjustment amplitude and the adjustment coefficient corresponding to the sensitivity adjustment direction, and the sensitivity of the target microphone is adjusted based on the actual control value and according to the sensitivity adjustment direction corresponding to the microphone.
[0092] Beneficial effects of the above technical solution: The present invention utilizes the distance between the speaker and the embedded book module and the actual sound receiving effect of each microphone to enhance the reception of sound from a specific direction, while suppressing interference noise from other directions, effectively improving the clarity of the collected voice signal.
[0093] Example 5:
[0094] On the basis of Example 1, the speech recognition and command parsing submodule, such as Figure 4 Shown, including:
[0095] The model recognition unit is used to identify high-quality voice signals based on a deep learning voice recognition model and convert them into text data. It also performs semantic analysis on the text data based on natural language technology to determine user semantics and user intent.
[0096] The command screening unit is used to screen high-quality voice signals based on user semantics and user intent, combined with a preset control command data set, to obtain effective control voice;
[0097] The command determination unit is used to determine the user control command according to the corresponding relationship between the effective control voice and the control vocabulary in the preset control instruction data set and the control instruction corresponding to the control vocabulary.
[0098] The beneficial effects of the above technical solution: The present invention uses a deep learning speech recognition model to convert high-quality voice signals into text data. The deep learning model has a powerful feature extraction capability and can automatically learn complex features in the voice signal, such as speech patterns under different accents, speaking speeds, intonations, etc. Then, the semantics of the text data is parsed by natural language processing technology to accurately determine the user control commands. This enables the system to understand the user's intentions, whether it is a simple instruction (such as "turn on the lights") or a complex instruction (such as "adjust the air conditioning temperature in the living room to 26 degrees and turn on the air purification mode") can be effectively parsed.
[0099] Example 6:
[0100] On the basis of Example 1, the control instruction generation submodule is as follows: Figure 5 Shown, including:
[0101] An instruction generation unit, configured to generate corresponding control instructions based on user commands and in combination with the communication protocol between the embedded speech recognition control module and the lower-level device;
[0102] The instruction sending unit is used to send control instructions to the lower-level device through a preset communication channel.
[0103] The beneficial effects of the above technical solution are: The present invention can generate corresponding device control instructions based on user control commands and send them to lower-level devices. The precise instruction generation and transmission mechanism ensures that the user's voice commands are accurately converted into actual device operations. For example, in a smart home system, the user's voice commands can directly control various smart appliances such as lamps, curtains, and electrical appliances, improving the convenience and intelligence of device control.
[0104] Example 7:
[0105] Based on Example 1, the adaptive learning and personalization submodule, such as Figure 6 Shown, including:
[0106] A user data processing unit is configured to obtain all valid historical wake-up records of the voice recognition control module, define the period from the initial wake-up to the successful recognition of the control command by the voice recognition control module as one wake-up cycle, divide the voice recognition control module into multiple successful control wake-up cycles, and obtain multiple successful control wake-up cycle data sets;
[0107] Respectively obtaining invalid voice contents of the user in a plurality of successful control wake-up cycles, and when the number of invalid voice contents is zero, eliminating the successful control wake-up cycle;
[0108] A data identification unit is configured to generate a corresponding instruction tag based on the successful control instruction corresponding to the successful voice wake-up cycle and add it to the successful control wake-up cycle data set, and monitor invalid voices, or true invalid voices and pseudo invalid voices, based on the keywords of the successful control instruction corresponding to the successful control wake-up cycle data set;
[0109] A usage habit analysis unit is used to respectively obtain unique voice features corresponding to each pseudo-invalid voice, compare the unique voice features with the voice features of the standard control instruction corresponding to the control command, and obtain the accent difference between the user's actual voice and the pre-selected voice and the user's text usage habits;
[0110] Based on the voiceprint features of the voice signals in different successful control wake-up cycles, the successful control wake-up cycles are clustered to obtain the voice control data clusters corresponding to the same user;
[0111] Compare the accent differences corresponding to the pseudo-invalid speech within the same voice control data cluster and the user's text usage habits to obtain the current user's personalized accent and usage habits, and compare the personalized accents and usage habits of different users to obtain user voice similarity;
[0112] a data merging unit, configured to merge voice control data clusters whose user voice similarity is greater than a preset value to obtain a merged data set, and generate different batches of training sets based on the merged data set and the unmerged voice control data clusters;
[0113] The adaptive training unit is used to input the corresponding training set into the deep learning speech recognition model for adaptive training according to the training set batches until all batches of training are completed and the adaptive training is ended.
[0114] The user data processing unit includes:
[0115] The voiceprint file creation subunit is used to create a voiceprint file based on the voiceprint features corresponding to the voice signal. When a new voiceprint feature is detected, a new voiceprint file is automatically generated.
[0116] When the current user's voice is confirmed to be a voiceprint object, determine whether the current user's voice control command is a single feedback command. If so, add an invalid label to the historical wake-up record of the current user;
[0117] Otherwise, add a valid tag to the historical wake-up record of the current user;
[0118] An intelligent comparison sub-unit is used to compare the current voice segment collected by the profiled user in the current successful control wake-up cycle with the historical voice segment in the invalid voiceprint file corresponding to the latest similar control command of the profiled user when a pseudo-invalid voice appears during the voice control command recognition process of the profiled user, and determine the usage habit and accent differences of the profiled user based on the comparison results;
[0119] The audio corresponding to the current speech segment is highlighted according to the usage habit difference or the pronunciation difference, and a training control group is generated based on the highlighted speech segment and the historical speech segment. The training control group is used to perform adaptive training on the deep learning speech recognition model.
[0120] In this embodiment, the unique voice features include the user's accent, speaking speed, pitch, word order, and other characteristics.
[0121] The beneficial effect of the above technical solution is that the present invention adaptively trains the deep learning speech recognition model based on the user's text usage habits and accent, enabling the speech recognition system to better adapt to the characteristics of each user. As the user uses the system more frequently, the system becomes increasingly familiar with the user's accent, vocabulary usage habits, etc., thereby continuously improving the accuracy of speech recognition.
[0122] Example 8:
[0123] The present invention provides a device interaction system of an embedded voice recognition control module, which establishes a communication connection between the embedded voice recognition control module described in any one of 1-7 and its corresponding lower-level device.
[0124] In this embodiment, the communication connection includes a wired communication connection and a wireless communication connection:
[0125] For wired communications, the communication protocols for interfaces such as SPI, I2C, and UART have been optimized. The SPI interface uses adaptive clock technology, which automatically adjusts the clock frequency according to the data transmission rate, improving data transmission efficiency. The I2C interface has added collision detection and automatic retransmission mechanisms to enhance communication stability. The UART interface supports multiple baud rate adaptations, facilitating connection with different devices.
[0126] In terms of wireless communications, the Wi-Fi 6 module supports MU-MIMO technology, which can simultaneously transmit high-speed data with multiple devices; the Bluetooth 5.0 module uses low-power broadcast technology to extend the device's battery life and enhance the signal coverage range; the ZigBee 3.0 module has self-organizing and self-healing functions to ensure network reliability in complex environments.
[0127] The beneficial effects of the above technical solution: The present invention establishes multiple communication connections between the embedded voice recognition control module and the lower-level devices through the device interaction system, so that the embedded voice recognition control module can be seamlessly connected and stably interact with various types of devices. It has strong compatibility and scalability, effectively expanding the scope of use of the embedded voice recognition control module while breaking the network's restrictions on voice interaction. Users can freely use voice commands to control the device whether the device is connected to the network or not.
[0128] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. An embedded speech recognition control module, characterized in that: include: A voice acquisition submodule consisting of one or more microphone arrays, used to collect the user's voice signal; The speech preprocessing submodule is used to preprocess the speech signal to obtain high-quality speech signals; The speech recognition and command parsing submodule is used to convert high-quality speech signals into text data based on a deep learning speech recognition model, and to perform semantic parsing of the text data through natural language processing technology to determine user control commands; The control instruction generation submodule is used to generate corresponding device control instructions based on user control commands and send them to the lower-level devices; The actuator and feedback submodule is used to generate corresponding feedback voice based on the control instruction execution results of the actuator of the lower device; Adaptive learning and personalization submodule, used to adaptively train the deep learning speech recognition model based on user text usage habits and accents; The voice collection submodule includes: The direction sensing unit is used to determine the source direction of each sound source signal based on the time difference and phase difference of the sound signals received by each microphone in the microphone array, and to determine the effective coverage range corresponding to each channel signal according to the size of the different channel signals received by each microphone; a sound data analysis unit, configured to determine an effective coverage area corresponding to a human voice signal as a target effective coverage range, and to confirm a target microphone array area corresponding to the target effective coverage range; Acquire a first target human voice signal collected by each microphone in each target microphone array area, perform target sound fusion based on the first target human voice signal to obtain a second target human voice signal, and send the second target human voice signal as the collected voice signal to the voice preprocessing submodule; a sensitivity control unit, configured to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value, obtain an acquisition error rate, and adjust the sensitivity of the microphone corresponding to each target human voice signal based on the acquisition error rate and the speaker position information; When any microphone is within the target microphone array area corresponding to two or more channel signals, the final sensitivity of the microphone is based on the sensitivity corresponding to the channel signal with greater sensitivity; The sensitivity control unit includes: The first intelligent control sub-unit is used to compare the total coverage area of all target effective coverage areas with the full coverage area of the microphone array, determine the vocal invalid area, and adjust the sensitivity value of the microphone corresponding to the vocal invalid area to a minimum threshold; The speaker localization subunit is used to obtain the distribution characteristics of the microphones in the target microphone array area corresponding to each human voice signal, and determine the relative distance between the speaker and the embedded speech recognition control module based on the distribution characteristics and the sound propagation characteristics; The second intelligent control subunit is configured to compare the relative distance with a preset standard value to obtain a position difference, and to compare the plurality of first target human voice signals contained in different second target human voice signals with a preset threshold value to obtain an acquisition error rate; Determine the microphone sensitivity adjustment direction and the corresponding adjustment coefficient based on the positive and negative corresponding position difference, and obtain the microphone sensitivity adjustment amplitude based on the acquisition error and the standard sensitivity of the microphone; An actual control value is obtained according to the sensitivity adjustment amplitude and the adjustment coefficient corresponding to the sensitivity adjustment direction, and the sensitivity of the target microphone is adjusted based on the actual control value and according to the sensitivity adjustment direction corresponding to the microphone.
2. The embedded speech recognition control module according to claim 1, characterized in that: Speech preprocessing submodule, including: A first preprocessing unit is used to perform preliminary processing on the speech signal based on Gaussian filtering and normalized least mean square algorithm to obtain a denoised speech signal; The second preprocessing unit is used to obtain the volume characteristics corresponding to each denoised speech signal, and perform volume normalization processing on the denoised speech signal based on the volume characteristics to obtain a high-quality speech signal.
3. The embedded speech recognition control module according to claim 1, characterized in that: Speech recognition and command parsing submodules include: The model recognition unit is used to identify high-quality voice signals based on a deep learning voice recognition model and convert them into text data. It also performs semantic analysis on the text data based on natural language technology to determine user semantics and user intent. The command screening unit is used to screen high-quality voice signals based on user semantics and user intent, combined with a preset control command data set, to obtain effective control voice; The command determination unit is used to determine the user control command according to the corresponding relationship between the effective control voice and the control vocabulary in the preset control instruction data set and the control instruction corresponding to the control vocabulary.
4. The embedded speech recognition control module according to claim 1, characterized in that: The control instruction generation submodule includes: An instruction generation unit, configured to generate corresponding control instructions based on user commands and in combination with the communication protocol between the embedded speech recognition control module and the lower-level device; The instruction sending unit is used to send control instructions to the lower-level device through a preset communication channel.
5. The embedded speech recognition control module according to claim 1, characterized in that: Adaptive learning and personalization submodules include: A user data processing unit is configured to obtain all valid historical wake-up records of the voice recognition control module, define the period from the initial wake-up to the successful recognition of the control command by the voice recognition control module as one wake-up cycle, divide the voice recognition control module into multiple successful control wake-up cycles, and obtain multiple successful control wake-up cycle data sets; Respectively obtaining invalid voice contents of the user in a plurality of successful control wake-up cycles, and when the number of invalid voice contents is zero, eliminating the successful control wake-up cycle; A data identification unit is configured to generate a corresponding instruction tag based on the successful control instruction corresponding to the successful voice wake-up cycle and add it to the successful control wake-up cycle data set, and monitor invalid voices, or true invalid voices and pseudo invalid voices, based on the keywords of the successful control instruction corresponding to the successful control wake-up cycle data set; A usage habit analysis unit is used to respectively obtain unique voice features corresponding to each pseudo-invalid voice, compare the unique voice features with the voice features of the standard control instruction corresponding to the control command, and obtain the accent difference between the user's actual voice and the pre-selected voice and the user's text usage habits; Based on the voiceprint features of the voice signals in different successful control wake-up cycles, the successful control wake-up cycles are clustered to obtain the voice control data clusters corresponding to the same user; Compare the accent differences corresponding to the pseudo-invalid speech within the same voice control data cluster and the user's text usage habits to obtain the current user's personalized accent and usage habits, and compare the personalized accents and usage habits of different users to obtain user voice similarity; a data merging unit, configured to merge voice control data clusters whose user voice similarity is greater than a preset value to obtain a merged data set, and generate different batches of training sets based on the merged data set and the unmerged voice control data clusters; The adaptive training unit is used to input the corresponding training set into the deep learning speech recognition model for adaptive training according to the training set batches until all batches of training are completed and the adaptive training is ended.
6. The embedded speech recognition control module according to claim 5, characterized in that: User data processing unit, including: The voiceprint file creation subunit is used to create a voiceprint file based on the voiceprint features corresponding to the voice signal. When a new voiceprint feature is detected, a new voiceprint file is automatically generated. When the current user's voice is confirmed to be a voiceprint object, determine whether the current user's voice control command is a single feedback command. If so, add an invalid label to the historical wake-up record of the current user; Otherwise, add a valid tag to the historical wakeup record of the current user.
7. The embedded speech recognition control module according to claim 6, characterized in that: The user data processing unit further includes: An intelligent comparison sub-unit is used to compare the current voice segment collected by the profiled user in the current successful control wake-up cycle with the historical voice segment in the invalid voiceprint file corresponding to the latest similar control command of the profiled user when a pseudo-invalid voice appears during the voice control command recognition process of the profiled user, and determine the usage habit and accent differences of the profiled user based on the comparison results; The audio corresponding to the current speech segment is highlighted according to the usage habit difference or the pronunciation difference, and a training control group is generated based on the highlighted speech segment and the historical speech segment. The training control group is used to perform adaptive training on the deep learning speech recognition model.
8. A device interaction system with an embedded voice recognition control module, characterized in that: A communication connection is established between the embedded speech recognition control module according to any one of claims 1 to 7 and its corresponding lower-level device.
Citation Information
Patent Citations
Interactive system for vehicle-mounted voice
CN101281745A
Interactive blind guiding system and method based on improved Yolov2 target detection and voice recognition
CN110728308A
Pickup sensitivity adjusting method and device
CN112015364A