Voice recognition device and method fusing mouth airflow, mouth shape and voice data

The speech recognition device integrates mouth airflow, mouth shape and voice data, and uses a head-mounted sensor array acquisition device and a neural network for feature fusion, thereby solving the problem of poor speech recognition in noisy environments, providing an effective language communication solution for users with impaired speech function, and achieving highly accurate and robust speech recognition.

CN120748409APending Publication Date: 2025-10-03SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511039261.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing speech recognition methods have poor recognition effects in noisy environments and cannot meet the language communication needs of users with impaired speech function.

Method used

A speech recognition device that integrates mouth airflow, mouth shape and voice data is used. Lip image information, airflow information and voice information are collected through a head-mounted sensor array acquisition device, and feature fusion and recognition are performed using a data processing device. Speech judgment is performed by combining a convolutional-recursive hybrid neural network and a recursive neural network.

Benefits of technology

It improves the accuracy of speech recognition in noisy environments, provides a non-speech-dependent language communication channel for users with impaired speech function, reduces environmental noise interference, improves the robustness of the recognition system and reduces the misjudgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748409A_ABST
    Figure CN120748409A_ABST
Patent Text Reader

Abstract

The invention relates to a voice recognition device and method fusing mouth airflow, mouth shape and voice data. The device comprises a head-mounted main body bracket, a sensor array acquisition device and a data processing device, the head-mounted main body bracket comprises a head-mounted bracket and a sensor bracket; the sensor array acquisition device comprises a support, an airflow acquisition module, an image acquisition module and a voice acquisition module; the airflow acquisition module is used for acquiring airflow information; the image acquisition module is used for acquiring lip image information; the voice acquisition module is used for acquiring voice information; and the data processing device is used for receiving the lip airflow information, the image information and the voice information, inputting the information into a voice recognition model arranged in the data processing device, obtaining a voice content recognition result of the user, and sending the voice content recognition result to the terminal equipment. The speech recognition device can effectively improve the accuracy of speech recognition in a noisy environment, and provides a non-speech-dependent alternative channel for the language communication of a user with a damaged speech function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and in particular to a speech recognition device and method that integrates mouth airflow, mouth shape and speech data. Background Art

[0002] With technological advancements, speech recognition technology has been widely used in many human-computer interaction scenarios. However, existing speech recognition methods mostly rely on sound signals in the air, which results in poor recognition in noisy environments. Furthermore, for those with impaired speech function, who cannot produce normal pronunciation, traditional speech recognition equipment cannot recognize their speech.

[0003] In summary, it is necessary to study a new speech recognition auxiliary method and device that can effectively improve the accuracy of speech recognition in noisy environments and provide a non-speech-dependent alternative channel for language communication for users with impaired speech function. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a speech recognition device that integrates mouth airflow, mouth shape and voice data. The speech recognition device can effectively improve the accuracy of speech recognition in noisy environments and provide a non-speech-dependent alternative channel for language communication for users with impaired speech function.

[0005] Another object of the present invention is to provide a speech recognition method that integrates mouth airflow, mouth shape and speech data.

[0006] The technical solution of the present invention to solve the above technical problems is:

[0007] A speech recognition device that integrates mouth airflow, mouth shape and speech data includes a head-mounted main body support, a sensor array acquisition device arranged on the head-mounted main body support, and a data processing device, wherein:

[0008] The head-mounted main body support includes a head-mounted support and a sensor support arranged on the head-mounted support, wherein the sensor support is rotatably connected to the head-mounted support; the data processing device is installed on the head-mounted support; and the sensor array acquisition device is installed on the sensor support;

[0009] The sensor array acquisition device includes a support provided on the sensor bracket, an airflow acquisition module provided on the support, an image acquisition module, and a voice acquisition module, wherein the image acquisition module is used to acquire lip image information of the user's mouth shape; the airflow acquisition module is used to acquire airflow information exhaled by the user; and the voice acquisition module is used to acquire voice information;

[0010] The data processing device is used to receive lip image information, airflow information and voice information, and input them into its built-in voice recognition model to obtain the user's voice content recognition data, and transmit it to the terminal device through the built-in wireless transmission module.

[0011] Preferably, the head-mounted bracket includes a headband and ear hooks arranged on both sides of the headband, wherein the headband is an elastic arc-shaped structure, which spans the user's head area, and its two ends extend above the user's ears, and are respectively connected to the ear hooks; a speaker is provided on the ear hook; an angle adjustment mechanism for adjusting the angle of the sensor bracket is provided between the head-mounted bracket and the sensor bracket, and the angle adjustment mechanism includes an adjustment knob arranged on the outside of the ear hook; the adjustment knob is rotatably connected to the ear hook, and the sensor bracket is fixed on the adjustment knob.

[0012] Preferably, the support comprises a cylindrical base and a support column arranged on the cylindrical base, wherein the image acquisition module, the airflow acquisition module and the voice acquisition module are all installed on the support column, wherein,

[0013] The image acquisition device is located at the upper end of the support column, and includes a mounting bracket provided on the support column and a camera module provided on the mounting bracket, wherein an angle adjustment module for adjusting the angle of the camera module is provided between the camera module and the mounting bracket;

[0014] The voice collection device is located at the lower end of the support column, and includes a flexible silicone tube provided on the support column and a microphone module provided at the end of the flexible silicone tube, wherein one end of the flexible silicone tube is mounted on the cylindrical base, and the other end is connected to the microphone module;

[0015] The airflow collection device is located between the image collection device and the voice collection device, and is coaxially arranged with the cylindrical base. The airflow collection device includes an external airflow blocking structure for blocking external airflow and a pine cone sensor array structure arranged in the external airflow blocking structure, wherein the external airflow blocking structure includes a truncated cone-shaped shell and a bottom crossbeam support member arranged on the truncated cone-shaped shell, wherein the bottom crossbeam support members are in multiple groups, and the multiple groups of bottom crossbeam support members are arranged in a circle; each group of bottom crossbeam support members is provided with a threaded hole; the truncated cone-shaped shell is fixed to the support column after a screw is passed through the threaded hole; the lower side of the truncated cone-shaped shell is provided with a gap extending along its axial direction.

[0016] Preferably, the pine cone-shaped sensor array structure includes a central main axis and a first-layer sensor mounting platform, a second-layer sensor mounting platform, and a third-layer sensor mounting platform arranged on the central main axis, wherein the first-layer sensor mounting platform is located at the end of the central main axis close to the user; the third-layer sensor mounting platform is located at the end of the central main axis away from the user; the first-layer sensor mounting platform is located between the first-layer sensor mounting platform and the third-layer sensor mounting platform, wherein,

[0017] The first layer sensor mounting platform includes three groups of first long strip brackets; the three groups of first long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the gaps between the three groups of first long strip brackets form a triangular pyramid structure; the first long strip brackets are oriented at an angle of 45 degrees to the central main axis;

[0018] The second-layer sensor mounting platform includes three groups of second long strip brackets; the three groups of second long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the distribution angle of the second long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the second long strip brackets and the central main axis is 67.5 degrees;

[0019] The second-layer sensor mounting platform includes three groups of second long strip brackets; the three groups of second long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the distribution angle of the second long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the second long strip brackets and the central main axis is 67.5 degrees;

[0020] The third-layer sensor mounting platform includes six groups of third long strip brackets; the six groups of third long strip brackets are evenly distributed circumferentially at intervals of 60 degrees in the circumferential direction of the central main axis; the distribution angle of the third long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the third long strip brackets and the central main axis is 0 degrees;

[0021] The first long strip bracket, the second long strip bracket and the third long strip bracket are all provided with air pressure sensors.

[0022] A speech recognition method that integrates mouth airflow, mouth shape, and speech data comprises the following steps:

[0023] Wearing the voice recognition device and detecting whether the user is wearing the voice recognition device correctly. If it is detected that the user is not wearing the voice recognition device correctly, reminding the user to adjust the wearing posture until the user is detected to be wearing the voice recognition device correctly;

[0024] The voice acquisition module, the image acquisition module, and the airflow acquisition module respectively collect the user's voice information, the lip image information, and the airflow information, and transmit the collected voice information, the lip image information, and the airflow information to the data processing device;

[0025] The data processing device pre-processes the voice information, lip image information, and airflow information, and performs time-series alignment on voice information, lip image information, and airflow information of different lengths. The device then performs matching analysis on the time-series aligned voice information, lip image information, and airflow information to determine whether the user is speaking. If it is determined that the user is speaking, the corresponding voice information, lip image information, and airflow information are saved. The voice information, lip image information, and airflow information are then windowed according to a predetermined time period, and feature fusion is performed on the voice information, lip image information, and airflow information in each window to obtain fused features. The fused features are input into a trained voice recognition model to output a voice content recognition result.

[0026] The voice content recognition structure is transmitted to the terminal device through the wireless transmission module.

[0027] Preferably, the step of determining whether the user is speaking is:

[0028] Airflow information, lip image information, and voice information are collected in real time using an airflow acquisition device, an image acquisition device, and a voice acquisition device. These information are then preprocessed. After preprocessing, features are extracted from the airflow information, lip image information, and voice information. The extracted features are input into corresponding neural network classifiers to obtain speech determination results based on different data sources. These speech determination results are output as a binary output of "speaking / not speaking";

[0029] The speech determination results output by each neural network classifier are aligned in time windows, wherein a fixed time window is used as the alignment period, and the number of valid samples in each modal data and the number of times the judgment result is speaking are counted within the alignment period, and the ratio of the number of valid samples to the number of times speaking is calculated; if the ratio of the number of valid samples to the number of times speaking of all modal data exceeds the set ratio threshold, a speaking state confirmation signal is output; if the ratio of the number of valid samples to the number of times speaking of any modal data is lower than the set ratio threshold, a silent state confirmation signal is output.

[0030] Preferably, the step of performing speech judgment on the lip image information by using a neural network classifier is:

[0031] Perform frame splitting, noise suppression, and image normalization on the video information collected by the image acquisition module, and output a standardized lip image sequence;

[0032] A convolutional-recursive hybrid neural network architecture is used for lip shape recognition. The convolutional-recursive hybrid neural network architecture consists of a convolutional unit and a recursive unit. The convolutional unit encodes a single-frame lip image into a fixed-dimensional feature vector to capture the spatial characteristics of the lips. The recursive unit is used to capture the temporal dependencies between consecutive frames and output an output vector that integrates spatiotemporal information.

[0033] Based on the output vector, the feature vector is nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

[0034] Preferably, the step of performing speech judgment on the voice information by using a neural network classifier is:

[0035] The voice collecting device collects voice signals within a predetermined time period, and performs frame processing on the collected voice signals to generate time-series voice segments;

[0036] Extracting time-frequency domain feature vectors from the time-series speech segments, and performing sliding window smoothing and normalization operations on the time-frequency domain feature vectors;

[0037] A preconfigured recursive neural network architecture is used to learn the temporal dependencies of the time-frequency domain feature vectors. The time-frequency domain feature vectors are nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

[0038] Preferably, the step of performing speech judgment on the airflow information by using a neural network classifier is:

[0039] The airflow acquisition device acquires time series data of air pressure, flow velocity, and airflow direction in real time; a sliding window segmentation operation is performed on the acquired time series data, and time domain statistical features and frequency domain energy distribution features are extracted; the time domain statistical features and frequency domain energy distribution features are input into a pre-built recursive neural network classifier, and the long-term dependencies are captured through its memory unit to output the airflow type classification result; based on the probability distribution of the recursive neural network classifier, it is determined whether the current airflow belongs to one of the categories of speech airflow, respiratory airflow, or cough airflow.

[0040] Preferably, the step of performing feature fusion on the voice information, lip image information and airflow information is:

[0041] The lip image information is cached and then normalized, resized, and feature enhanced. The airflow information is low-pass filtered and outliers are removed. The speech information is standardized.

[0042] A two-dimensional convolutional neural network is used to extract spatiotemporal features of lip movements from image information; a one-dimensional convolutional neural network is used to extract acoustic features from speech information; and a recurrent neural network is used to extract dynamic airflow pattern features from airflow information; the two-dimensional convolutional neural network, the one-dimensional convolutional neural network, and the recurrent neural network output a feature vector of uniform dimension.

[0043] Different weight coefficients are assigned to the spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement, and the spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement are spliced ​​using the feature splicing method to obtain multimodal fusion features;

[0044] The multimodal fusion features are input into a pre-trained speech recognition model; the speech recognition model contains at least one layer of long short-term memory units and gated recurrent units; the long-term contextual associations in the multimodal fusion features are extracted by the long short-term memory units, and the gated recurrent units are then used to strengthen the modeling of local temporal dependencies; after the speech recognition model processes the multimodal fusion features time step by time, it completes nonlinear mapping through a fully connected layer, and finally outputs the probability distribution of the speech content through a softmax activation function, and the text corresponding to the maximum probability is used as the speech content recognition result.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. The speech recognition device of the present invention, which integrates mouth airflow, mouth shape and voice data, is modeled with the help of an airflow acquisition device, an image acquisition device and a voice acquisition device, and uses an artificial intelligence algorithm to improve the accuracy of the recognition task; at the same time, the speech recognition method using speaking airflow makes the speech recognition device of the present invention less affected by environmental noise and can help user groups with impaired pronunciation function.

[0047] 2. The speech recognition device of the present invention, which integrates mouth airflow, mouth shape and voice data, matches the speaking airflow with the speaking content. By using the airflow information, lip image information and voice information during speaking as multimodal fusion features for speech recognition, it can effectively cope with external interference such as noisy environments and light changes, thereby constructing a speech recognition system with stronger robustness and lower error rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 Schematic diagram of the structure of the speech recognition device of the present invention that integrates mouth airflow, mouth shape and speech data.

[0049] Figure 2 This is the main view of the sensor array acquisition device.

[0050] Figure 3 This is the right side view of the sensor array acquisition device.

[0051] Figure 4 This is a structural diagram of the pine cone-shaped sensor array structure.

[0052] Figure 5 Schematic diagram of the flow of the speech recognition method of the present invention that integrates mouth airflow, mouth shape and speech data. DETAILED DESCRIPTION

[0053] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0054] See also Figures 1-4 The speech recognition device of the present invention that integrates mouth airflow, mouth shape and speech data includes a head-mounted main body support, a sensor array acquisition device and a data processing device arranged on the head-mounted main body support, wherein:

[0055] The head-mounted main support includes a head-mounted support (similar to an earphone structure) and a sensor support 3 provided on the head-mounted support, wherein the sensor support 3 is rotatably connected to the head-mounted support; the data processing device is installed on the head-mounted support; and the sensor array acquisition device is installed on the sensor support 3;

[0056] The sensor array acquisition device includes a support provided on the sensor bracket 3, an airflow acquisition module provided on the support, an image acquisition module, and a voice acquisition module, wherein the image acquisition module is used to acquire lip image information of the user's mouth shape; the airflow acquisition module is used to acquire airflow information exhaled by the user; and the voice acquisition module is used to acquire voice information;

[0057] The data processing device is connected to the sensor array acquisition device via a multi-channel flat communication cable, and both ends of the multi-channel flat communication cable are reinforced and stabilized to ensure a stable and reliable connection. The data processing device is used to receive lip image information, airflow information, and voice information, and input them into its built-in voice recognition model to obtain user voice content recognition data, and transmit it to the terminal device via a built-in wireless transmission module;

[0058] In this embodiment, the data processing device includes a single-chip microcomputer, a memory, a power supply module and a communication cable module, wherein the single-chip microcomputer adopts a high-performance single-chip microcomputer and can run a real-time operating system of multi-threaded tasks, thereby realizing the synchronous collection and rapid processing of airflow, image and voice data; the memory has a large storage capacity and can open up a large buffer space for temporary storage of data.

[0059] See also Figures 1-4The head-mounted bracket includes a headband 1 and ear hooks 4 arranged on both sides of the headband 1, wherein the headband 1 is an elastic arc-shaped structure, which spans the user's head area, and its two ends extend above the user's ears respectively, and are connected to the ear hooks 4 respectively; a speaker 5 is provided on the ear hook 4; the sensor bracket 3 is a suspension structure, and the sensor bracket 3 is a rigid tubular structure, and a channel for the communication cable to pass through is provided in its lumen; an angle adjustment mechanism for adjusting the angle of the sensor bracket 3 is provided between the head-mounted bracket and the sensor bracket 3, and the angle adjustment mechanism includes an adjustment knob 2 arranged on the outside of the ear hook 4; the adjustment knob 2 is rotatably connected to the ear hook 4, and the sensor bracket 3 is fixed on the adjustment knob 2; the sensor bracket 3 can be driven to rotate by the adjustment knob 2 to adjust the spatial angle of the sensor bracket 3.

[0060] See also Figures 1-4 The support includes a cylindrical base 6 and a support column 7 provided on the cylindrical base 6; the support column 7 is installed on the cylindrical base 6 and is coaxially arranged with the cylindrical base 6; the image acquisition module, the airflow acquisition module and the voice acquisition module are all installed on the support column 7, wherein,

[0061] The image acquisition device 8 is located at the upper end of the support column 7, and includes a mounting bracket provided on the support column 7 and a camera module provided on the mounting bracket. The camera module has a small overall volume and is easy to integrate and fix. An angle adjustment module for adjusting the angle of the camera module is provided between the camera module and the mounting bracket, for example, by changing the shooting angle of the camera module through a knob mechanism.

[0062] The voice collection device is located at the lower end of the support column 7 and includes a flexible silicone tube 9 provided on the support column 7 and a microphone module 10 provided at the end of the flexible silicone tube 9, wherein one end of the flexible silicone tube 9 is mounted on the cylindrical base 6 and the other end is connected to the microphone module 10; the plasticity and deformation characteristics of the flexible silicone tube 9 ensure that the distance between the microphone module 10 and the user's mouth is at an appropriate distance;

[0063] The airflow collection device 11 is located between the image collection device 8 and the voice collection device and is coaxially arranged with the cylindrical base 6. The airflow collection device 11 includes an external airflow blocking structure 12 for blocking external airflow and a pine cone sensor array structure 13 arranged in the external airflow blocking structure 12, wherein the external airflow blocking structure 12 includes a truncated cone-shaped shell 123 and a bottom crossbeam support member 122 arranged on the truncated cone-shaped shell 123, wherein the truncated cone-shaped shell 123 is coaxially arranged with the support column 7, and the end with a larger diameter is installed on the support column 7, and the end with a smaller diameter is installed on the support column 7. Then it faces the user; the bottom cross beam support members 122 are multiple groups, and the multiple groups of bottom cross beam support members 122 are arranged in a circle; each group of bottom cross beam support members 122 is provided with a threaded hole 121; the truncated cone-shaped shell 123 is fixed to the support column 7 after the screw is passed through the threaded hole 121; the lower side of the truncated cone-shaped shell 123 is provided with a gap extending along its axial direction; the gap is used to provide an airflow diffusion channel connected to the outside world, that is, the airflow exhaled by the user can only flow out to the outside through the gap, which can avoid airflow retention and interference, and isolate the external airflow in other directions, such as isolating the nasal airflow from the external environment.

[0064] See also Figures 1-4 The pine cone-shaped sensor array structure 13 includes a central main axis 19 and a first-layer sensor mounting platform 16, a second-layer sensor mounting platform 17, a third-layer sensor mounting platform 18 and an airflow diversion device 15 arranged on the central main axis 19, wherein the first-layer sensor mounting platform 16 is located at the end of the central main axis 19 close to the user; the third-layer sensor mounting platform 18 is located at the end of the central main axis 19 away from the user; the first-layer sensor mounting platform 16 is located between the first-layer sensor mounting platform 16 and the third-layer sensor mounting platform 18; the spacing between two adjacent groups of sensor mounting platforms is 2 cm; the airflow diversion device 15 is a triangular pyramid structure, and the angle between its adjacent surfaces is 60°; wherein,

[0065] The first sensor mounting platform 16 includes three groups of first long strip brackets; the three groups of first long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis 19; the gaps between the three groups of first long strip brackets form a triangular pyramid structure; the first long strip brackets are oriented at an angle of 45 degrees to the central main axis 19;

[0066] The second sensor mounting platform 17 includes three groups of second elongated brackets; the three groups of second elongated brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis 19; the distribution angle of the second elongated brackets is offset by 60 degrees relative to the distribution angle of the first elongated brackets, and the angle between the orientation of the second elongated brackets and the central main axis 19 is 67.5 degrees;

[0067] The second sensor mounting platform 17 includes three groups of second elongated brackets; the three groups of second elongated brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis 19; the distribution angle of the second elongated brackets is offset by 60 degrees relative to the distribution angle of the first elongated brackets, and the angle between the orientation of the second elongated brackets and the central main axis 19 is 67.5 degrees;

[0068] The third sensor mounting platform 18 includes six groups of third elongated brackets; the six groups of third elongated brackets are evenly distributed circumferentially at intervals of 60 degrees in the circumferential direction of the central main axis 19; the distribution angle of the third elongated brackets is offset by 60 degrees relative to the distribution angle of the first elongated brackets, and the angle between the orientation of the third elongated brackets and the central main axis 19 is 0 degrees;

[0069] In this embodiment, the first elongated bracket, the second elongated bracket and the third elongated bracket are all provided with grooves; the size of the groove is slightly larger than the size of the air pressure sensor, so that the air pressure sensor module can be embedded in the groove, and a shell protection structure is placed above the groove to fix the air pressure sensor; holes are left on the shell to guide the external air pressure changes into the air pressure sensor; in addition, the sizes of the first elongated bracket, the second elongated bracket and the third elongated bracket increase successively in the direction away from the user; the central main shaft 19 is also a stepped shaft structure, and three step surfaces with increasing outer diameters are provided on the central main shaft 19 in the direction away from the user; the first layer sensor mounting platform 16, the second layer sensor mounting platform 17, and the third layer sensor mounting platform 18 are respectively arranged on the corresponding shaft segments; a hollow channel 14 is provided inside the central main shaft 19 to shuttle the communication cable of the air pressure sensor;

[0070] By adopting the above structure, the airflow collection device 11 of the present invention can utilize the characteristics of different sizes of each layer of the conical structure and combine the sensing signal of the air pressure sensor to achieve multi-plane respiratory airflow signal measurement with little interference to the airflow.

[0071] See also Figure 1-Figure 5 The speech recognition method of the present invention, which integrates mouth airflow, mouth shape and speech data, comprises the following steps:

[0072] Wearing the voice recognition device and detecting whether the user is wearing the voice recognition device correctly. If it is detected that the user is not wearing the voice recognition device correctly, reminding the user to adjust the wearing posture until the user is detected to be wearing the voice recognition device correctly;

[0073] In this embodiment, the head-mounted main body bracket is worn on the user's head, and the position of the sensor bracket 3 is adjusted by the adjustment knobs 2 on both sides of the head-mounted bracket so that the airflow acquisition device is in full contact with the user's respiratory airflow; after the device is worn, the wearing quality is automatically detected, specifically: the image acquisition device 8 starts to acquire lip image information and transmits the lip image information to the data processing device; in the data processing device, the relative distance between the image center of the lip image information and the mouth contour is determined, and the distance is used to determine whether the orientation of the image acquisition device 8 at this time is appropriate. If the relative distance is less than the set value, it is considered that the wearing quality is qualified. Otherwise, a voice is played through the speaker 5 to remind the user to make adjustments; when it is detected that the wearing quality is qualified, the data acquisition program is started and voice recognition begins.

[0074] The voice acquisition module, the image acquisition module, and the airflow acquisition module respectively collect the user's voice information, the lip image information, and the airflow information, and transmit the collected voice information, the lip image information, and the airflow information to the data processing device;

[0075] The data processing device pre-processes the voice information, lip image information and airflow information, and performs time-series alignment on the voice information, lip image information and airflow information of different lengths, performs matching analysis on the voice information, lip image information and airflow information after time-series alignment, and determines whether the user is speaking. If it is determined that the user is speaking, the time period of the user's speaking is recorded, and the multimodal data within the speaking time period is intercepted from the data stream, and the corresponding voice information, lip image information and airflow information are saved; then the voice information, lip image information and airflow information are windowed according to a predetermined time, and the voice information, lip image information and airflow information of each divided window are subjected to feature fusion to obtain fused features, and the fused features are input into the trained speech recognition model to output the speech content recognition results; if it is determined that the user is not speaking, the collected multimodal data can be deleted, and the subsequent speech recognition work is not performed;

[0076] In the above process, an update threshold may be set, and when the time period is greater than the set update threshold, a speech recognition result is output.

[0077] The voice content recognition structure is transmitted to the terminal device through the wireless transmission module.

[0078] In this embodiment, the image acquisition device 8 can acquire video information of the lips and obtain lip image information through the video information.

[0079] See also Figure 1-Figure 5 , the steps to determine whether the user is speaking are:

[0080] Airflow information, lip image information, and voice information are collected in real time using an airflow acquisition device, an image acquisition device, and a voice acquisition device. These information are then preprocessed. After preprocessing, features are extracted from the airflow information, lip image information, and voice information. The extracted features are input into corresponding neural network classifiers to obtain speech determination results based on different data sources. These speech determination results are output as a binary output of "speaking / not speaking";

[0081] The speech determination results output by each neural network classifier are aligned in time windows, wherein a fixed time window is used as the alignment period, and the number of valid samples in each modal data and the number of times the judgment result is speaking are counted within the alignment period, and the ratio of the number of valid samples to the number of times speaking is calculated; if the ratio of the number of valid samples to the number of times speaking of all modal data exceeds the set ratio threshold, a speaking state confirmation signal is output; if the ratio of the number of valid samples to the number of times speaking of any modal data is lower than the set ratio threshold, a silent state confirmation signal is output.

[0082] The steps of using a neural network classifier to perform speech judgment on lip image information are as follows:

[0083] Perform frame splitting, noise suppression, and image normalization on the video information collected by the image acquisition module, and output a standardized lip image sequence;

[0084] A convolutional-recursive hybrid neural network architecture is used for lip shape recognition. The convolutional-recursive hybrid neural network architecture consists of a convolutional unit and a recursive unit. The convolutional unit encodes a single-frame lip image into a fixed-dimensional feature vector to capture the spatial characteristics of the lips. The recursive unit is used to capture the temporal dependencies between consecutive frames and output an output vector that integrates spatiotemporal information.

[0085] Based on the output vector, the feature vector is nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

[0086] The steps of using a neural network classifier to judge speech information are as follows:

[0087] The voice collecting device collects voice signals within a predetermined time period, and performs frame processing on the collected voice signals to generate time-series voice segments;

[0088] Extracting time-frequency domain feature vectors from the time-series speech segments, and performing sliding window smoothing and normalization operations on the time-frequency domain feature vectors;

[0089] A preconfigured recursive neural network architecture is used to learn the temporal dependencies of the time-frequency domain feature vectors. The time-frequency domain feature vectors are nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

[0090] The steps of using a neural network classifier to perform speech judgment on airflow information are as follows:

[0091] The airflow acquisition device 11 is used to obtain time series data of air pressure, flow velocity, and airflow direction in real time; a sliding window segmentation operation is performed on the obtained time series data, and time domain statistical features and frequency domain energy distribution features are extracted; the time domain statistical features and frequency domain energy distribution features are input into a pre-built recursive neural network classifier, and the long-term dependencies are captured through its memory unit to output the airflow type classification result; based on the probability distribution of the recursive neural network classifier, it is determined whether the current airflow belongs to one of the categories of speech airflow, respiratory airflow, or cough airflow.

[0092] See also Figure 1-Figure 5 ,The steps of feature fusion of speech information, lip image information and airflow information are;

[0093] The lip image information is cached and then normalized, resized, and feature enhanced. The airflow information is low-pass filtered and outliers are removed. The speech information is standardized.

[0094] A two-dimensional convolutional neural network is used to extract spatiotemporal features of lip movements from image information; a one-dimensional convolutional neural network is used to extract acoustic features from speech information; and a recurrent neural network is used to extract dynamic airflow pattern features from airflow information; the two-dimensional convolutional neural network, the one-dimensional convolutional neural network, and the recurrent neural network output a feature vector of uniform dimension.

[0095] In order to increase the weight of factors with greater influence, different weight coefficients are assigned to the spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement. The spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement are then spliced ​​together using the feature splicing method to obtain multimodal fusion features.

[0096] In this embodiment, different weight coefficients are assigned to the spatiotemporal characteristics, acoustic characteristics and airflow dynamic pattern characteristics of lip movement, and the sum of the products of the spatiotemporal characteristics, acoustic characteristics and airflow dynamic pattern characteristics of lip movement and the corresponding weight coefficients are obtained to obtain multimodal fusion features; the initial values ​​of the weight coefficients are all 1, and can be flexibly adjusted according to actual conditions to meet the operation accuracy requirements of different environments.

[0097] The multimodal fusion features are input into a pre-trained speech recognition model; the speech recognition model contains at least one layer of long short-term memory units and gated recurrent units; the long-term contextual associations in the multimodal fusion features are extracted by the long short-term memory units, and the gated recurrent units are then used to strengthen the modeling of local temporal dependencies; after the speech recognition model processes the multimodal fusion features time step by time, it completes nonlinear mapping through a fully connected layer, and finally outputs the probability distribution of the speech content through a softmax activation function, and the text corresponding to the maximum probability is used as the speech content recognition result.

[0098] The above is a preferred embodiment of the present invention, but the embodiment of the present invention is not limited to the above content. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A speech recognition device that integrates mouth airflow, mouth shape and voice data, characterized in that: It includes a head-mounted main body support, a sensor array acquisition device and a data processing device arranged on the head-mounted main body support, wherein: The head-mounted main body support includes a head-mounted support and a sensor support arranged on the head-mounted support, wherein the sensor support is rotatably connected to the head-mounted support; the data processing device is installed on the head-mounted support; and the sensor array acquisition device is installed on the sensor support; The sensor array acquisition device includes a support provided on the sensor bracket, an airflow acquisition module provided on the support, an image acquisition module, and a voice acquisition module, wherein the airflow acquisition module is used to acquire airflow information exhaled by the user; the image acquisition module is used to acquire lip image information of the user's mouth shape; and the voice acquisition module is used to acquire voice information; The data processing device is used to receive mouth airflow information, image information and voice information, and input them into its built-in voice recognition model to obtain the user's voice content recognition data, and transmit it to the terminal device through the built-in wireless transmission module.

2. The speech recognition device according to claim 1, wherein: The head-mounted bracket includes a headband and ear hooks arranged on both sides of the headband, wherein the headband is an elastic arc-shaped structure, which spans the user's head area, and its two ends extend above the user's ears respectively and are connected to the ear hooks respectively; a speaker is provided on the ear hook; an angle adjustment mechanism for adjusting the angle of the sensor bracket is provided between the head-mounted bracket and the sensor bracket, and the angle adjustment mechanism includes an adjustment knob arranged on the outside of the ear hook; the adjustment knob is rotatably connected to the ear hook, and the sensor bracket is fixed on the adjustment knob.

3. The speech recognition device according to claim 2, wherein: The support includes a cylindrical base and a support column arranged on the cylindrical base, wherein the airflow acquisition module, the image acquisition module and the voice acquisition module are all installed on the support column, wherein, The airflow collection device is located between the image collection device and the voice collection device, and is coaxially arranged with the cylindrical base. The airflow collection device includes an external airflow blocking structure for blocking external airflow and a pine cone sensor array structure arranged in the external airflow blocking structure, wherein the external airflow blocking structure includes a truncated cone-shaped shell and a bottom crossbeam support member arranged on the truncated cone-shaped shell, wherein the bottom crossbeam support members are in multiple groups, and the multiple groups of bottom crossbeam support members are arranged in a circle; each group of bottom crossbeam support members is provided with a threaded hole; the truncated cone-shaped shell is fixed to the support column after a screw is passed through the threaded hole; the lower side of the truncated cone-shaped shell is provided with a gap extending along its axial direction. The image acquisition device is located at the upper end of the support column, and includes a mounting bracket provided on the support column and a camera module provided on the mounting bracket, wherein an angle adjustment module for adjusting the angle of the camera module is provided between the camera module and the mounting bracket; The voice collection device is located at the lower end of the support column, and includes a flexible silicone tube arranged on the support column and a microphone module arranged at the end of the flexible silicone tube, wherein one end of the flexible silicone tube is installed on the cylindrical base, and the other end is connected to the microphone module.

4. The speech recognition device according to claim 2, wherein: The pine cone-shaped sensor array structure includes a central main axis and a first-layer sensor mounting platform, a second-layer sensor mounting platform, and a third-layer sensor mounting platform arranged on the central main axis, wherein the first-layer sensor mounting platform is located at the end of the central main axis close to the user; the third-layer sensor mounting platform is located at the end of the central main axis away from the user; the first-layer sensor mounting platform is located between the first-layer sensor mounting platform and the third-layer sensor mounting platform, wherein, The first layer sensor mounting platform includes three groups of first long strip brackets; the three groups of first long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the gaps between the three groups of first long strip brackets form a triangular pyramid structure; the first long strip brackets are oriented at an angle of 45 degrees to the central main axis; The second-layer sensor mounting platform includes three groups of second long strip brackets; the three groups of second long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the distribution angle of the second long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the second long strip brackets and the central main axis is 67.5 degrees; The second-layer sensor mounting platform includes three groups of second long strip brackets; the three groups of second long strip brackets are evenly distributed circumferentially at intervals of 120 degrees in the circumferential direction of the central main axis; the distribution angle of the second long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the second long strip brackets and the central main axis is 67.5 degrees; The third-layer sensor mounting platform includes six groups of third long strip brackets; the six groups of third long strip brackets are evenly distributed circumferentially at intervals of 60 degrees in the circumferential direction of the central main axis; the distribution angle of the third long strip brackets is offset by 60 degrees relative to the distribution angle of the first long strip brackets, and the angle between the orientation of the third long strip brackets and the central main axis is 0 degrees; The first long strip bracket, the second long strip bracket and the third long strip bracket are all provided with air pressure sensors.

5. A speech recognition method that integrates mouth airflow, mouth shape and voice data, characterized in that: The following steps are involved: Wearing the voice recognition device and detecting whether the user is wearing the voice recognition device correctly. If it is detected that the user is not wearing the voice recognition device correctly, reminding the user to adjust the wearing posture until the user is detected to be wearing the voice recognition device correctly; The airflow acquisition module, the image acquisition module, and the voice acquisition module respectively collect the user's airflow information, lip image information, and voice information, and transmit the collected airflow information, lip image information, and voice information to the data processing device; The data processing device pre-processes the airflow information, lip image information, and voice information, and performs time-series alignment on the airflow information, lip image information, and voice information of different lengths. The device then performs matching analysis on the time-series aligned airflow information, lip image information, and voice information to determine whether the user is speaking. If it is determined that the user is speaking, the corresponding airflow information, lip image information, and voice information are saved. The airflow information, lip image information, and voice information are then windowed according to a predetermined time period, and feature fusion is performed on the airflow information, lip image information, and voice information in each window to obtain a fused feature. The fused feature is input into a trained speech recognition model to output a speech content recognition result. The voice content recognition structure is transmitted to the terminal device through the wireless transmission module.

6. The speech recognition method according to claim 5, wherein: The steps to determine whether the user is speaking are: The airflow information, lip image information and voice information are collected in real time by the airflow collection device, the image collection device and the voice collection device, and the airflow information, the lip image information and the voice information are preprocessed respectively; After preprocessing, feature extraction is performed on airflow information, lip image information, and speech information. The extracted features are input into the corresponding neural network classifier to obtain speech determination results based on different data sources. The speech determination results are output as a binary output of "speaking / not speaking"; The speech determination results output by each neural network classifier are aligned in time windows, wherein a fixed time window is used as the alignment period, and the number of valid samples in each modal data and the number of times the judgment result is speaking are counted within the alignment period, and the ratio of the number of valid samples to the number of times speaking is calculated; if the ratio of the number of valid samples to the number of times speaking of all modal data exceeds the set ratio threshold, a speaking state confirmation signal is output; if the ratio of the number of valid samples to the number of times speaking of any modal data is lower than the set ratio threshold, a silent state confirmation signal is output.

7. The speech recognition method according to claim 6, wherein: The steps for using a neural network classifier to determine speech from lip image information are as follows: Perform frame splitting, noise suppression, and image normalization on the video information collected by the image acquisition module, and output a standardized lip image sequence; A convolutional-recursive hybrid neural network architecture is used for lip shape recognition. The convolutional-recursive hybrid neural network architecture consists of a convolutional unit and a recursive unit. The convolutional unit encodes a single-frame lip image into a fixed-dimensional feature vector to capture the spatial characteristics of the lips. The recursive unit is used to capture the temporal dependencies between consecutive frames and output an output vector that integrates spatiotemporal information. Based on the output vector, the feature vector is nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

8. The speech recognition method according to claim 6, wherein: The steps for judging speech information through a neural network classifier are as follows: The voice collecting device collects voice signals within a predetermined time period, and performs frame processing on the collected voice signals to generate time-series voice segments; Extracting time-frequency domain feature vectors from the time-series speech segments, and performing sliding window smoothing and normalization operations on the time-frequency domain feature vectors; A preconfigured recursive neural network architecture is used to learn the temporal dependencies of the time-frequency domain feature vectors. The time-frequency domain feature vectors are nonlinearly mapped through a fully connected classification layer to generate a binary probability distribution of speaking / non-speaking states, and the state corresponding to the maximum probability is used as the final recognition result.

9. The speech recognition method according to claim 6, wherein: The steps for judging the speech of airflow information through the neural network classifier are as follows: The air pressure, flow velocity and air flow direction time series data are acquired in real time through an airflow acquisition device; a sliding window segmentation operation is performed on the acquired time series data, and time domain statistical features and frequency domain energy distribution features are extracted; the time domain statistical features and frequency domain energy distribution features are input into a pre-built recursive neural network classifier, and the long-term dependencies are captured through its memory unit to output the airflow type classification result; according to the probability distribution of the recursive neural network classifier, the state corresponding to the maximum probability is used as the final recognition result.

10. The speech recognition method according to claim 6, wherein: The steps of feature fusion of speech information, lip image information and airflow information are as follows: The lip image information is cached and then normalized, resized, and feature enhanced. The airflow information is low-pass filtered and outliers are removed. Standardize voice information; A two-dimensional convolutional neural network is used to extract spatiotemporal features of lip movements from image information; a one-dimensional convolutional neural network is used to extract acoustic features from speech information; and a recurrent neural network is used to extract dynamic airflow pattern features from airflow information; the two-dimensional convolutional neural network, the one-dimensional convolutional neural network, and the recurrent neural network output a feature vector of uniform dimension. Different weight coefficients are assigned to the spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement, and the spatiotemporal features, acoustic features, and airflow dynamic pattern features of lip movement are spliced ​​using the feature splicing method to obtain multimodal fusion features; Inputting the multimodal fusion features into a pre-trained speech recognition model; the speech recognition model comprises at least one layer of long short-term memory units and gated recurrent units; extracting long-term contextual associations from the multimodal fusion features through the long short-term memory units, and then strengthening the modeling of local temporal dependencies through the gated recurrent units; After processing the multimodal fusion features time-step by time, the speech recognition model completes nonlinear mapping through the fully connected layer, and finally outputs the probability distribution of the speech content through the softmax activation function, and the text corresponding to the maximum probability is used as the speech content recognition result.